Company

Hugging Face

別名: Hugging Face

huggingface.co

Overview

最終更新: 2026年7月11日

Hugging Faceは、2016年にニューヨーク州ブルックリンで設立されたアメリカ合衆国の企業であり、機械学習モデルやデータセットの共有プラットフォームの運営を中核事業とする。自然言語処理(NLP)を起点に発展し、現在は画像・音声・マルチモーダルを含む幅広いAIモデルのホスティングとオープンソースライブラリの提供で知られる。子会社にHugging Face SASを持つ。

概要

Hugging Faceのプラットフォームは、研究者や開発者が事前学習済みモデルをアップロード・ダウンロードし、データセットを共有できる「モデルハブ」としての機能を中心に構築されている。Transformersライブラリはその代表的なOSSプロダクトであり、PyTorchやTensorFlowといった主要なフレームワークと連携して動作する。機械学習の専門知識を持つ者だけでなく、より広い開発者コミュニティがAIモデルを利用・再配布できる環境を整えることを志向している。企業形態としては仮想共同体的な側面も持ち合わせており、コミュニティ主導の文化が特徴的である。

技術的位置づけ

Hugging Faceは、AIモデルの「共有・配布インフラ」としての役割を担っている。大規模言語モデル(LLM)や画像生成モデルなど多様なアーキテクチャのモデルがプラットフォーム上で公開されており、オープンウェイトモデルの流通拠点として機能している。GitHubがソースコードの共有基盤であるのと同様に、Hugging FaceはAIモデルとデータセットの共有基盤としての地位を確立している。また、推論APIやSpacesと呼ばれるデモ環境の提供など、モデルの活用を支援する周辺サービスも展開している。

主要な動向

XenoSpectrumの関連記事では、Hugging Faceのプラットフォームが各種オープンモデルの公開・配布先として繰り返し登場している。2026年6月にGoogleが公開したGemma 4 12B Unifiedは、Hugging Face上でモデルウェイトが配布され、ローカル環境でのマルチモーダルエージェント構築に活用できるものとして紹介されている。同月、GoogleはGemma 4向けのMulti-Token Prediction対応ドラフトモデルも公開しており、こちらもHugging Faceを通じて提供された。

2026年5月には、Stability AIが最長6分20秒の音楽生成に対応した「Stable Audio 3.0」をオープンウェイトで公開し、Hugging Faceでのモデル配布が行われた。同月、DeepSeek-AIが1.6兆パラメータのMixture-of-Expertsモデル「DeepSeek-V4」のプレビュー版を公開した際も、同プラットフォームが配布チャネルの一つとして機能している。また、OpenAIがApache 2.0ライセンスで公開した「OpenAI Privacy Filter」や、中国のZhipu AIによる「GLM-5」の公開においても、Hugging Faceは主要な配布・公開プラットフォームとして位置づけられている。

2026年4月にGoogleが発表したオープンソース翻訳モデル「TranslateGemma」についても、Hugging Faceを通じたアクセスが提供されている。これらの動向は、オープンウェイトモデルのエコシステムにおいてHugging Faceが事実上の標準的配布拠点となっていることを示している。大規模な商用プレイヤーから中国発のスタートアップ、欧米の研究機関まで、幅広いアクターがHugging Faceを通じてモデルを公開・共有しており、AIモデルの流通インフラとしての存在感は引き続き大きい。

よくある質問

Hugging Faceとは何ですか?
2016年に米国ブルックリンで設立された企業で、機械学習モデルやデータセットの共有プラットフォームを運営している。TransformersをはじめとするオープンソースAIライブラリの開発・配布でも知られる。
Hugging Faceのプラットフォームはどのように使われますか?
研究者や開発者が事前学習済みモデルやデータセットをアップロード・ダウンロードできるモデルハブとして機能する。GoogleやMeta、各国のAIスタートアップなど幅広いアクターがモデルの公開先として利用している。
Hugging FaceはGitHubとどう違いますか?
GitHubがソースコードの共有基盤であるのに対し、Hugging FaceはAIモデルのウェイトやデータセットの共有に特化したプラットフォームである。推論APIやデモ環境(Spaces)など、モデルの活用を支援する付加機能も提供している。
最近Hugging Faceで公開された主なモデルにはどのようなものがありますか?
2026年には、GoogleのGemma 4シリーズ、Stability AIのStable Audio 3.0、DeepSeek-AIのDeepSeek-V4プレビュー版、OpenAIのPrivacy Filter、中国Zhipu AIのGLM-5など、多様なオープンウェイトモデルが同プラットフォームを通じて公開されている。
Hugging Faceの専門分野はどこにありますか?
機械学習全般を専門分野とし、特に自然言語処理(NLP)分野で発展してきた。現在は画像・音声・マルチモーダルモデルの配布にも対応しており、AIモデルの流通インフラとして幅広い領域をカバーしている。

Mentioned Articles

29 件(最新 20 件を表示)

Research Papers

5 件
  • HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face

    Yongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li, Weiming Lu, Y. Zhuang

    20231,578 件引用Semantic Scholar

    Solving complicated AI tasks with different domains and modalities is a key step toward artificial general intelligence. While there are numerous AI models available for various domains and modalities, they cannot handle complicated AI tasks autonomously. Considering large language models (LLMs) have exhibited exceptional abilities in language understanding, generation, interaction, and reasoning, we advocate that LLMs could act as a controller to manage existing AI models to solve complicated AI tasks, with language serving as a generic interface to empower this. Based on this philosophy, we present HuggingGPT, an LLM-powered agent that leverages LLMs (e.g., ChatGPT) to connect various AI models in machine learning communities (e.g., Hugging Face) to solve AI tasks. Specifically, we use ChatGPT to conduct task planning when receiving a user request, select models according to their function descriptions available in Hugging Face, execute each subtask with the selected AI model, and summarize the response according to the execution results. By leveraging the strong language capability of ChatGPT and abundant AI models in Hugging Face, HuggingGPT can tackle a wide range of sophisticated AI tasks spanning different modalities and domains and achieve impressive results in language, vision, speech, and other challenging tasks, which paves a new way towards the realization of artificial general intelligence.

  • NuminaMath: The largest public dataset in AI4Maths with 860k pairs of competition math problems and solutions

    Jia Li, E. Beeching, Lewis Tunstall, Ben Lipkin, Roman Soletskyi, Shengyi Huang, K. Rasul, Long Yu, Albert Qiaochu Jiang, Ziju Shen, Zihan Qin, Bin Dong, Li Zhou, Y. Fleureau, Guillaume Lample, Stanislas Polu, Hugging Face, AI Mistral

    289 件引用Semantic Scholar
  • The AI community building the future? A quantitative analysis of development activity on Hugging Face Hub

    Cailean Osborne, Jennifer Ding, Hannah Rose Kirk

    202482 件引用Semantic Scholar

    Open model developers have emerged as key actors in the political economy of artificial intelligence (AI), but we still have a limited understanding of collaborative practices in the open AI ecosystem. This paper responds to this gap with a three-part quantitative analysis of development activity on the Hugging Face (HF) Hub, a popular platform for building, sharing, and demonstrating models. First, various types of activity across 348,181 model, 65,761 dataset, and 156,642 space repositories exhibit right-skewed distributions. Activity is extremely imbalanced between repositories; for example, over 70% of models have 0 downloads, while 1% account for 99% of downloads. Furthermore, licenses matter: there are statistically significant differences in collaboration patterns in model repositories with permissive, restrictive, and no licenses. Second, we analyse a snapshot of the social network structure of collaboration in model repositories, finding that the community has a core-periphery structure, with a core of prolific developers and a majority of isolate developers (89%). Upon removing these isolates from the network, collaboration is characterised by high reciprocity regardless of developers’ network positions. Third, we examine model adoption through the lens of model usage in spaces, finding that a minority of models, developed by a handful of companies, are widely used on the HF Hub. Overall, the findings show that various types of activity across the HF Hub are characterised by Pareto distributions, congruent with open source software development patterns on platforms like GitHub. We conclude with recommendations for researchers, and practitioners to advance our understanding of open AI development.

  • How do Hugging Face Models Document Datasets, Bias, and Licenses? An Empirical Study

    Federica Pepe, Vittoria Nardone, A. Mastropaolo, G. Bavota, G. Canfora, Massimiliano Di Penta

    202450 件引用Semantic Scholar

    Pre-trained Machine Learning (ML) models help to create ML-intensive systems without having to spend conspicuous resources on traimng a new model from the ground up. However, the lack of transparency for such models could lead to undesired consequences in terms of bias, fairness, trustworthiness of the underlying data, and, potentially even legal implications. Taking as a case study the transformer models hosted by Hugging Face, a popular hub for pre-trained ML models, this paper empirically investigates the transparency of pre-trained transformer models. We look at the extent to which model descriptions (i) specify the datasets being used for their pre-training, (ii) discuss their possible training bias, (iii) declare their license, and whether projects using such models take these licenses into account. Results indicate that pre-trained models still have a limited exposure of their traimng datasets, possible biases, and adopted licenses. Also, we found several cases of possible licensing violations by client projects. Our findings motivate further research to improve the transparency of ML models, which may result in the definition, generation, and adoption of Artificial Intelligence Bills of Materials. CCS CONCEPTS • Software and its engineering $\rightarrow$ Software libraries and repositories.

  • Navigating Dataset Documentations in AI: A Large-Scale Analysis of Dataset Cards on Hugging Face

    Xinyu Yang, Weixin Liang, James Zou

    202448 件引用Semantic Scholar

    Advances in machine learning are closely tied to the creation of datasets. While data documentation is widely recognized as essential to the reliability, reproducibility, and transparency of ML, we lack a systematic empirical understanding of current dataset documentation practices. To shed light on this question, here we take Hugging Face -- one of the largest platforms for sharing and collaborating on ML models and datasets -- as a prominent case study. By analyzing all 7,433 dataset documentation on Hugging Face, our investigation provides an overview of the Hugging Face dataset ecosystem and insights into dataset documentation practices, yielding 5 main findings: (1) The dataset card completion rate shows marked heterogeneity correlated with dataset popularity. (2) A granular examination of each section within the dataset card reveals that the practitioners seem to prioritize Dataset Description and Dataset Structure sections, while the Considerations for Using the Data section receives the lowest proportion of content. (3) By analyzing the subsections within each section and utilizing topic modeling to identify key topics, we uncover what is discussed in each section, and underscore significant themes encompassing both technical and social impacts, as well as limitations within the Considerations for Using the Data section. (4) Our findings also highlight the need for improved accessibility and reproducibility of datasets in the Usage sections. (5) In addition, our human annotation evaluation emphasizes the pivotal role of comprehensive dataset content in shaping individuals' perceptions of a dataset card's overall quality. Overall, our study offers a unique perspective on analyzing dataset documentation through large-scale data science analysis and underlines the need for more thorough dataset documentation in machine learning research.

External Mentions

10 件