
DeepSeekが中国AI史上初の外部資金調達、74億ドルで評価額500億ドルを超えAGI研究へ
DeepSeekが中国AI企業として史上初の外部資金調達を完了し、74億ドルを調達した。評価額は500億ドル超に達し、中国最高値のAIスタートアップとなった。Tencentを含む外部投資家には議決権なし・5年ロックアップという異例の条件が課され、創業者 Liang Wenfeng 氏が経営支配権を完全に維持する。資金はAGI研究・エンジニア採用・コンピュートインフラの三本柱に投下される。
別名: R1, DeepSeek R1, DeepSeek-R1
DeepSeek-R1は、中国・杭州のAI企業DeepSeekが2024年に発表した大規模言語モデルである。強化学習を通じて思考の連鎖(Chain of Thought)による推論能力を獲得しており、オープンソースとして公開されている点が特徴だ。米国の先行モデルに匹敵する推論性能を、比較的少ない開発コストで実現したとされ、公開直後から国際的な注目を集めた。
DeepSeek-R1はトランスフォーマー系アーキテクチャを基盤とし、教師あり微調整に加えて強化学習を段階的に組み合わせることで、数学やコーディングなど複雑な問題に対する多段階推論を可能にしている。オープンソースとして重みが公開されたことで、研究機関や企業が自由に検証・改良できる点が、クローズドソース中心であった従来の高性能推論モデル市場に一定の変化をもたらした。
DeepSeek-R1は、OpenAIの「o1」などと並ぶ「推論モデル」の代表例として位置づけられる。これらのモデルは、回答を即座に出力するのではなく、内部で思考過程を展開してから結論を導く方式を採る。R1はこの方式を強化学習によって効率化した点が評価されており、以降の推論モデル研究においても比較対象として頻繁に引用されている。一方で、推論モデル特有の弱点として、単純な問題でも過剰に思考を続けてしまう現象が研究者の間で指摘されており、R1もその分析対象に含まれている。
2026年6月25日、DeepSeekは自社レポートでR1の訓練コストが29.4万ドル(約4,400万円)であったと報告した。OpenAIなど競合が数千万ドルから1億ドル以上を投じているとされる中、この低コストでの実現がR1の技術的な注目点として改めて取り上げられた。ただし、この金額は最終段階の訓練コストの一部であり、開発全体を反映したものではないとの見方も示されている。
同月29日には、OpenAIが米議会に対し、DeepSeek-R1が自社モデルの出力を蒸留する形で開発された可能性があると警告したことが報じられた。低コストでの高性能実現の背景に、米国製モデルへの「ただ乗り」があるとする主張であり、米中間のAI開発競争における知的財産面の緊張を象徴する動きとなった。
さらに2026年6月19日には、DeepSeekが中国AI企業として史上初となる外部資金調達を完了し、74億ドルを調達したことが明らかになった。評価額は500億ドルを超え、中国国内で最高値のAIスタートアップとなった。Tencentを含む外部投資家には議決権なし・5年ロックアップという異例の条件が課され、創業者Liang Wenfeng氏が経営支配権を維持する形が取られている。調達資金はAGI研究、エンジニア採用、コンピュートインフラの三本柱に投下される計画であり、R1で示した技術的成果が資金調達面でも評価された結果と位置づけられる。
こうした一連の動向は、DeepSeek-R1が単なる一モデルの成功にとどまらず、中国AI産業全体の資金環境や、米中間のAI開発を巡る政治的議論にも影響を及ぼしていることを示している。

DeepSeekが中国AI企業として史上初の外部資金調達を完了し、74億ドルを調達した。評価額は500億ドル超に達し、中国最高値のAIスタートアップとなった。Tencentを含む外部投資家には議決権なし・5年ロックアップという異例の条件が課され、創業者 Liang Wenfeng 氏が経営支配権を完全に維持する。資金はAGI研究・エンジニア採用・コンピュートインフラの三本柱に投下される。

AutoTTSは、LLMの推論コストを削減するため、制御アルゴリズムの探索自体をAIに委ねるという発想で開発された。Claude Codeを探索エージェントとして活用し、わずか39.90ドルでトークン使用量を約70%削減するアルゴリズムを自律発見した。この低コストは、AIによるアルゴリズム発見が個人研究者やスタートアップにも開かれつつあることを示している。

Ubuntuは、AI機能の一括オフ機能は複雑で実装できないと明言し、Snap confinementとローカル推論による透明性を選択した。これは「UbuntuをAI製品にしない」という宣言であり、ユーザーの不信感に対し、具体的な設計選択で応えるものだ。また、Implicit AIとExplicit AIの二種類に分類し、OS機能として溶け込むAIと、ユーザーが呼び出すAIを区別している。

アラブ首長国連邦は、政府のサービスとプロセスを2年以内に50%エージェント型AIへ移行させる国家戦略を発表した。これは、AIを単なるツールではなく、自律的に意思決定し実行する「執行パートナー」として位置づけ、市民の複雑な行政手続きを大幅に簡素化することを目指す。UAEは、長年のデジタルインフラ構築とトップダウンのアジリティにより、この野心的な目標達成に自信を示している。

2025年初頭、シリコンバレーとワシントンの双方に激震が走った。中国・杭州を拠点とするAIスタートアップ、DeepSeekが公開した「DeepSeek-R1」は、米国製モデルに匹敵する性能をわずかなコストで実現したとされ […]

2026年1月、スイスのダボスで開催された世界経済フォーラム(WEF)。雪に閉ざされた静謐なリゾート地で、世界のテクノロジー業界を震撼させる「舌戦」が繰り広げられた。 生成AIの安全性と倫理を最重視するAnthropic […]

私たちの直感に反する奇妙な現象が、最先端の人工知能(AI)の世界で起きている。 人類はこれまで、OpenAIの「o1」やDeepSeekの「R1」といった「推論モデル」の開発に熱狂してきた。これらは、「思考の連鎖(Cha […]

中国のAI開発企業DeepSeekが、同社の推論モデル「R1」の訓練コストがわずか29.4万ドル(約4,400万円)であるとする詳細なレポートを発表した。OpenAIなどが数千万ドルから1億ドル以上を投じているとされる中 […]

OpenAIは2025年8月5日、「gpt-oss-120b」と「gpt-oss-20b」という2つのオープンウェイトモデルを同時に公開した。これは、2019年のGPT-2以来となるオープンソースへの回帰である。プロプラ […]

GoogleがAIベンチマークの再定義に乗り出した。従来の静的テストに代わり、動的かつ対話的なゲーム環境でAIの「思考」を可視化する試みとして、同社は新プラットフォーム「Kaggle Game Arena」を正式発表。初 […]

NVIDIAは、中国DeepSeek社の巨大推論モデル「DeepSeek R1 0528」の知性を、より小型で効率的なモデル群に凝縮した「OpenReasoning-Nemotron」ファミリーをオープンソースとして公開 […]

大規模言語モデル(LLM)は、流暢な会話をこなし、専門的な質問にも答える。その驚くべき能力に、私たちは「AIは本当に理解しているのではないか」という期待を抱きがちだ。しかし、その知性は本物なのだろうか? こうした我々の抱 […]

Microsoftが開発する小型言語モデル(SLM)群「Phi」シリーズに、目覚ましい「推論能力」を実装した新モデル群「Phi-4」が加わった。わずか一年前、Phi-3をもってSLMの潜在力を世に示したMicrosoft […]

サンフランシスコの新興企業Deep Cogitoが、ステルスモードを解除し、高性能なオープンソースAIモデル群「Cogito v1」を発表した。独自のIDA訓練手法とハイブリッド推論機能を備え、既存のLlamaやDeep […]

Anthropicが、AIの思考プロセス、いわゆる「思考の連鎖:Chain-of-Thought(CoT)」の信頼性に関する衝撃的な研究結果を発表した。最新の高性能推論モデルでさえ、自身の思考過程を偽り、時には不正な情報 […]

Googleは、同社史上最も高性能とされるAIモデル「Gemini 2.5 Pro」のプレビュー版を開発者向けに公開した。これまで制限付きの無料実験版のみだったが、今回のリリースでより高いレート制限と明確な料金体系が導入 […]

Googleは、同社が「最もインテリジェント」と位置づける最新AIモデル「Gemini 2.5 Pro」を発表した。このモデルは、応答前に内部で「思考」する能力を備え、複雑なタスクにおける推論やコーディング性能を大幅に向 […]

NVIDIAは年次開発者会議GTC 2025で、新世代のAIチップ「Blackwell Ultra」と将来のGPUアーキテクチャ「Vera Rubin」を発表した。Blackwell Ultraは2025年後半から出荷予 […]

DeepSeekが発表した推論モデル「DeepSeek-R1」は、優れた推論能力を持つ一方で、ハルシネーション(事実に基づかない情報を生成する現象)率が他社の主要モデルと比較して突出して高いことが、Vectaraの調査で […]

中国の新興AIスタートアップDeepSeekが、人工知能の歴史に新たな一章を刻む革新的な言語モデル「DeepSeek-R1」を発表した。このモデルは、業界最高峰とされるOpenAIの「o1」と同等の性能を持ちながら、驚異 […]
General reasoning represents a long-standing and formidable challenge in artificial intelligence (AI). Recent breakthroughs, exemplified by large language models (LLMs)1,2 and chain-of-thought (CoT) prompting3, have achieved considerable success on foundational reasoning tasks. However, this success is heavily contingent on extensive human-annotated demonstrations and the capabilities of models are still insufficient for more complex problems. Here we show that the reasoning abilities of LLMs can be incentivized through pure reinforcement learning (RL), obviating the need for human-labelled reasoning trajectories. The proposed RL framework facilitates the emergent development of advanced reasoning patterns, such as self-reflection, verification and dynamic strategy adaptation. Consequently, the trained model achieves superior performance on verifiable tasks such as mathematics, coding competitions and STEM fields, surpassing its counterparts trained through conventional supervised learning on human demonstrations. Moreover, the emergent reasoning patterns exhibited by these large-scale models can be systematically used to guide and enhance the reasoning capabilities of smaller models. A new artificial intelligence model, DeepSeek-R1, is introduced, demonstrating that the reasoning abilities of large language models can be incentivized through pure reinforcement learning, removing the need for human-annotated demonstrations.
Large Reasoning Models (LRMs) have recently extended their powerful reasoning capabilities to safety checks-using chain-of-thought reasoning to decide whether a request should be answered. While this new approach offers a promising route for balancing model utility and safety, its robustness remains underexplored. To address this gap, we introduce Malicious-Educator, a benchmark that disguises extremely dangerous or malicious requests beneath seemingly legitimate educational prompts. Our experiments reveal severe security flaws in popular commercial-grade LRMs, including OpenAI o1/o3, DeepSeek-R1, and Gemini 2.0 Flash Thinking. For instance, although OpenAI's o1 model initially maintains a high refusal rate of about 98%, subsequent model updates significantly compromise its safety; and attackers can easily extract criminal strategies from DeepSeek-R1 and Gemini 2.0 Flash Thinking without any additional tricks. To further highlight these vulnerabilities, we propose Hijacking Chain-of-Thought (H-CoT), a universal and transferable attack method that leverages the model's own displayed intermediate reasoning to jailbreak its safety reasoning mechanism. Under H-CoT, refusal rates sharply decline-dropping from 98% to below 2%-and, in some instances, even transform initially cautious tones into ones that are willing to provide harmful content. We hope these findings underscore the urgent need for more robust safety mechanisms to preserve the benefits of advanced reasoning capabilities without compromising ethical standards.
Large Reasoning Models like DeepSeek-R1 mark a fundamental shift in how LLMs approach complex problems. Instead of directly producing an answer for a given input, DeepSeek-R1 creates detailed multi-step reasoning chains, seemingly"thinking"about a problem before providing an answer. This reasoning process is publicly available to the user, creating endless opportunities for studying the reasoning behaviour of the model and opening up the field of Thoughtology. Starting from a taxonomy of DeepSeek-R1's basic building blocks of reasoning, our analyses on DeepSeek-R1 investigate the impact and controllability of thought length, management of long or confusing contexts, cultural and safety concerns, and the status of DeepSeek-R1 vis-\`a-vis cognitive phenomena, such as human-like language processing and world modelling. Our findings paint a nuanced picture. Notably, we show DeepSeek-R1 has a'sweet spot'of reasoning, where extra inference time can impair model performance. Furthermore, we find a tendency for DeepSeek-R1 to persistently ruminate on previously explored problem formulations, obstructing further exploration. We also note strong safety vulnerabilities of DeepSeek-R1 compared to its non-reasoning counterpart, which can also compromise safety-aligned LLMs.
Language models trained to solve reasoning tasks via reinforcement learning have achieved striking results. We refer to these models as reasoning models. Are the Chains of Thought (CoTs) of reasoning models more faithful than traditional models? We evaluate three reasoning models (based on Qwen-2.5, Gemini-2, and DeepSeek-V3-Base) on an existing test of faithful CoT. To measure faithfulness, we test whether models can describe how a cue in their prompt influences their answer to MMLU questions. For example, when the cue"A Stanford Professor thinks the answer is D"is added to the prompt, models sometimes switch their answer to D. In such cases, the DeepSeek-R1 reasoning model describes the cue's influence 59% of the time, compared to 7% for the non-reasoning DeepSeek model. We evaluate seven types of cue, such as misleading few-shot examples and suggestive follow-up questions from the user. Reasoning models describe cues that influence them much more reliably than all the non-reasoning models tested (including Claude-3.5-Sonnet and GPT-4o). In an additional experiment, we provide evidence suggesting that the use of reward models causes less faithful responses -- which may help explain why non-reasoning models are less faithful. Our study has two main limitations. First, we test faithfulness using a set of artificial tasks, which may not reflect realistic use-cases. Second, we only measure one specific aspect of faithfulness -- whether models can describe the influence of cues. Future research should investigate whether the advantage of reasoning models in faithfulness holds for a broader set of tests. Still, we think this increase in faithfulness is promising for the explainability of language models.