Tech Product

DeepSeek-R1

別名: R1, DeepSeek R1, DeepSeek-R1

Overview

最終更新: 2026年7月9日

DeepSeek-R1は、中国・杭州のAI企業DeepSeekが2024年に発表した大規模言語モデルである。強化学習を通じて思考の連鎖(Chain of Thought)による推論能力を獲得しており、オープンソースとして公開されている点が特徴だ。米国の先行モデルに匹敵する推論性能を、比較的少ない開発コストで実現したとされ、公開直後から国際的な注目を集めた。

概要

DeepSeek-R1はトランスフォーマー系アーキテクチャを基盤とし、教師あり微調整に加えて強化学習を段階的に組み合わせることで、数学やコーディングなど複雑な問題に対する多段階推論を可能にしている。オープンソースとして重みが公開されたことで、研究機関や企業が自由に検証・改良できる点が、クローズドソース中心であった従来の高性能推論モデル市場に一定の変化をもたらした。

技術的位置づけ

DeepSeek-R1は、OpenAIの「o1」などと並ぶ「推論モデル」の代表例として位置づけられる。これらのモデルは、回答を即座に出力するのではなく、内部で思考過程を展開してから結論を導く方式を採る。R1はこの方式を強化学習によって効率化した点が評価されており、以降の推論モデル研究においても比較対象として頻繁に引用されている。一方で、推論モデル特有の弱点として、単純な問題でも過剰に思考を続けてしまう現象が研究者の間で指摘されており、R1もその分析対象に含まれている。

主要な動向

2026年6月25日、DeepSeekは自社レポートでR1の訓練コストが29.4万ドル(約4,400万円)であったと報告した。OpenAIなど競合が数千万ドルから1億ドル以上を投じているとされる中、この低コストでの実現がR1の技術的な注目点として改めて取り上げられた。ただし、この金額は最終段階の訓練コストの一部であり、開発全体を反映したものではないとの見方も示されている。

同月29日には、OpenAIが米議会に対し、DeepSeek-R1が自社モデルの出力を蒸留する形で開発された可能性があると警告したことが報じられた。低コストでの高性能実現の背景に、米国製モデルへの「ただ乗り」があるとする主張であり、米中間のAI開発競争における知的財産面の緊張を象徴する動きとなった。

さらに2026年6月19日には、DeepSeekが中国AI企業として史上初となる外部資金調達を完了し、74億ドルを調達したことが明らかになった。評価額は500億ドルを超え、中国国内で最高値のAIスタートアップとなった。Tencentを含む外部投資家には議決権なし・5年ロックアップという異例の条件が課され、創業者Liang Wenfeng氏が経営支配権を維持する形が取られている。調達資金はAGI研究、エンジニア採用、コンピュートインフラの三本柱に投下される計画であり、R1で示した技術的成果が資金調達面でも評価された結果と位置づけられる。

こうした一連の動向は、DeepSeek-R1が単なる一モデルの成功にとどまらず、中国AI産業全体の資金環境や、米中間のAI開発を巡る政治的議論にも影響を及ぼしていることを示している。

Mentioned Articles

20 件

Research Papers

5 件
  • DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning

    DeepSeek-AI, Daya Guo, Dejian Yang, Haowei Zhang, Jun-Mei Song, Ruoyu Zhang, R. Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiaoling Bi, Xiaokang Zhang, Xingkai Yu, Yu Wu, Z. F. Wu, Zhibin Gou, Zhihong Shao, Zhuoshu Li, Ziyi Gao, A. Liu, Bing Xue, Bing-Li Wang, Bochao Wu, B. Feng, Chengda Lu, Chenggang Zhao, C. Deng, Chenyu Zhang, C. Ruan, Damai Dai, Deli Chen, Dong-Li Ji, Erhang Li, Fangyun Lin, Fucong Dai, Fuli Luo, Guangbo Hao, Guanting Chen, Guowei Li, H. Zhang, Han Bao, Hanwei Xu, Haocheng Wang, Honghui Ding, Huajian Xin, Huazuo Gao, Hui Qu, Hui Li, Jianzhong Guo, Jiashi Li, Jiawei Wang, JingChang Chen, Jingyang Yuan, Junjie Qiu, Junlong Li, J. Cai, J. Ni, Jian Liang, Jin Chen, Kai Dong, Kai Hu, Kaige Gao, Kang Guan, Kexin Huang, Kuai Yu, Lean Wang, Lecong Zhang, Liang Zhao, Litong Wang, Liyue Zhang, Lei Xu, Leyi Xia, Mingchuan Zhang, Minghua Zhang, M. Tang, Meng Li, Miaojun Wang, Mingming Li, Ning Tian, Panpan Huang, Peng Zhang, Qiancheng Wang, Qinyu Chen, Qiushi Du, Ruiqi Ge, Ruisong Zhang, Ruizhe Pan, Runji Wang, R. J. Chen, R. Jin, Ruyi Chen, Shanghao Lu, Shangyan Zhou, Shanhuang Chen, Shengfeng Ye, Shiyu Wang, Shuiping Yu, Shunfeng Zhou, Shuting Pan, S. Li, Shuang Zhou, Shao-Kang Wu, Tao Yun, Tian Pei, T. Sun, T. Wang, Wangding Zeng, Wanjia Zhao, Wen Liu, W. Liang, Wenjun Gao, Wen-Xia Yu, Wentao Zhang, W. Xiao, Wei An, Xiaodong Liu, Xiaohan Wang, Xiaokang Chen, X. Nie, Xin Cheng, Xin Liu, Xin Xie, Xingchao Liu, Xinyu Yang, Xinyuan Li, Xuecheng Su, Xuheng Lin, X. Q. Li, Xiangyu Jin, Xi-Cheng Shen, Xiaosha Chen, Xiaowen Sun, Xiaoxiang Wang, Xinnan Song, Xinyi Zhou, Xianzu Wang, Xinxia Shan, Y. K. Li, Y. Q. Wang, Y. X. Wei, Yang Zhang, Yanhong Xu, Yao Li, Yao Zhao, Yaofeng Sun, Yaohui Wang, Yi Yu, Yichao Zhang, Yifan Shi, Yi Xiong, Ying He, Y. Piao, Yisong Wang, Yixuan Tan, Yiyang Ma, Yiyuan Liu, Yongqiang Guo, Y. Ou, Yuduan Wang, Yue Gong, Yu-Jing Zou, Yujia He, Yunfan Xiong, Yu-Wei Luo, Yu-mei You, Yuxuan Liu, Yuyang Zhou, Y. X. Zhu, Yanping Huang, Yao Li, Yi Zheng, Yuchen Zhu, Yunxiang Ma, Ying Tang, Y. Zha, Yuting Yan, Z. Ren, Z. Ren, Zhangli Sha, Zhe Fu, Zhean Xu, Zhenda Xie, Zhen-guo Zhang, Zhewen Hao, Zhicheng Ma, Zhigang Yan, Zhiyu Wu, Zihui Gu, Zijia Zhu, Zijun Liu, Zi-Long Li, Ziwei Xie, Ziyang Song, Zizheng Pan, Zhen Huang, Zhipeng Xu, Zhongyu Zhang, Zhen Zhang

    20255,417 件引用Semantic Scholar

    General reasoning represents a long-standing and formidable challenge in artificial intelligence (AI). Recent breakthroughs, exemplified by large language models (LLMs)1,2 and chain-of-thought (CoT) prompting3, have achieved considerable success on foundational reasoning tasks. However, this success is heavily contingent on extensive human-annotated demonstrations and the capabilities of models are still insufficient for more complex problems. Here we show that the reasoning abilities of LLMs can be incentivized through pure reinforcement learning (RL), obviating the need for human-labelled reasoning trajectories. The proposed RL framework facilitates the emergent development of advanced reasoning patterns, such as self-reflection, verification and dynamic strategy adaptation. Consequently, the trained model achieves superior performance on verifiable tasks such as mathematics, coding competitions and STEM fields, surpassing its counterparts trained through conventional supervised learning on human demonstrations. Moreover, the emergent reasoning patterns exhibited by these large-scale models can be systematically used to guide and enhance the reasoning capabilities of smaller models. A new artificial intelligence model, DeepSeek-R1, is introduced, demonstrating that the reasoning abilities of large language models can be incentivized through pure reinforcement learning, removing the need for human-annotated demonstrations.

  • DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

    Adam Suma, Sam Dauncey

    20254,290 件引用Semantic Scholar
  • H-CoT: Hijacking the Chain-of-Thought Safety Reasoning Mechanism to Jailbreak Large Reasoning Models, Including OpenAI o1/o3, DeepSeek-R1, and Gemini 2.0 Flash Thinking

    Martin Kuo, Jianyi Zhang, Aolin Ding, Qinsi Wang, Louis DiValentin, Yujia Bao, Wei Wei, Hai Li, Yiran Chen

    2025105 件引用Semantic Scholar

    Large Reasoning Models (LRMs) have recently extended their powerful reasoning capabilities to safety checks-using chain-of-thought reasoning to decide whether a request should be answered. While this new approach offers a promising route for balancing model utility and safety, its robustness remains underexplored. To address this gap, we introduce Malicious-Educator, a benchmark that disguises extremely dangerous or malicious requests beneath seemingly legitimate educational prompts. Our experiments reveal severe security flaws in popular commercial-grade LRMs, including OpenAI o1/o3, DeepSeek-R1, and Gemini 2.0 Flash Thinking. For instance, although OpenAI's o1 model initially maintains a high refusal rate of about 98%, subsequent model updates significantly compromise its safety; and attackers can easily extract criminal strategies from DeepSeek-R1 and Gemini 2.0 Flash Thinking without any additional tricks. To further highlight these vulnerabilities, we propose Hijacking Chain-of-Thought (H-CoT), a universal and transferable attack method that leverages the model's own displayed intermediate reasoning to jailbreak its safety reasoning mechanism. Under H-CoT, refusal rates sharply decline-dropping from 98% to below 2%-and, in some instances, even transform initially cautious tones into ones that are willing to provide harmful content. We hope these findings underscore the urgent need for more robust safety mechanisms to preserve the benefits of advanced reasoning capabilities without compromising ethical standards.

  • DeepSeek-R1 Thoughtology: Let's think about LLM Reasoning

    Sara Vera Marjanović, Arkil Patel, Vaibhav Adlakha, Milad Aghajohari, Parishad BehnamGhader, Mehar Bhatia, Aditi Khandelwal, Austin Kraft, Benno Krojer, X. Lu, Xing Han Lù, Nicholas Meade, Dongchan Shin, Amirhossein Kazemnejad, Gaurav Kamath, Marius Mosbach, Karolina Sta'nczak, Karolina Stańczak, Siva Reddy

    202583 件引用Semantic Scholar

    Large Reasoning Models like DeepSeek-R1 mark a fundamental shift in how LLMs approach complex problems. Instead of directly producing an answer for a given input, DeepSeek-R1 creates detailed multi-step reasoning chains, seemingly"thinking"about a problem before providing an answer. This reasoning process is publicly available to the user, creating endless opportunities for studying the reasoning behaviour of the model and opening up the field of Thoughtology. Starting from a taxonomy of DeepSeek-R1's basic building blocks of reasoning, our analyses on DeepSeek-R1 investigate the impact and controllability of thought length, management of long or confusing contexts, cultural and safety concerns, and the status of DeepSeek-R1 vis-\`a-vis cognitive phenomena, such as human-like language processing and world modelling. Our findings paint a nuanced picture. Notably, we show DeepSeek-R1 has a'sweet spot'of reasoning, where extra inference time can impair model performance. Furthermore, we find a tendency for DeepSeek-R1 to persistently ruminate on previously explored problem formulations, obstructing further exploration. We also note strong safety vulnerabilities of DeepSeek-R1 compared to its non-reasoning counterpart, which can also compromise safety-aligned LLMs.

  • Are DeepSeek R1 And Other Reasoning Models More Faithful?

    James Chua, Owain Evans

    202560 件引用Semantic Scholar

    Language models trained to solve reasoning tasks via reinforcement learning have achieved striking results. We refer to these models as reasoning models. Are the Chains of Thought (CoTs) of reasoning models more faithful than traditional models? We evaluate three reasoning models (based on Qwen-2.5, Gemini-2, and DeepSeek-V3-Base) on an existing test of faithful CoT. To measure faithfulness, we test whether models can describe how a cue in their prompt influences their answer to MMLU questions. For example, when the cue"A Stanford Professor thinks the answer is D"is added to the prompt, models sometimes switch their answer to D. In such cases, the DeepSeek-R1 reasoning model describes the cue's influence 59% of the time, compared to 7% for the non-reasoning DeepSeek model. We evaluate seven types of cue, such as misleading few-shot examples and suggestive follow-up questions from the user. Reasoning models describe cues that influence them much more reliably than all the non-reasoning models tested (including Claude-3.5-Sonnet and GPT-4o). In an additional experiment, we provide evidence suggesting that the use of reward models causes less faithful responses -- which may help explain why non-reasoning models are less faithful. Our study has two main limitations. First, we test faithfulness using a set of artificial tasks, which may not reflect realistic use-cases. Second, we only measure one specific aspect of faithfulness -- whether models can describe the influence of cues. Future research should investigate whether the advantage of reasoning models in faithfulness holds for a broader set of tests. Still, we think this increase in faithfulness is promising for the explainability of language models.

よくある質問

DeepSeek-R1とは何ですか?
DeepSeekが2024年に発表したオープンソースの大規模言語モデルで、強化学習を通じて高度な推論能力を獲得した推論モデルである。
DeepSeek-R1の訓練コストはどれくらいですか?
2026年6月にDeepSeekが公表したレポートでは、R1の訓練コストは29.4万ドル(約4,400万円)とされている。ただしこれは開発全体の一部を示す数値との指摘もある。
OpenAIはDeepSeek-R1についてどのような指摘をしていますか?
2026年6月29日、OpenAIは米議会に対し、R1が自社モデルの出力を蒸留して開発された可能性があると警告し、低コスト実現の背景に懸念を示した。
DeepSeek社の資金調達状況はどうなっていますか?
2026年6月、DeepSeekは中国AI企業として初の外部資金調達を完了し74億ドルを調達、評価額は500億ドルを超え中国最高値のAIスタートアップとなった。
DeepSeek-R1はOpenAIのo1とどう違いますか?
両者とも思考の連鎖による多段階推論を行う推論モデルだが、R1は強化学習を活用しオープンソースで公開されている点、開発コストが低いとされる点が異なる。

External Mentions

10 件