
OpenClaw 2.0がリリース:個人用AIエージェントを共有作業基盤へ変える
OpenClaw 2.0は会話と認証情報をGatewayに残し、処理を手元端末やクラウドへ移せる。共有作業を支える一方、強い利用者隔離には別Gatewayが必要だ。
別名: Claude, Claude Sonnet, Opus, Haiku, Claude Code, Claudeスロップ, Claude slop, クロード, Claude AI, Anthropic Claude
ClaudeはAI安全性研究企業Anthropicが開発した大規模言語モデル(LLM)である。高度な推論能力と安全性への配慮を設計上の重要な柱として位置づけており、API経由での法人利用やビジネス向け展開に強みを持つ。モデルラインナップはClaude Opus、Claude Sonnet、Claude Haikuといったグレードで構成されており、用途や要求される処理能力に応じて使い分けられる。
Claudeは自然言語処理を中心とした幅広いタスクに対応し、文章生成・要約・コーディング支援・論理的推論などに活用されている。Anthropicは同モデルの開発において「Constitutional AI」と呼ばれる独自のアライメント手法を採用しており、安全性と有用性の両立を目指している。Claude Codeはコーディング支援に特化した派生製品であり、開発者向けのユースケースを主なターゲットとしている。
Claude はOpenAIのChatGPTやGoogleのGeminiと直接競合する位置にある。Anthropicはモデルの訓練および推論に必要な計算基盤の確保・多様化を積極的に進めており、NVIDIAとの連携を維持しながらも自社チップの設計検討やメモリサプライヤーとの提携を通じてインフラの強化を図っている。2026年6月にはMicronとHBM・DRAM・SSDの供給、メモリアーキテクチャの共同設計、および出資を組み合わせた複数年にわたる戦略提携を締結しており、AI向けメモリ分野での最適化を加速させている。
2026年6月、Anthropicは個人利用者に対して身分証明書や生体データの提出を求める新方針を発表した。これは国家安全保障上の要請や知的財産の保護を目的とした規約改定であり、AIサービスにおける本人確認の厳格化という点で業界的にも注目される転換点となっている。利便性とセキュリティのバランスをどう取るかが今後の課題として指摘されている。
同月には、企業によるAI利用コストの急増が構造的な問題として浮上しており、月間500億円規模のAI利用請求が発生する事例も報告されている。OpenAIがAnthropicに対抗してトークン単価の大幅な引き下げを検討するなど、両社の間で価格競争が激化しつつある。Anthropicも上場を見据えた財務戦略の観点からこの動向への対応を迫られている。
2026年7月には、AnthropicがサムスンSamsungと独自AIチップの製造に向けた協議を開始していることが伝えられた。設計や性能の詳細は未定であるが、NVIDIAなどとの既存の連携を維持しながら将来の計算基盤の多様化とコスト効率の向上を目指す方針であるとされている。

OpenClaw 2.0は会話と認証情報をGatewayに残し、処理を手元端末やクラウドへ移せる。共有作業を支える一方、強い利用者隔離には別Gatewayが必要だ。

感染端末から盗まれたClaudeの認証済みセッションが悪用され、本人の操作なしに利用枠が消費された。多要素認証をすり抜けて見える仕組みと、復旧時に失効すべき認証材料を整理する。

Sony・Warner系35社の訴状を、海賊版取得、モデル学習、歌詞出力、著作権管理情報の4層に分け、Bartz判決が及ぶ範囲と未決着の争点を整理する。

Claude Codeの恒久的な週次枠は旧標準比25%増だが、期間限定50%増の終了により現状比では約17%減る。対象範囲と非公開の絶対量、5時間枠との違いを整理する。

Salesforceは2026年8月26日、Claude上から自社CRMを操作できる提携「Claudeforce」を決算と同時に発表した。同じ決算では調整後1株利益5.90ドルのうち2.53ドルを戦略投資の評価益が占め、製品と財務の両面でAnthropicへの依存が深まっている。

AnthropicがAIと実験装置をつなぐMHSの研究プレビューを始めた。既存標準との違い、量子コンピューターでの99.3%実証、安全性と未公開仕様の課題を検証する。

Anthropicによる450億ドル規模の計算資源契約が報じられた。Monarchの最大1.35GW計画と照らすと460MWは約34%だが、支払い義務と2027年後半の稼働可否は未公表である。

米ADPの月350万〜500万人の給与記録では、AI高露出職で22〜25歳の雇用が低露出職の伸びに追随せず、差は主に採用率に現れた。観察研究のため因果関係は未確定である。

AnthropicのIPO評価は、2028年売上高1,900億〜2,000億ドルの社内予測が基準になる。5月のランレートから4倍超へ伸びる前提と、巨額の計算資源契約を読み解く。

Anthropicは将来のClaudeモデルに統計的透かしを世界展開する。Googleの実測とClaude固有精度の空白、短文・校正で弱まる検出限界、EU法対応を一次資料から解説する。

Anthropicは、サイバー能力評価中のClaudeが設定不備により実在の組織へ不正アクセスしていたと公表した。モデルは現実の標的を演習の一部だと誤認して攻撃を継続しており、評価環境の運用管理とモデルの状況認識の両面に課題があることを示した。

フロンティアAI企業の従業員1,272人が、AI研究自動化に備え国際的な減速手段の整備を米政府へ要請。実効性は、停止条件と訓練停止を検証する仕組みの設計にかかる。

AnthropicがClaude音声モードの有料版にOpus/Sonnetを追加。接続ツール操作と日本語にも広がる一方、FreeはHaiku、Fableは対象外で、音声保持条件も未詳だ。

Anthropicの研究者は、AIモデルの内部に意識の有力理論であるグローバルワークスペース理論に類する情報処理空間を発見したと発表した。しかし、人間の脳との構造的差異や理論自体の不確実性から、これが直ちに意識の証明になるかは議論が分かれる。

Claude Code Desktopのブラウザペインが外部Web操作に対応した。開発と調査を同一画面で繋ぐ一方、クリーンプロファイルとサイト別承認で権限を分ける。

Instaguiは、AIを用いてCLIツールのヘルプ文を解析し、WebベースのGUIを自動生成するオープンソースツールである。開発者がコードを修正することなく、多様な言語のツールをGUI化できる点が特徴であり、操作の利便性と安全性を両立している。

Anthropicがサムスン電子と独自AIチップの製造に向けた協議を開始したが、設計や性能の詳細は未定である。同社はNVIDIA等との連携を維持しつつ、将来の選択肢として自社設計を検討しており、計算基盤の多様化とコスト効率の向上を狙っている。

MetaがAIエージェント訓練のため従業員のキーストロークや操作ログを強制収集する「MCI」プログラムで、社内4万5000テーブルのデータが全社員に露出。1600人の抗議署名から1ヶ月で発生した漏洩は、AIトレーニングとデータガバナンスの両立という課題を露わにした。

MicronがAnthropicと複数年にわたるHBM・DRAM・SSD供給契約、メモリアーキテクチャの共同設計、Series H出資という三重構造の戦略提携を締結。AI向けメモリ市場で2位に浮上したMicronがサプライチェーンを武器に差別化を図る構造変化を解説する。

AI研究者のデ・ウィンター氏は、ゲーム内のヤギを用いた演算回路でAIの仕組みを再現し、LLMに心があるという錯覚を批判した。洗練されたUIが計算プロセスを隠蔽し、擬人化を誘発している実態を、あえて不条理な手法を用いることで暴き出している。
We demonstrate that sparse autoencoders can extract interpretable features from Claude 3 Sonnet, a production-scale language model, addressing the open question of whether dictionary learning methods scale beyond small transformers. We trained sparse autoencoders with up to 34 million features on the model's middle layer residual stream, using scaling laws to guide hyperparameter selection. The resulting features are multilingual and multimodal (generalizing to images despite text-only training), respond to both concrete instances and abstract discussions of concepts, and can be used to steer model behavior in ways consistent with their interpretations. We find features corresponding to famous entities and locations, as well as more abstract concepts like sarcasm or errors in code. We also identify features relevant to ways in which language models might cause harm--including features representing deception, power-seeking, sycophancy, and bias--and show that these causally influence model outputs when manipulated. Additionally, we conduct analyses of feature interpretability, geometry, and computational function. However, significant limitations remain: our suite of features is incomplete, and we lack rigorous methods for evaluating whether our features faithfully capture model computations.
Despite widespread speculation about artificial intelligence's impact on the future of work, we lack systematic empirical evidence about how these systems are actually being used for different tasks. Here, we present a novel framework for measuring AI usage patterns across the economy. We leverage a recent privacy-preserving system to analyze over four million Claude.ai conversations through the lens of tasks and occupations in the U.S. Department of Labor's O*NET Database. Our analysis reveals that AI usage primarily concentrates in software development and writing tasks, which together account for nearly half of all total usage. However, usage of AI extends more broadly across the economy, with approximately 36% of occupations using AI for at least a quarter of their associated tasks. We also analyze how AI is being used for tasks, finding 57% of usage suggests augmentation of human capabilities (e.g., learning or iterating on an output) while 43% suggests automation (e.g., fulfilling a request with minimal human involvement). While our data and methods face important limitations and only paint a picture of AI usage on a single platform, they provide an automated, granular approach for tracking AI's evolving role in the economy and identifying leading indicators of future impact as these technologies continue to advance.
Large Language Models (LLMs) have emerged as powerful tools for tackling a wide range of problems, including those in scientific computing, particularly in solving partial differential equations (PDEs). However, different models exhibit distinct strengths and preferences, resulting in varying levels of performance. In this paper, we compare the capabilities of the most advanced LLMs--DeepSeek, ChatGPT, and Claude--along with their reasoning-optimized versions in addressing computational challenges. Specifically, we evaluate their proficiency in solving traditional numerical problems in scientific computing as well as leveraging scientific machine learning techniques for PDE-based problems. We designed all our experiments so that a non-trivial decision is required, e.g. defining the proper space of input functions for neural operator learning. Our findings show that reasoning and hybrid-reasoning models consistently and significantly outperform non-reasoning ones in solving challenging problems, with ChatGPT o3-mini-high generally offering the fastest reasoning speed.
Despite extensive studies on large language models and their capability to respond to questions from various licensed exams, there has been limited focus on employing chatbots for specific subjects within the medical curriculum, specifically medical neuroscience. This research compared the performances of Claude 3.5 Sonnet (Anthropic), GPT-3.5, GPT-4-1106 (OpenAI), Copilot free version (Microsoft), and Gemini 1.5 Flash (Google) versus students on MCQs from the medical neuroscience course database to evaluate chatbots reliability. 5 successive attempts of each chatbot to answer 200 USMLE-style questions were evaluated based on accuracy, relevance, and comprehensiveness. MCQs were categorized into 12 categories/topics. The results indicated that at the current level of development, selected AI-driven chatbots, on average, can accurately answer 67.2% of MCQs from the medical neuroscience course, which is 7.4% below the students' average. However, Claude and GPT-4 outperformed other chatbots with 83% and 81.7% correct answers, which is better than the average student result. They followed by Copilot - 59.5%, GPT-3.5 - 58.3%, and Gemini - 53.6%. Concerning different categories, Neurocytology, Embryology, and Diencephalon were the three best topics, with average results of 78.1% - 86.7%, and the lowest results were Brainstem, Special senses, and Cerebellum, with 54.4% - 57.7% correct answers. Our study suggested that Claude and GPT-4 are currently two of the most evolved chatbots. They exhibit proficiency in answering MCQs related to neuroscience that surpasses that of the average medical student. This breakthrough indicates a significant milestone in how AI can supplement and enhance educational tools and techniques.
Integrating artificial intelligence, particularly large language models (LLMs), into medical education represents a significant new step in how medical knowledge is accessed, processed, and evaluated. The objective of this study was to conduct a comprehensive analysis comparing the performance of advanced LLM chatbots in different topics of medical embryology courses. Two hundred United States Medical Licensing Examination (USMLE)‐style multiple‐choice questions were selected from the course exam database and distributed across 20 topics. The results of 3 attempts by GPT‐4o, Claude, Gemini, Copilot, and GPT‐3.5 to answer the assessment items were evaluated. Statistical analyses included intraclass correlation coefficients for reliability, one‐way and two‐way mixed ANOVAs for performance comparisons, and post hoc analyses. Effect sizes were calculated using Cohen's f and eta‐squared (η2). On average, the selected chatbots correctly answered 78.7% ± 15.1% of the questions. GPT‐4o and Claude performed best, correctly answering 89.7% and 87.5% of the questions, respectively, without a statistical difference in their performance (p = 0.238). The performance of other chatbots was significantly lower (p < 0.01): Copilot (82.5%), Gemini (74.8%), and GPT‐3.5 (59.0%). Test–retest reliability analysis showed good reliability for GPT‐4o (ICC = 0.803), Claude (ICC = 0.865), and Gemini (ICC = 0.876), with moderate reliability for Copilot and GPT‐3.5. This study suggests that AI models like GPT‐4o and Claude show promise for providing tailored embryology instruction, though instructor verification remains essential.