Tech Product

Google Gemini

別名: Gemini, Gemini AI, Gemini 3, Gemini Nano, Google Gemini

gemini.google.com

Overview

最終更新: 2026年7月9日

Google Geminiは、Googleが開発したマルチモーダル生成AIモデルおよびそれを基盤とするチャットサービスである。テキスト・画像・音声・動画など複数のモダリティを処理する能力を持ち、Google検索やAndroid、各種クラウドサービスなど同社のプロダクトエコシステム全体に統合が進んでいる。

概要

GeminiはGoogleのAI研究部門が開発した大規模言語モデルを中核とし、単一モデルとしてのAPIアクセスから、エンドユーザー向けの対話インターフェースまで幅広い形態で提供されている。モデルの規模や用途に応じて複数のバリアントが存在し、クラウドからエッジデバイスまでの展開を想定した設計がとられている。GoogleのクラウドサービスであるGoogle CloudとFirebaseにも組み込まれ、開発者向けのAPI提供も行われている。

技術的位置づけ

Geminiは、テキスト生成にとどまらず、画像理解や音声処理などを統合したマルチモーダルモデルとして設計されている点が特徴だ。GeminiモデルはGoogle独自のTPU(Tensor Processing Unit)インフラで学習・推論が行われており、同社のAIインフラ戦略と密接に結びついている。ただし、2026年6月時点でGoogleはSpaceXとの間にGPU計算容量の調達契約を締結したことが報じられており、AI需要の急増によりGoogleも外部の計算リソースを調達せざるを得ない状況にあることが明らかになった。この契約は月額9.2億ドル規模とされ、業界全体における計算資源の逼迫を示す動向として注目されている。

主要な動向

2026年のGoogle I/Oでは、GeminiをOS全体に統合した新デバイスカテゴリ「Googlebook」が発表された。ChromeOSとAndroidを融合させた新OSにGeminiをOS全層に組み込み、カーソルが指す対象をAIが読み取って次の操作を提案する「Magic Pointer」機能などを搭載するとされ、2026年秋の出荷が予定されている。従来のChromebookとは別カテゴリのプレミアム志向AI前提PCとして位置づけられる。

同じくGoogle I/O 2026では、AIエージェントが複数の小売サービスを横断して買い物かごを一元管理する「Universal Cart」も発表された。この構想ではGeminiベースのAIが価格監視や互換性検証を担い、オープン規格「Universal Commerce Protocol(UCP)」および「Agent Payments Protocol(AP2)」と連携することで、自律的な購買支援を実現するとされている。

セキュリティ領域でもGeminiの活用が進んでいる。Google Cloudは2026年5月28日に「AI Threat Defense」を発表し、脆弱性の発見・評価・修正・監視を一本の自動パイプラインで処理する基盤を提供すると述べた。これは、攻撃者がAIを使って脆弱性の発見と武器化を加速させている状況に対応するもので、アラートではなく適用可能なパッチそのものを出荷することを目標としている。

一方でGoogleは、AIがゼロデイ脆弱性の発見と悪用に利用された初の確認事例を含むレポートも公表しており、Geminiのような生成AIが防御と攻撃の両面でサイバーセキュリティの構造を変えつつある現状を認識した上での対策強化を進めている。

Mentioned Articles

20 件

Research Papers

5 件
  • Comparative Analysis Based on DeepSeek, ChatGPT, and Google Gemini: Features, Techniques, Performance, Future Prospects

    Anichur Rahman, Shahariar Hossain Mahir, Md. Tanjum An Tashrif, A. Aishi, Md. Ahsan Karim, Dipanjali Kundu, Tanoy Debnath, Md. Abul Ala Moududi, MD. Zunead Abedin Eidmum

    202563 件引用Semantic Scholar

    Nowadays, DeepSeek, ChatGPT, and Google Gemini are the most trending and exciting Large Language Model (LLM) technologies for reasoning, multimodal capabilities, and general linguistic performance worldwide. DeepSeek employs a Mixture-of-Experts (MoE) approach, activating only the parameters most relevant to the task at hand, which makes it especially effective for domain-specific work. On the other hand, ChatGPT relies on a dense transformer model enhanced through reinforcement learning from human feedback (RLHF), and then Google Gemini actually uses a multimodal transformer architecture that integrates text, code, and images into a single framework. However, by using those technologies, people can be able to mine their desired text, code, images, etc, in a cost-effective and domain-specific inference. People may choose those techniques based on the best performance. In this regard, we offer a comparative study based on the DeepSeek, ChatGPT, and Gemini techniques in this research. Initially, we focus on their methods and materials, appropriately including the data selection criteria. Then, we present state-of-the-art features of DeepSeek, ChatGPT, and Gemini based on their applications. Most importantly, we show the technological comparison among them and also cover the dataset analysis for various applications. Finally, we address extensive research areas and future potential guidance regarding LLM-based AI research for the community.

  • Performance of the ChatGPT-3.5, ChatGPT-4, and Google Gemini large language models in responding to dental implantology inquiries.

    Noha Taymour, Shaimaa M. Fouda, Hams H. Abdelrahaman, Mohamed G. Hassan

    202543 件引用Semantic Scholar

    STATEMENT OF PROBLEM Artificial intelligence (AI) chatbots have been proposed as promising resources for oral health information. However, the quality and readability of existing online health-related information is often inconsistent and challenging. PURPOSE This study aimed to compare the reliability and usefulness of dental implantology-related information provided by the ChatGPT-3.5, ChatGPT-4, and Google Gemini large language models (LLMs). MATERIAL AND METHODS A total of 75 questions were developed covering various dental implant domains. These questions were then presented to 3 different LLMs: ChatGPT-3.5, ChatGPT-4, and Google Gemini. The responses generated were recorded and independently assessed by 2 specialists who were blinded to the source of the responses. The evaluation focused on the accuracy of the generated answers using a modified 5-point Likert scale to measure the reliability and usefulness of the information provided. Additionally, the ability of the AI-chatbots to offer definitive responses to closed questions, provide reference citation, and advise scheduling consultations with a dental specialist was also analyzed. The Friedman, Mann Whitney U and Spearman Correlation tests were used for data analysis (α=.05). RESULTS Google Gemini exhibited higher reliability and usefulness scores compared with ChatGPT-3.5 and ChatGPT-4 (P<.001). Google Gemini also demonstrated superior proficiency in identifying closed questions (25 questions, 41%) and recommended specialist consultations for 74 questions (98.7%), significantly outperforming ChatGPT-4 (30 questions, 40.0%) and ChatGPT-3.5 (28 questions, 37.3%) (P<.001). A positive correlation was found between reliability and usefulness scores, with Google Gemini showing the strongest correlation (ρ=.702). CONCLUSIONS The 3 AI Chatbots showed acceptable levels of reliability and usefulness in addressing dental implant-related queries. Google Gemini distinguished itself by providing responses consistent with specialist consultations.

  • Artificial intelligence in healthcare education: evaluating the accuracy of ChatGPT, Copilot, and Google Gemini in cardiovascular pharmacology

    I. Salman, O. Ameer, Mohammad A. Khanfar, Y. Hsieh

    202535 件引用Semantic Scholar

    Background Artificial intelligence (AI) is revolutionizing medical education; however, its limitations remain underexplored. This study evaluated the accuracy of three generative AI tools—ChatGPT-4, Copilot, and Google Gemini—in answering multiple-choice questions (MCQ) and short-answer questions (SAQ) related to cardiovascular pharmacology, a key subject in healthcare education. Methods Using free versions of each AI tool, we administered 45 MCQs and 30 SAQs across three difficulty levels: easy, intermediate, and advanced. AI-generated answers were reviewed by three pharmacology experts. The accuracy of MCQ responses was recorded as correct or incorrect, while SAQ responses were rated on a 1–5 scale based on relevance, completeness, and correctness. Results ChatGPT, Copilot, and Gemini demonstrated high accuracy scores in easy and intermediate MCQs (87–100%). While all AI models showed a decline in performance on the advanced MCQ section, only Copilot (53% accuracy) and Gemini (20% accuracy) had significantly lower scores compared to their performance on easy-intermediate levels. SAQ evaluations revealed high accuracy scores for ChatGPT (overall 4.7 ± 0.3) and Copilot (overall 4.5 ± 0.4) across all difficulty levels, with no significant differences between the two tools. In contrast, Gemini’s SAQ performance was markedly lower across all levels (overall 3.3 ± 1.0). Conclusion ChatGPT-4 demonstrates the highest accuracy in addressing both MCQ and SAQ cardiovascular pharmacology questions, regardless of difficulty level. Copilot ranks second after ChatGPT, while Google Gemini shows significant limitations in handling complex MCQs and providing accurate responses to SAQ-type questions in this field. These findings can guide the ongoing refinement of AI tools for specialized medical education.

  • Evaluating ChatGPT and Google Gemini Performance and Implications in Turkish Dental Education

    Ipek Kinikoglu

    202533 件引用Semantic Scholar

    Artificial intelligence (AI) has emerged as a transformative tool in education, particularly in specialized fields such as dentistry. This study evaluated the performance of four advanced AI models - ChatGPT-4o (San Francisco, CA: OpenAI), ChatGPT-o1, Gemini 1.5 Pro (Mountain View, CA: Google LLC), and Gemini 2.0 Advanced, in the Turkish Dental Specialty Examination (DUS) for 2020 and 2021. A total of 240 questions, comprising 120 questions per year from basic and clinical sciences, were analyzed. AI models were assessed based on their accuracy in providing correct answers compared to the official answer keys. For the 2020 DUS, ChatGPT-o1 and Gemini 2.0 Advanced achieved the highest accuracy rates of 93.70% and 96.80%, respectively, with net scores of 112.50 and 115 out of 120 questions. ChatGPT-4o and Gemini 1.5 Pro followed with accuracy rates of 83.33% and 85.40%. For the 2021 DUS, ChatGPT-o1 again demonstrated the highest accuracy at 97.88% (115.50 net score), closely followed by Gemini 2.0 Advanced at 96.82% (114.25 net score). Overall, ChatGPT-4o and Gemini 1.5 Pro scored lower for 2021, achieving accuracy rates of 88.35% and 93.64%, respectively. Combining results from both years (238 total questions), ChatGPT-o1 and Gemini 2.0 Advanced achieved accuracy rates of 97.46% (230 correct answers, 95% CI: 94.62%, 100.00%) and 97.90% (231 correct answers, 95% CI: 94.62%, 100.00%), respectively, significantly outperforming ChatGPT-4o (88.66%, 211 correct answers, 95% CI: 85.43%, 91.89%) and Gemini 1.5 Pro (91.60%, 218 correct answers, 95% CI: 87.75%, 95.45%). Statistical analysis revealed significant differences among the models (p = 0.0002). Pairwise comparisons demonstrated that ChatGPT-4o underperformed significantly compared to ChatGPT-o1 (p = 0.0016) and Gemini 2.0 Advanced (p = 0.0007) after Bonferroni correction. The consistently high accuracy rates and narrow confidence intervals for the top-performing models underscore their superior reliability and performance in answering the DUS questions. Generative AI modules such as ChatGPT-01 and Gemini 2.0 have the potential to enhance dental board exam preparation through question evaluation. While the AI modules appear to outperform humans on DUS questions, the study raises a concern about the ethical uses of AI and the true justification and value of DUS examinations as dental competency examinations. A higher level of knowledge evaluation should be considered. This research contributes to the growing body of literature on AI applications in specialized knowledge domains and provides a foundation for further exploration of its integration into dental education.

  • Political Bias in Large Language Models: A Comparative Analysis of ChatGPT-4, Perplexity, Google Gemini, and Claude

    Tavishi Choudhary

    202533 件引用Semantic Scholar

    Artificial Intelligence large language models have rapidly gained widespread adoption, sparking discussions on their societal and political impact, especially for political bias and its far-reaching consequences on society and citizens. This study explores the political bias in large language models by conducting a comparative analysis across four popular AI models—ChatGPT-4, Perplexity, Google Gemini, and Claude. This research systematically evaluates their responses to politically charged prompts and questions from the Pew Research Center’s Political Typology Quiz, Political Compass Quiz, and ISideWith Quiz. The findings revealed that ChatGPT-4 and Claude exhibit a liberal bias, Perplexity is more conservative, while Google Gemini adopts more centrist stances based on their training data sets. The presence of such biases underscores the critical need for transparency in AI development and the incorporation of diverse training datasets, regular audits, and user education to mitigate any of these biases. The most significant question surrounding political bias in AI is its consequences, particularly its influence on public discourse, policy-making, and democratic processes. The results of this study advocate for ethical implications for the development of AI models and the need for transparency to build trust and integrity in AI models. Additionally, future research directions have been outlined to explore and address the complex AI bias issue.

よくある質問

Google Geminiとは何ですか?
GoogleがGoogleが開発したマルチモーダル生成AIモデルおよびチャットサービスだ。テキスト・画像・音声・動画など複数の入力形式を処理する能力を持ち、Google検索やAndroid、Google Cloudなど同社製品群に幅広く統合されている。
GeminiはどのようなGoogleのサービスや製品に組み込まれていますか?
Google検索、Android、ChromeOS、Google Cloudなど幅広いサービスに統合されている。2026年には新OS搭載の「Googlebook」にもGeminiがOS全層に組み込まれることが発表されており、開発者向けAPIも提供されている。
GeminiはOpenAIのChatGPTなどの競合モデルと何が違いますか?
Geminiはテキスト生成だけでなく画像・音声・動画を統合したマルチモーダル設計を持ち、Google独自のTPUインフラや検索・クラウド等の既存サービスとの深い統合が特徴だ。ただし大規模言語モデル間の性能差は業界全体で縮小傾向にある。
GeminiはセキュリティやAI防御にどのように活用されていますか?
Google Cloudは2026年5月に「AI Threat Defense」を発表し、Geminiを活用して脆弱性の発見・評価・修正・監視を自動パイプラインで処理する仕組みを提供するとしている。攻撃者によるAIの悪用が加速する中、パッチ生成までを自動化することを目標としている。
2026年のGoogle I/OではGeminiに関してどのような発表がありましたか?
ChromeOSとAndroidを融合しGeminiをOS全層に組み込んだ新デバイス「Googlebook」の2026年秋出荷が発表された。また、Geminiが複数サイトの買い物かごを一元管理する「Universal Cart」構想も公表され、AIエージェントによる自律的な購買支援を実現するとしている。

External Mentions

10 件