Tech Product

Sora

別名: Sora

comune.sora.fr.it

Overview

最終更新: 2026年7月9日

SoraはOpenAIが開発した動画生成AIモデルであり、テキストプロンプトを入力することで高品質な動画を自動生成する能力を持つ。拡散モデルとトランスフォーマーアーキテクチャを組み合わせた技術基盤を持ち、テキストから映像への変換において高い表現力を示す。

概要

Soraは、ユーザーが自然言語で記述したシナリオや情景をもとに、最大数十秒規模の動画コンテンツを生成できるAIモデルである。OpenAIがこれまで構築してきた大規模言語モデルおよび画像生成技術の延長線上に位置づけられており、映像のフレーム一貫性や物体の物理的な挙動を学習した上で動画を合成する点が特徴とされている。生成AIの応用領域を静止画・テキストから動画へと拡張する取り組みとして、業界内外から注目を集めた。

技術的位置づけ

Soraの技術的な核心は、映像を時空間パッチに分割して扱うトランスフォーマーベースのアーキテクチャにある。この手法により、異なる解像度・アスペクト比・長さの動画をある程度柔軟に扱えるとされる。従来の動画生成モデルが短時間のクリップ生成に留まりがちであったのに対し、Soraはより長尺かつ複雑な映像シーンの生成を目指した設計となっている。また、テキストのみならず、静止画や既存動画を入力として受け付ける機能も備えており、映像編集・拡張のユースケースへの応用も想定される。

主要な動向

2026年6月には、OpenAIのAIコーディングツール「Codex」がGPT-5を活用した自己進化的開発プロセスに移行しているという事実が報じられた際、SoraのモバイルアプリケーションがAIエージェントによってわずか18日間で構築されたという事例が明らかになった。これはAIが別のAIシステム向けソフトウェアを短期間で実装できる段階に達しつつあることを示す具体的な事例として注目された。また、OpenAIのAPIが処理するトークン量は2025年10月の毎分60億から2026年3月末には毎分150億へと約2.5倍に増加しており、Soraを含むOpenAIの各種サービスへの需要拡大が計算資源の逼迫という課題を生んでいる状況が続いている。動画生成AIの競合環境も激化しており、GoogleやMeta、さらには複数のスタートアップが同領域への参入を進める中、Soraはその先行者優位をどこまで維持できるかが引き続き問われている。

Mentioned Articles

20 件

Research Papers

5 件
  • Open-Sora: Democratizing Efficient Video Production for All

    Zangwei Zheng, Xiangyu Peng, Tianji Yang, Chenhui Shen, Shenggui Li, Hongxin Liu, Yukun Zhou, Tianyi Li, Yang You

    2024664 件引用Semantic Scholar

    Vision and language are the two foundational senses for humans, and they build up our cognitive ability and intelligence. While significant breakthroughs have been made in AI language ability, artificial visual intelligence, especially the ability to generate and simulate the world we see, is far lagging behind. To facilitate the development and accessibility of artificial visual intelligence, we created Open-Sora, an open-source video generation model designed to produce high-fidelity video content. Open-Sora supports a wide spectrum of visual generation tasks, including text-to-image generation, text-to-video generation, and image-to-video generation. The model leverages advanced deep learning architectures and training/inference techniques to enable flexible video synthesis, which could generate video content of up to 15 seconds, up to 720p resolution, and arbitrary aspect ratios. Specifically, we introduce Spatial-Temporal Diffusion Transformer (STDiT), an efficient diffusion framework for videos that decouples spatial and temporal attention. We also introduce a highly compressive 3D autoencoder to make representations compact and further accelerate training with an ad hoc training strategy. Through this initiative, we aim to foster innovation, creativity, and inclusivity within the community of AI content creation. By embracing the open-source principle, Open-Sora democratizes full access to all the training/inference/data preparation codes as well as model weights. All resources are publicly available at: https://github.com/hpcaitech/Open-Sora.

  • Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models

    Yixin Liu, Kai Zhang, Yuan Li, Zhiling Yan, Chujie Gao, Ruoxi Chen, Zhengqing Yuan, Yue Huang, Hanchi Sun, Jianfeng Gao, Lifang He, Lichao Sun

    2024625 件引用Semantic Scholar

    Sora is a text-to-video generative AI model, released by OpenAI in February 2024. The model is trained to generate videos of realistic or imaginative scenes from text instructions and show potential in simulating the physical world. Based on public technical reports and reverse engineering, this paper presents a comprehensive review of the model's background, related technologies, applications, remaining challenges, and future directions of text-to-video AI models. We first trace Sora's development and investigate the underlying technologies used to build this"world simulator". Then, we describe in detail the applications and potential impact of Sora in multiple industries ranging from film-making and education to marketing. We discuss the main challenges and limitations that need to be addressed to widely deploy Sora, such as ensuring safe and unbiased video generation. Lastly, we discuss the future development of Sora and video generation models in general, and how advancements in the field could enable new ways of human-AI interaction, boosting productivity and creativity of video generation.

  • Open-Sora Plan: Open-Source Large Video Generation Model

    Bin Lin, Yunyang Ge, Xinhua Cheng, Zongjian Li, Bin Zhu, Shaodong Wang, Xianyi He, Yang Ye, Shenghai Yuan, Liuhan Chen, Tanghui Jia, Junwu Zhang, Zhenyu Tang, Yatian Pang, Bin She, Cen Yan, Zhiheng Hu, Xiao-wen Dong, Lin Chen, Zhang Pan, Xing Zhou, Shaoling Dong, Yonghong Tian, Li Yuan

    2024271 件引用Semantic Scholar

    We introduce Open-Sora Plan, an open-source project that aims to contribute a large generation model for generating desired high-resolution videos with long durations based on various user inputs. Our project comprises multiple components for the entire video generation process, including a Wavelet-Flow Variational Autoencoder, a Joint Image-Video Skiparse Denoiser, and various condition controllers. Moreover, many assistant strategies for efficient training and inference are designed, and a multi-dimensional data curation pipeline is proposed for obtaining desired high-quality data. Benefiting from efficient thoughts, our Open-Sora Plan achieves impressive video generation results in both qualitative and quantitative evaluations. We hope our careful design and practical experience can inspire the video generation research community. All our codes and model weights are publicly available at \url{https://github.com/PKU-YuanGroup/Open-Sora-Plan}.

  • Open-Sora 2.0: Training a Commercial-Level Video Generation Model in $200k

    Xiangyu Peng, Zangwei Zheng, Chenhui Shen, Tom Young, Xinying Guo, Binluo Wang, Hang Xu, Hongxin Liu, M. Jiang, Wenjun Li, Yuhui Wang, Anbang Ye, G. Ren, Qianran Ma, Wanying Liang, Xiangru Lian, Xiwen Wu, Yu Zhong, Zhuangyan Li, Chaoyu Gong, Guojun Lei, Lei Cheng, Liming Zhang, Minghao Li, Ruijie Zhang, Silan Hu, Shijie Huang, Xiaokang Wang, Yuanheng Zhao, Yuqi Wang, Ziang Wei, Yang You

    2025129 件引用Semantic Scholar

    Video generation models have achieved remarkable progress in the past year. The quality of AI video continues to improve, but at the cost of larger model size, increased data quantity, and greater demand for training compute. In this report, we present Open-Sora 2.0, a commercial-level video generation model trained for only $200k. With this model, we demonstrate that the cost of training a top-performing video generation model is highly controllable. We detail all techniques that contribute to this efficiency breakthrough, including data curation, model architecture, training strategy, and system optimization. According to human evaluation results and VBench scores, Open-Sora 2.0 is comparable to global leading video generation models including the open-source HunyuanVideo and the closed-source Runway Gen-3 Alpha. By making Open-Sora 2.0 fully open-source, we aim to democratize access to advanced video generation technology, fostering broader innovation and creativity in content creation. All resources are publicly available at: https://github.com/hpcaitech/Open-Sora.

  • Is Sora a World Simulator? A Comprehensive Survey on General World Models and Beyond

    Zheng Zhu, Xiaofeng Wang, Wangbo Zhao, Chen Min, Nianchen Deng, Min Dou, Yuqi Wang, Botian Shi, Kai Wang, Chi Zhang, Yang You, Zhaoxiang Zhang, Dawei Zhao, Liang Xiao, Jian Zhao, Jiwen Lu, Guan Huang

    2024107 件引用Semantic Scholar

    General world models represent a crucial pathway toward achieving Artificial General Intelligence (AGI), serving as the cornerstone for various applications ranging from virtual environments to decision-making systems. Recently, the emergence of the Sora model has attained significant attention due to its remarkable simulation capabilities, which exhibits an incipient comprehension of physical laws. In this survey, we embark on a comprehensive exploration of the latest advancements in world models. Our analysis navigates through the forefront of generative methodologies in video generation, where world models stand as pivotal constructs facilitating the synthesis of highly realistic visual content. Additionally, we scrutinize the burgeoning field of autonomous-driving world models, meticulously delineating their indispensable role in reshaping transportation and urban mobility. Furthermore, we delve into the intricacies inherent in world models deployed within autonomous agents, shedding light on their profound significance in enabling intelligent interactions within dynamic environmental contexts. At last, we examine challenges and limitations of world models, and discuss their potential future directions. We hope this survey can serve as a foundational reference for the research community and inspire continued innovation. This survey will be regularly updated at: https://github.com/GigaAI-research/General-World-Models-Survey.

よくある質問

Soraとは何ですか?
SoraはOpenAIが開発したAIモデルで、テキストプロンプトを入力することで高品質な動画を自動生成できる。映像のフレーム一貫性や物体の物理的挙動を学習した上で動画を合成する点が特徴とされている。
Soraはどのような技術で動作していますか?
映像を時空間パッチに分割して処理するトランスフォーマーベースのアーキテクチャを採用している。テキストだけでなく、静止画や既存動画を入力として受け付ける機能も持ち、映像編集や拡張への応用も想定されている。
Soraアプリ開発に関して最近の動きはありますか?
2026年6月の報道によれば、OpenAIのAIエージェント「Codex」がSoraのモバイルアプリをわずか18日間で構築したことが明らかになった。AIが別のAIシステム向けソフトウェアを短期間で実装できる段階に達した事例として注目された。
Soraの競合となるサービスはありますか?
動画生成AI分野にはGoogleやMetaなど大手テック企業に加え、複数のスタートアップが参入しており、競合環境は激化している。Soraはこの分野における先行サービスの1つとして位置づけられている。
Soraの利用拡大はOpenAIのインフラにどう影響していますか?
OpenAIのAPIが処理するトークン量は2025年10月の毎分60億から2026年3月末には毎分150億へと約2.5倍に増加しており、Soraを含む各種サービスへの需要拡大が計算資源の逼迫という課題を生んでいる。

External Mentions

10 件