ByteDance has announced a new model, "Seedance 2.5," that generates audio and video simultaneously. A single generation can now produce up to 30 seconds of video, and up to 50 reference materials can be passed into one clip. Along with the phased rollout, Dreamina will also offer a feature that lets users select and revise specific timestamps or specific people within a video. This update pushes the tool beyond a generator for repeatedly redoing short clips and toward a production tool for assembling longer scenes while preserving direction.
What deserves emphasis is not simply that new input formats have been added. The previous generation, Seedance 2.0, already handled text, images, audio, and video within a single generation system. Version 2.5 expands how much material can be fed into that system and narrows down what can be fixed after generation. Three changes—length, references, and editing—are connected within the same production workflow.
From 15 Seconds to 30 Seconds: The Generation Ceiling Doubles
According to the technical report for Seedance 2.0, that model could directly generate audio-visual video between 4 and 15 seconds long, at native resolutions of 480p or 720p. Version 2.5 can generate up to 30 seconds in a single pass. The upper time limit has doubled from 15 to 30 seconds.
With the move to 30 seconds, the stretch over which consistency must be maintained in a single generation also grows longer. Beyond a person's appearance and motion, continuity of lighting and camera work must now be sustained over a longer range than before. Dreamina states that scene transitions are smoother and character consistency has improved, and Seed's model page also claims improvements in motion stability and visual realism. However, this announcement includes no figures or comparison conditions indicating the degree of improvement. At this stage, the quality gains are ByteDance's own assessment.
Seed's model page also states clearly that a generated video can be extended twice. It does not specify how many seconds each extension adds, so the total possible length cannot be determined. The single-pass 30-second generation and the later extension feature need to be considered separately.
Bringing 50 References into a Single Clip
Seedance 2.5 can accept up to 50 references per clip, combining images, video, audio, and text. Product photos or storyboards can be combined with motion samples and audio, all fed into the same generation together with directorial instructions. This reduces the burden of having to reproduce shooting intent through text alone, making it easier for advertisements and short-form videos to convey camera and sound intentions while preserving the appearance of characters and products.
The previous generation's public platform allowed references of up to 3 videos, 9 images, and 3 audio clips—totaling 15 non-text materials. Version 2.5 has expanded the total to 50, but it has not disclosed how many images or videos can be used individually. Moreover, since text is also counted within the 50, it cannot be concluded that input capacity has simply more than tripled.
Even so, the direction of the change is clear. If 2.0 was the generation that unified multiple formats into a single model, 2.5 is the generation that bundles numerous production materials into a single clip. As more is entrusted to generative AI, there is a growing need for functions that determine priority when references conflict with one another, and that let creators verify which material actually influenced the result. ByteDance explains that the model understands the intent, composition, and visual expression of reference videos more accurately, but has not yet disclosed how this is evaluated.
Fixing Timestamps and People Before Regenerating Everything
The editing feature being added to Dreamina lets users specify the timestamp, person, or other element they want to revise. With conventional generation, fixing even a minor flaw often meant redoing the entire clip, sometimes altering motion or lighting that the creator had liked. If this partial specification works as intended, users can iterate by narrowing in on the problematic section alone.
The ability to edit a specified person or similar element was already documented in the 2.0 technical report, which described modifications targeting elements such as clips, actions, and story elements, as well as video extension. What comes to the forefront with 2.5 is the granularity of directly selecting a timestamp or person within Dreamina and revising it interactively. This should be understood not as the debut of the concept of partial editing itself, but as an evolution that makes it possible to select a target directly from the production screen.
Seed's model page explains that the process has evolved from simply transferring motion from a reference video to a creative transformation that reads composition and camera expression. In addition to green-screen editing, it claims to support professional camera work and staged performance as well. ByteDance positions these as features that support complex video production and professional production workflows.
In production settings, value is determined not only by how good the initial output looks, but by whether revisions can be made repeatedly. Will a person's appearance hold up over 30 seconds? Can fixing one specific part preserve the surrounding audio and lighting? Will important instructions be correctly interpreted out of 50 references without confusion? Version 2.5 expands the scope of what must be tested—from whether a generated result is usable, to whether it can be brought closer to a finished deliverable.
A Phased Rollout, with Performance Conditions Still to Come
In its announcement dated July 31, 2026, Dreamina stated that Seedance 2.5 would be rolled out in phases starting the following week. The target regions include Europe and Asia, as well as the Middle East and South America. A subscription account is required, and users must be over 16 years old. As of the announcement, not all accounts will be able to use it simultaneously; availability will need to be confirmed by region and by account.
As a transparency measure for generated content, Dreamina combines invisible watermarking, C2PA Content Credentials, and visible AI labeling. For harmful content or the unauthorized use of a person's likeness, the company says it will respond through prompt blocking, review of generated outputs, and user reporting. It also touts technology intended to prevent the unauthorized generation of intellectual property. None of these measures guarantee complete prevention of harm; they are designed around repeated detection and enforcement.
The model page for 2.5 and Dreamina's announcement do not include model scale, training data, or comparative benchmarks. Native resolution and generation speed also remain unknown. Pricing, credit consumption per generation, and API availability conditions have not been disclosed either. Because the 2.0 technical report clearly stated output conditions of 480p/720p and 4–15 seconds, once a technical report or API specification for 2.5 is released, it should become possible to compare the relationship between computational cost and image quality that comes with extending output to 30 seconds.
There are also gaps in how evaluation methods are disclosed. The 2.0 technical report described a framework in which motion stability was evaluated automatically, while aesthetic quality was passed to blind evaluation by human experts. For editing, it measured how well unmodified regions were preserved, calling this "edit consistency." Version 2.5 markets exactly this kind of partial editing as a selling point, but the corresponding scores or testing conditions have not been shown this time.
Whether Seedance 2.5 establishes itself as a genuine production tool will be determined not by how impressive public demos look, but by the rate at which the same instructions can be reproduced and the success rate of partial edits. What should be checked once the phased rollout is underway is consistency of people and sound throughout a full 30 seconds, how well areas outside the edited section are preserved, and the time and cost required to complete a single piece of video.
