On September 4, 2026, Google rolled out its music-generation model "Lyria 3.5" to the Gemini app and the Gemini API. In the app, users can select a genre or describe one in text, choose vocal or instrumental, and use templates to create either short or long tracks. Google is promoting more expressive vocals and richer musical arrangements as the model's selling points.
However, calling this the debut of a new model would get the timeline wrong. Lyria 3.5 was already previewed on July 29 through Google Flow Music. What actually changed this time isn't the model's birthdate but where music generation now lives. By entering Gemini's image-and-video input flow and its developer-facing API, the model can now be tried through the same pathway for everything from short playful outputs to material for video, games, and podcasts.
What moved into Gemini after the 30-second clip?
Looking at the timeline, what's new about Lyria 3.5 this time isn't the model's initial appearance but the expansion of how it can be accessed. Google released Lyria 3.5 via Google Flow Music on July 29, then extended it to the Gemini app and Gemini API on September 4.
The starting point was Lyria 3 in February. In the Gemini app, users could create 30-second tracks—with lyrics or instrumental—using text, photos, or video as cues, and even add cover art with Nano Banana. At the time, Google framed the use case around sharing with friends and everyday self-expression, not replacing production workflows with finished tracks.
July's Lyria 3.5 arrived first in Google Flow Music as a model with improved control over musicality, lyrics, vocals, tempo, and length. This is an environment aimed at musicians and AI creators. Since 2023, Google has been developing its Music AI Sandbox together with musicians, producers, and songwriters, gathering feedback through experimental tools like Create, Extend, and Edit.
The September rollout connects this production-oriented model to the general user's chat interface and the developer's API entry point. In the Gemini app, templates shorten the distance to a first track; in the API, it becomes easier to embed the model into video editing, games, or in-app notification sounds. The change in distribution matters as much as any improvement in model performance.
How much of the "songwriting" process does Lyria 3.5 actually handle?
What sets Lyria 3.5 apart isn't stretching out audio fragments—it's the ability to specify a song's temporal structure directly within the input. Google's API documentation lists vocals, timed lyrics, and full instrumental arrangements, and explains that section tags like [Verse], [Chorus], and [Bridge], along with timestamps, can guide a song's progression.
Images can also serve as input. Through the API, users can pass up to 10 images alongside text, building music from colors, landscapes, or the mood of a subject. The Gemini app's "create a song from a photo" experience packages image understanding and music generation into a single action. Adding instrumentation, tempo, key, or mood to the text prompt lets users fine-tune the direction of the result.
According to the model card, Lyria 3.5 uses an approach that applies latent diffusion to a time-directional latent representation of audio. More important than the method's name is the fact that Google lists musical quality, vocals, audio fidelity, and prompt adherence as evaluation criteria. Google claims audio fidelity has improved compared to Lyria 2, and that adherence to complex instructions has improved for lyric-based generation—but it doesn't provide independent numerical benchmarks.
In short, Lyria 3.5 aims to move beyond simply turning text into sound, targeting output with actual song structure. But being able to specify structure and getting a finished result that matches your intent on the first try are not the same thing. Google itself advises checking whether generated tracks match the intended outcome.
The API's branching matters more than the 44.1kHz spec
According to the API specifications, Lyria 3 Clip generates a fixed 30-second MP3, while Lyria 3.5 generates full-length songs running several minutes, with the option to request WAV format for the latter.
These two are distinguished by different model IDs even though both belong to the "Lyria 3" family. lyria-3-clip-preview returns fixed-length output suited for short loops, previews, and social media material, while lyria-3.5 creates songs spanning several minutes that include intros, verses, choruses, and bridges. While Lyria 3.5's length can be specified via prompt, the API documentation describes it only as "several minutes"—not a maximum duration that applies uniformly across all conditions.
Output is 44.1kHz stereo audio, defaulting to MP3, with WAV available as an option for Lyria 3.5. Up to 10 images can be used as input, and users can write their own lyrics, guiding the song's flow using section tags and timestamps. For developers, the real design decision isn't comparing audio quality through marketing language—it's deciding which model ID to assign to quick short-form prototyping versus full-length production generation.
Meanwhile, the current version of Lyria 3.5 completes its work in a single generation pass. It doesn't support interactive editing, where a generated clip is gradually refined through additional prompts. Even with the same prompt, results can vary between calls. While the Music AI Sandbox has experimented with Extend and Edit functions for existing audio, Lyria 3.5 as called through the Gemini API currently requires generation and editing to remain separate processes.
Watermarking and prohibitions are not a substitute for rights clearance
Google embeds SynthID into every piece of audio generated by Lyria 3.5. This is an inaudible digital watermark that, according to Google, is designed to remain detectable even after MP3 compression, added noise, or playback speed changes. Google also provides a feature that lets users upload audio to Gemini to check whether it was generated or edited using Google AI.
There are also restrictions on the input side. The API documentation explicitly states that requests to replicate a specific artist's voice, and requests to generate copyrighted lyrics, are blocked by safety filters. The model card also lists uses that infringe on others' intellectual property rights or privacy as subject to its prohibited-use policy.
A distinction needs to be drawn here. SynthID is a mechanism for identifying generated content, while input filters reduce clearly imitative requests. Neither is described as a mechanism that automatically grants users clearance for third-party rights or commercial-use permissions. If you plan to publish or distribute generated audio, you need to separately check the service's terms of use and the rights conditions relevant to your specific use case.
What the model card discloses extends only to a description of the audio training data and preprocessing steps—not a list of specific tracks or licenses.
Google explains that it added text captions to audio data and performed deduplication along with safety and quality filtering. This is an explanation of data processing, not a disclosure that lets readers verify, track by track, under which contracts or legal basis each piece of audio was used for training. This gap remains a separate evaluation axis from the generation quality of AI music itself.
What determines usability is editing flexibility, not quality
For general users, Lyria 3.5 is easiest to understand as a feature for creating temporary video background music, birthday songs, short podcast intros, or personal ringtones. Google's Japanese help documentation lists usage conditions including being 18 or older, having activity saving turned on, and a cap on the number of generations allowed, and it directs users to upgrade to Pro in order to create full-length tracks. Downloads can be chosen as either an MP4 with cover art or an MP3 audio-only file.
This workflow shortens the time between having an idea and turning it into sound. Google AI Studio also offers a developer-facing demo that analyzes video to generate a description and then uses Lyria to generate background music. The more that video, text, images, and music are connected within the same suite of services, the easier it becomes to treat music not as an independent work but as one step in a content-production pipeline.
That said, using it for a long-form finished piece raises additional considerations: Can you edit only the part you want to fix? Can you reproduce the same lyrics or melody? Does the cap on the number of generations fit your production schedule? Do the publishing platform's terms and commercial-use conditions align? While the current version of Lyria 3.5 expands the model's generative capability, it leaves editing and rights verification to the user.
What Google needs to prove next isn't a matter of adjectives describing sound quality. Even after moving from short prototypes to longer songs, the real test is whether Google can consistently explain—across the model, the app, the API, and the terms of service—which parts humans need to fix and under what conditions the results can be published. Integration into Gemini has widened the entry point. Whether it remains useful beyond that point will be determined by post-generation editing and rights conditions.
