DocsQuick StartAI News
AI NewsLyria 3.5 Makes AI Songs Sound More Like Real Songs
New Model

Lyria 3.5 Makes AI Songs Sound More Like Real Songs

2026-07-29T18:04:10.049Z
Lyria 3.5 Makes AI Songs Sound More Like Real Songs

Google has launched Lyria 3.5, with Flow Music becoming the first to integrate it. The new version focuses on improving melodic coherence, lyrics, vocals, and creative control, but details on a standalone API, pricing, and model specifications have yet to be announced.

Lyria 3.5 Makes AI Songs Sound More Like Songs

Google DeepMind recently released its music generation model Lyria 3.5, initially making it available in Google Flow Music, a product aimed at music creators. Rather than focusing on longer durations or higher sample rates, this update tackles the four hardest-to-hide problems in AI music: whether the melody works, whether the lyrics can be sung clearly, whether the vocals sound natural, and whether creators can truly control the result.

As of July 29, 2026, Google has not disclosed Lyria 3.5’s model size, training data composition, standalone API model ID, pricing, or usage quotas in its announcement. In other words, this is clearly a product release, but not yet a full developer platform release. What can currently be confirmed is that Lyria 3.5 is available in Flow Music, while the existing Lyria 3 remains accessible through channels such as the Gemini API.

Interface for generating and editing songs with Lyria 3.5 in Google Flow Music, showing lyrics, song structure, and creative controls

This Upgrade Is Not Primarily About “Making Better Sounds”

Over the past two years, music generation models have already crossed the threshold of being able to generate a song at all. Entering a short description and receiving audio with vocals, lyrics, and an arrangement tens of seconds later is no longer a novel capability. What truly determines usability is whether the generated result can be integrated into the creative process.

Many AI-generated songs sound impressive on the first listen, only to reveal problems on the second:

  • The verse has a decent melody, but the chorus is not memorable;
  • The lyrics rhyme on paper, but their stresses are completely misplaced when set to music;
  • The voice sounds like it is singing, but the articulation, breathing, and emotion are inconsistent;
  • The user asks for “a stronger second chorus,” but the model also changes the chords, tempo, or even the singer;
  • Repeating the same prompt feels like drawing a new random card every time, making stable iteration impossible.

Lyria 3.5 aims to move music generation from “one-click finished-track roulette” toward “a creative system that can be revised repeatedly.” Google summarizes the improvements across four dimensions: musicality, lyrics, vocals, and creative control. Put more plainly, the goal is to make the music more coherent, the performance more believable, and the system more responsive to instructions.

Melody: From Locally Pleasing to Coherent as a Whole

Music generation can easily create an illusion: if a few seconds sound good, the entire song must be good.

In reality, a short clip can easily feel “finished” as long as the timbre is polished and the harmony is pleasant. A complete song, however, must handle longer-range relationships: how the verse builds emotion, how the chorus fulfills expectations, why the bridge belongs where it does, and how the second chorus can repeat without sounding copied and pasted.

This is why Google’s emphasis on improved musicality in Lyria 3.5 matters. The goal is not merely to make individual notes more pleasant, but to reduce structural looseness, melodic drift, and disconnected sections. Developers can think of it as the musical counterpart to contextual consistency: a large language model must remember what was established earlier, while a music model must remember motifs, tonality, rhythmic patterns, and the function of each section.

The earlier Lyria 3 could already analyze prompts before generation and infer structures such as intros, verses, choruses, bridges, and outros. For version 3.5 to represent a clear generational leap, this structural planning must influence the melody itself more deeply, rather than merely dividing the audio into sections according to labels.

This is also why Lyria 3.5 deserves more attention than another incremental improvement in audio quality. A 44.1 kHz stereo output is already sufficient for most previews, short videos, and online distribution scenarios. Further increases in output specifications offer limited marginal value; creating a chorus with a genuinely memorable hook is far more valuable.

Lyrics and Vocals Are Two Sides of the Same Problem

Lyrics cannot be evaluated as text alone. Words that read smoothly on a screen may not sing naturally.

Singable lyrics must simultaneously satisfy requirements involving meaning, syllables, stress, rhyme, and melodic placement. Chinese also presents conflicts between lexical tones and melodic contours, while English commonly suffers from errors in connected speech, swallowed sounds, and stress placement in multisyllabic words. If a model first writes ordinary text and then forcibly fits it into a melody, the result often features crammed words, excessively prolonged line endings, or an unimportant function word sung on the highest note.

Lyria 3.5 strengthens both lyrics and vocals, suggesting that Google is not treating them as independent modules. A more natural approach is to jointly generate lyric planning, melody, and performance: the model must decide not only “what to sing,” but also “how these words should be sung.”

Natural-sounding vocals are also about more than whether the timbre resembles a real person. The most obvious flaws emerge over time:

  1. Whether sustained notes remain stable or develop unnecessary wavering;
  2. Whether breaths occur in appropriate places;
  3. Whether the same singer maintains a consistent timbre across different sections;
  4. Whether the emotion develops with the arrangement instead of remaining at the same intensity throughout;
  5. Whether consonants, word endings, and backing vocals align with the beat.

These issues have limited impact on short-video soundtracks, but they are critical for complete songs. A 30-second clip can succeed on the strength of a catchy chorus, while a finished track lasting several minutes must make the singer feel “alive” throughout the entire piece.

Creative Control Determines Whether the Model Is a Toy or a Tool

The most important Lyria 3.5 upgrade to watch is actually creative control.

No matter how high the generation quality is, if users can only enter a prompt and wait for the result, the music model is still essentially a vending machine. Professional creation requires localized edits: keep the verse but redo only the chorus; preserve the melody while replacing the lyrics; maintain the singer and tempo while changing the arrangement from acoustic folk to synth-pop; or add another drum layer to the second section without changing the chords.

This kind of control is more difficult than “rewrite the third paragraph” in text generation because musical elements are tightly coupled. Changing the rhythm may affect the number of syllables in the lyrics, modifying the chords may force changes to the melody, and replacing an instrument may alter the overall dynamics. Control does not simply mean offering more style keywords. It means freezing other dimensions as much as possible when one dimension is modified.

Google has not yet fully disclosed Lyria 3.5’s control parameters or underlying mechanisms, so it is too early to claim that it supports digital audio workstation-style track-by-track editing. However, the decision to launch the model first in Flow Music rather than merely offering a generation button in a chat box already indicates the product direction: Google is targeting not ordinary users who occasionally generate birthday songs, but creators who repeatedly audition, revise, and compare versions.

A prompt better suited to Lyria should not simply say “generate a sad pop song.” It should resemble a production brief:

Style: Restrained electropop without a cinematic soundtrack feel
Tempo: Mid-tempo with a steady four-on-the-floor beat
Structure: Short intro → verse → pre-chorus → chorus → verse → chorus → bridge → final chorus
Vocals: Intimate female vocals, minimal vibrato in the verses, greater intensity in the chorus
Lyrics: Chinese, about leaving a city in the early hours of the morning; avoid abstract slogans
Arrangement: Verses led by electric piano and low-frequency pulses; add wide stereo synthesizers in the second chorus
Constraints: Maintain a consistent primary melodic motif; do not abruptly switch to rock in the bridge

The point is not how long the prompt is, but whether it clearly states the non-negotiable requirements. Lyria 3 already supports structural tags such as [Intro], [Verse], [Chorus], [Bridge], and [Outro]. The value of version 3.5 will depend on whether the model can follow these constraints more consistently, rather than merely recognizing the tag names.

Do Not Conflate Lyria 3.5 With Lyria 3—At Least for Now

Developers should pay particular attention to the distinction between product versions and API versions.

The existing Lyria 3 series mainly includes:

| Model | Model ID | Typical Use | Output Duration | | --- | --- | --- | --- | | Lyria 3 Clip | lyria-3-clip-preview | Short clips, loops, and concept previews | About 30 seconds | | Lyria 3 Pro | lyria-3-pro-preview | Complete songs containing verses, choruses, and bridges | Several minutes, controllable through prompts |

These two models can accept text or image inputs through the Gemini API’s Interactions API and generate 44.1 kHz stereo MP3 files. Image input is suitable for mapping visuals to musical moods. For example, users can upload a photo of a rainy street and have the model extract its cool color palette, sense of space, and rhythmic atmosphere.

However, Lyria 3.5 becoming available in Flow Music does not mean that lyria-3.5 is now a public API model ID. Until Google announces a standalone endpoint, SDK support, and billing rules, developers should not directly insert the product name into production code or assume that existing Lyria 3 Pro requests will automatically switch to version 3.5.

There is also a straightforward reason why this article does not provide a request example for a supposed “compatible API”: the currently available public information is insufficient to confirm the form of the Lyria 3.5 API. Music generation involves asynchronous tasks, binary audio, duration controls, and copyright safety metadata. You cannot simply replace a chat model’s name with Lyria 3.5 and treat it as an OpenAI-compatible API call. OpenAI Hub can be used to access general-purpose models such as GPT, Claude, and Gemini through a unified interface, helping generate lyrics, structural plans, and music prompts. As for a Lyria 3.5 audio endpoint, developers should wait for official integration rather than misleading others by including a model ID that does not yet exist in sample code.

Google’s Advantages Extend Beyond the Model Itself

Compared with music generation products such as Suno and Udio, Google may not lead in every individual aspect of the experience. Mature AI music products have already established workflows around communities, song extension, version management, and rapid regeneration. Users can enter a single sentence and receive a shareable finished track, making the consumer experience highly complete.

Google’s strengths lie elsewhere:

  • Broader entry points. Lyria is already integrated into products such as Gemini, YouTube Dream Track, Google Vids, AI Studio, and Flow Music;
  • A complete multimodal pipeline. Images, video, text, lyrics, and music can be connected within the same model ecosystem;
  • Clear distribution scenarios. Generated music can directly support Shorts, video production, and enterprise content;
  • A mature safety-marking system. The Lyria family follows the SynthID approach, embedding imperceptible but detectable watermarks in AI-generated audio;
  • An established developer platform. Once version 3.5 receives API access, the Gemini SDK, authentication systems, and cloud infrastructure can support it quickly.

This means Google does not necessarily need to win the market through a standalone “AI music app.” It can make music generation part of Gemini’s multimodal capabilities, then channel the resulting content into YouTube and workplace video tools. This kind of vertical integration is the hardest part for pure-play music startups to replicate.

However, large-company ecosystems also have a typical problem: the same model family is scattered across multiple products, with availability varying by region, account, and subscription tier. Product launches and API releases are often out of sync. Lyria 3.5 is currently at precisely this stage—creators can already access it in Flow Music, but developers still lack sufficiently complete API information.

The Real Test Is Not “Can It Generate?” but “Can It Be Edited?”

Judging from industry progress, the next stage of competition in AI music will not revolve solely around the quality of the initial generation. Generating a work that “sounds like a song” is gradually becoming commoditized. The real barriers will emerge in three areas:

  • Editability: Whether specific parts can be modified without damaging the entire song;
  • Identity consistency: Whether the same virtual singer can maintain the same timbre and performance habits across multiple songs;
  • Rights and provenance: Whether enforceable rules can be established around training data, style imitation, voice authorization, and generated-content labeling.

Lyria 3.5’s concentrated improvements to melody, lyrics, vocals, and control are moving in the right direction. They show that Google has recognized that the bottleneck in music generation is no longer decoding a piece of audio, but building a feedback loop that resembles a real creative process.

However, until Google discloses more specific details about control granularity, long-song consistency, generation speed, pricing, and API availability, Lyria 3.5 is better understood as a substantial product upgrade than as a foundation-model generation that developers can migrate to immediately.

In one sentence: Lyria 3.5 is transforming AI music from “randomly giving you a song” into “writing a song with you according to your requirements.” The direction matters more than the parameters, but whether it ultimately becomes a true productivity tool will depend on when Flow Music’s full control capabilities reach the API.

References

Google DeepMind’s official Lyria 3.5 announcement and the Gemini API documentation are the primary sources for this article. Due to link-domain restrictions, the corresponding external links are not included here. The following official SDK repositories are accessible from within China and can be used to understand the Gemini API’s client implementations.

  • Google Gen AI Python SDK: Google’s official Python SDK, including client implementations and examples for the Gemini API, Interactions, and other interfaces.
  • Google Gen AI JavaScript SDK: Google’s official JavaScript/TypeScript SDK, which can be used to track model interfaces, asynchronous tasks, and multimodal input support.

Related Articles

View All

Contact Us

We usually reply quickly during business hours

Scan WeChat

Support: Hub Assistant

WeChat ID: