Google DeepMind has released the Lyra 3.5 music generation model, which supports generating full songs of up to 3 minutes in length.
July 30, Google DeepMind released the new generation music generation model Lyria 3.5, focusing on improving music structure, lyric quality, instruction compliance, vocal performance, and song duration control capabilities. It can directly generate full songs up to 3 minutes long, rather than just audio snippets lasting a few seconds.
According to the introduction, Lyria 3.5 still adopts the latent diffusion architecture, generating diffusion in the time-audio latent space, with training data consisting of audio with text annotations at different granularities. It underwent post-training through SFT combined with reinforcement learning incorporating human and Critic feedback, and all generated content is embedded with SynthID watermark. However, Google did not disclose the scale, source, and quantified benchmark of the training data, only indicating a significant improvement in audio clarity and lyric instruction compliance compared to Lyria 2.
Furthermore, Lyria 3.5 debuted integrated into Flow Music, which supports conversational music creation, stem separation, remixing, music distribution, playlist generation, and can collaborate with Veo to create music videos. It also supports audio plugin development, music games, and custom DAW, further integrating Lyria, Veo, and Gemini to build a comprehensive ecosystem covering music creation, editing, publishing, and distribution.