MiniMax H3
MiniMax H3 is a general-purpose omni-modal generation model that jointly understands multimodal contexts spanning text, images, video, and audio. It generates videos with native stereo sound at up to 2K resolution and 15 seconds in length, delivering content for advertising, branding, ecommerce, product design, UI/UX, gaming, and creative workflows. Users can combine reference types in one instruction, for example, transferring camera movement from a video, placing a character from an image into the scene, and matching vocals from an audio clip, while describing the relationships in natural language. H3 supports text-to-image, text-to-video with jointly generated audio, multi-shot modeling, text-to-audio, and generalized reference and editing across images, videos, and audio. Voice, sound effects, and music are modeled together. The model excels at instruction following, accurate text and brand presentation, and video-to-video motion transfer.
Learn more
Grok Imagine Video 1.5
Grok Imagine Video 1.5 is xAI’s improved image-to-video model, built for better quality at faster speeds. Now generally available on the Imagine API as grok-imagine-video-1.5, it gives creators and developers a way to start from an image, describe the motion, and choose the resolution and duration for the generated video. Grok Imagine Video 1.5 and Video 1.5 Fast are described as xAI’s best image-to-video models yet, with better motion, better physics, better audio, and faster generation for real creative work. Audio and speech are generated in the same pass as the visuals, so sound effects, ambience, and dialogue land on the action, while speech is clearer and better synchronized. Motion and physics are also improved, helping movement hold together across the length of a clip with fewer warps and more believable weight and momentum. Grok Imagine Video 1.5 Fast almost doubles generation speed, producing 6-second, 720p videos in about 25 seconds.
Learn more
Lyria 3
Lyria 3 is Google DeepMind’s most advanced AI music generation model, designed to create high-fidelity, professional-grade audio from simple prompts. It enables users to describe a track in natural language and refine details such as tempo, vocal style, and instrumentation for greater creative control. The model can generate cohesive songs that flow naturally from start to finish across a wide range of genres and global languages. Lyria 3 also supports image-to-music composition, allowing users to upload visuals and transform them into custom soundtracks. Built with input from musicians and producers, it understands rhythm, arrangement, and musical structure at a deeper level. Users can export crisp, polished tracks suitable for background ambience, content creation, or mainstage productions. Integrated into Gemini and other creative tools, Lyria 3 empowers creators to explore, experiment, and express ideas through AI-driven music.
Learn more
Pika SFX
Pika SFX is a sound-effects generation model that turns natural-language direction into focused, ready-to-use audio for video, games, editing workflows, and creative tools. Users simply describe the sound they want, whether it is a glass shattering, a metal door slamming in an empty warehouse, a cork popping and fizzing into a glass, or something more stylized. The model can handle both a single crisp event and longer sequences, with control over material, space, perspective, timing, texture, and mood. It can generate natural Foley, cartoon-style effects, designed fantasy sounds, environmental ambience, and other prompt-driven audio while following the brief closely. Unless requested, it avoids adding unrelated speech, music, or background noise, helping creators generate clean effects that can be dropped directly into a project.
Learn more