Gemini Omni
Gemini Omni
Google DeepMind 于 2026 年 5 月 19 日发布了新一代多模态大模型 Gemini Omni。该模型整合了文本、图像、音频与视频的理解与生成能力,旨在实现更自然的人机交互。在同期 Hacker News 讨论中获得 112 点赞,显示出技术社区对多模态融合趋势的持续关注。这标志着大模型从单模态向全模态感知与响应能力的进一步演进。
Gemini Omni 把视频编辑变成自然语言对话,多轮编辑和物理理解让它从玩具变成创作工具,做视频的值得一试。
Gemini Omni is where Gemini’s ability to reason meets the ability to create. It delivers a leap in world understanding, multimodality, and editing.
Prompt: Make it look like the weird shape of my hand hole super zooms and magnifies the ground it's looking at in sharper quality.
Prompt: When the finger in <video> touches the animal toy play the sound the animal makes
Prompt: The lights of the apartments start turning on in sync with the music.
Input video
Prompt: Transport the violinist to the image environment
Prompt: Make the violin invisible
Prompt: Change the camera angle to be over the violinist’s shoulder.
Prompt: Change spaceship to <object>
Prompt: A marble rolling fast on a chain reaction style track, continuous smooth shot
Prompt: claymation explainer of protein folding, everything is made out of clay, no hands, stop motion, accurate
Prompt: A skeuomorphism stop motion explainer about how the brain hippocampus works with a compelling voiceover. Don’t add seahorses. No voice cuts at the end. Don’t add text.
Prompt: The video shows items of the alphabet. An unusual item starting with each letter is shown sitting on a table (like a Capybara for C, disco globe for D and Lava Lamp for L). All 26 letters must be represented by 26 items with matching lower thirds displaying the letter. Only one item and lower third at a time. Each lower third must look like a black marker written on a slip of paper in the bottom left. Rapid fire, roughly 9 frames per item at 24FPS. Last frame is a slip of paper "THE END". The whole video is accompanied by calm smooth music.
Prompt: word by word, one word on a the screen at a time: did, you, know, that, this, model, can, do, pretty, good, text!? each word appears with a different animated style, perfect pacing to a rhythm, sizzle reel
Prompt: The camera zooms into the TV screen, where we see the same woman and the same scene from the beginning. Seamless video. One continuous shot, no jump cuts.
Prompt: A close-up low-angle shot of a stylish drummer in a beige suit playing a red drum kit in a grand hall transitions as the camera whip-pans to the side, revealing an older saxophonist playing alongside a ballet dancer spinning in a white outfit under soft purple stage lights. One continuous shot, no jump cuts.
Creating your prompts
Use our prompt guide to create realistic, coherent, and creative output.
Slide 1 of 5
[1] Human raters conducted direct side-by-side comparisons across 504 diverse examples (each including text prompts for editing, evaluating one generated video per example).
Editing 3P 10s
Video Editing
Omni's Editing capability has achieved leading results for: Overall Preference and Instruction Following in head-to-head comparisons by human raters against other leading video generation models on internal benchmarks [1].
Movie Bench Audio T2VA Public Benchmark
Text to Video (Overall Preference & Instruction Following)
Participants viewed 1,003 prompts and respective videos on MovieGenBench, a benchmark dataset released by Meta. Omni performs best on Overall Preference and Instruction Following.
T2V Fast Motion
Text to Video (Fast Motion)
The Fast Motion evaluation dataset consists of 500 detailed text prompts designed to describe a wide variety of sports, athletic performances, and other activities. This dataset serves as a comprehensive benchmark for assessing the ability of AI models to interpret and generate highly dynamic, high-energy physical actions and diverse sporting events.
I2V VBench
Image to Video
Participants viewed 355 image and text pairs from the VBench I2V benchmark, outputs across Gemini Omni Flash, Grok-Imagine-Video, and Kling tied, leading over other models.
[1] Human raters conducted direct side-by-side comparisons across 468 diverse examples (each including a mix of images, audio, and text prompts as references, evaluating one generated video per example).
I+AR2V Image & Audio
Reference to Video
Omni's Reference to Video capability has achieved leading results for: Overall Preference and Speech Adherence in head-to-head comparisons by human raters against other leading video generation models on internal benchmarks [1].
Training/development evaluations including automated and human evaluations carried out continuously throughout and after the model’s training, to monitor its progress and performance
Human red teaming conducted by specialist teams who sit outside of the model development team, across the policies and desiderata, deliberately trying to spot weaknesses and ensure the model adheres to safety policies and desired outcomes
Automated red teaming to dynamically evaluate Gemini Omni Flash for safety and security considerations at scale, complementing human red teaming and static evaluations
Ethics and safety reviews conducted ahead of the model’s release
Content created or edited with Omni in the Gemini app, Google Flow or YouTube includes our imperceptible SynthID digital watermark and C2PA Content Credentials. You can easily verify content through the Gemini app and coming soon to Chrome and Search. You can find out more about how we're expanding our content transparency and verification tools to help you understand how content was created and edited across the web in our blog post.
Gemini
Supercharge your creativity and productivity
Google Flow
An AI creative studio built with and for creatives
YouTube Shorts
A shorter way to discover, watch, and create on YouTube
Google Vids
AI-powered video creation for work
Google AI Studio
The fastest path from prompt to production
Gemini API
Get started building with cutting-edge AI models
Google Enterprise Agent Platform
Build, scale, and govern agents
Google AI subscription required. Features vary by tier and geography.
来源:Hacker News 热门(buzzing.cc 中文翻译) · deepmind.google