Google DeepMind 推出 Gemini 3.8 Live with Live Avatar,为对话模型加入实时视频形象
Google DeepMind 发布 Gemini 3.8 Live with Live Avatar,把近实时视频生成与语音结合,为实时对话模型加入能听、能看、能说的可视化形象,现已在 Gemini Enterprise 提供。
Google DeepMind 发布 Gemini 3.8 Live with Live Avatar,把近实时视频生成与语音结合,为实时对话模型加入能听、能看、能说的可视化形象,现已在 Gemini Enterprise 提供。
Sarvam AI 发布新一代语音识别模型 Saaras V4,采用音频编码器加自研 3B 混合状态空间 LLM 解码器,同一模型可输出 transcribe、translate、verbatim、translit、codemix 五种转写格式。