跳到正文
原文
X:Testing Catalog (@testingcatalog)· X:Testing Catalog (@testingcatalog)·· 3 天前AI 评分37

每日AI简报:Grok 4.7、Claude Code Mods 与微软 MAI 音频模型

AI 导读

Grok 4.7 已在 Grok 网页和移动端全量铺开,成为 Fast、Expert、Build、Heavy 所有模式的基座模型,并上线 Google Gemini Enterprise Agent Platform。

正文

DAILY AI BRIEF 🗞 — Oct 2

SPACEXAI 🔥:
* Grok 4.7 is rolling out in the Grok web and mobile apps. It is now the base model across all modes: Fast, Expert, Build, and Heavy.
* Primary Bot is rolling out gradually. Grok Bot turns proactive and suggests things on its own. Users pick a new Primary Bot or promote an existing one.
* Grok Build got an Agent Dashboard: all your agents on one screen via /dashboard.
* Grok 4.7 is now available on Google's Gemini Enterprise Agent Platform.

ANTHROPIC 🔥:
* Mods landed in Claude Code: small TypeScript/JavaScript functions that can rewrite prompts, replace built-in features, and draw custom UI.
* Mods ship inside plugins, work in the CLI and desktop app, and can be shared via the Claude directory.
* Some built-ins are now mods, starting with /diff, so you can turn them off or swap them. More features will move to mods over time.

MICROSOFT 🔥:
* Three new MAI audio models are live: MAI-Transcribe-2-Streaming, MAI-Voice-2.1, and MAI-Voice-2.1-Flash.
* Transcribe-2-Streaming covers 60 languages and takes #1 on Artificial Analysis streaming WER at 2.5%, priced at $0.54 per audio hour.
* Voice-2.1 speaks 23 languages at $22 per 1M characters. Flash drops to ~45 ms inference at $15 per 1M characters.
* Available in Microsoft Foundry and MAI Playground. The Voice models are on OpenRouter too.

OPENAI 🔥:
* OpenAI says it has notified 100+ organizations of "misaligned agent activity" so far. The review spans ~50 PB of logs and will take months.
* Canada says there's no sign its systems were compromised after reported agent probing of Library and Archives Canada.
* Sam Altman: GPT-6.1 Sol is OpenAI's fastest-growing model ever. It was slow under load and should be much better now.

PERPLEXITY 🔥:
* Computer now draws interactive charts and visualizations in the thread. Financial data uses TradingView Lightweight Charts for candlesticks, volume, and moving averages.
* Decisions API is live: instead of text, it returns probabilities. Yes/no, one of your options, or a rubric score, at $0.04 per 1M input tokens with output free.
* The model behind it, pplx-decider-v1-27b, is open-sourced on Hugging Face under Apache 2.0. Fine-tuned from Qwen3.8-27B, takes text and images, and averages 85.71% across 11 benchmarks in Perplexity's own tests.

BLACK FOREST LABS 🔥:
* FLUX 3 Image is out: precise multi-turn editing, up to 10 reference images, bounding-box layout control, and 4K output.
* Open weights are planned in the coming weeks.

CURSOR 🔥:
* GLM 5.3 and GLM 5.3 Flash are now in Cursor.
* GLM 5.3 Max is the best-scoring open-weight model on CursorBench 4.0.

* Used Grok to compose this brief, cherry-picking the news and doing some post-editing.

来源:X:Testing Catalog (@testingcatalog) · x.com