← 首页
25JUL
AI前沿每日脉动
AI Frontier Pulse · 中英双语版 Bilingual Edition
2026.07.25 · 周六刊
21 位 Builder48 条推文1 期播客1 篇博客
Open ModelsOpus 5Prompt EngineeringAgent ContainmentWork AI
Richard Liu · 2026
Curated by Richard Liu · Snapshot follow-builders-2026-07-25-v2
今日头条 · Open And Frontier Model Competition01 / 15
今天最强信号:开放模型与专有模型不再是路线之争,而是同一场 AI 竞争中的互补战略。The strongest signal: open and proprietary models are becoming complementary strategies in the same AI race.
@samaX 原文11.5K ❤ · 736 RT · 1.8K 💬
Sam Altman:美国 AI 竞争必须同时赢得开放与专有模型Open And Proprietary Models Both Matter
Sam Altman 表示,美国要赢得 AI 竞争,开放模型与专有模型两条路线都不能缺席。今天的模型格局不再是二选一,而是能力前沿、生态扩散与开发者控制权的并行竞赛。
Sam Altman argues that winning in AI requires strength in both open-source and proprietary models. The contest now spans frontier capability, ecosystem reach, and developer control.
@samaX 原文11.5K ❤ · 736 RT · 1.8K 💬
Claude Code 为新模型删掉约 80% 系统提示词Claude Code Cuts Its System Prompt By About 80%
Thariq 透露,面向最新模型,Claude Code 移除了约 80% 的 system prompt。模型能力上升后,提示词、skills 与 Claude.md 的设计重点从堆叠规则转向提供清晰目标、必要上下文与可复用流程。
Thariq says Claude Code removed roughly 80% of its system prompt for newer models. As models improve, prompts, skills, and Claude.md files can emphasize clear goals, essential context, and reusable procedures.
@trq212X 原文10.2K ❤ · 1.1K RT · 255 💬
02 / 15
模型发布 · Opus 5 Lands02 / 15
Opus 5:跨知识工作升级,注入抗性成为更重要指标Opus 5 Makes Prompt-injection Resistance A Headline Metric
Boris Cherny 认为 Opus 5 在编码、数据分析、设计和知识工作上都很强,但最值得关注的是它成为 Anthropic 迄今最不易受 prompt injection 影响的模型。安全表现正在从系统卡附注走向产品选择指标。
Boris Cherny highlights Opus 5 across coding, analysis, design, and knowledge work, but calls its prompt-injection resistance the more important advance. Safety behavior is becoming a product-selection metric.
@bchernyX 原文5K ❤ · 383 RT · 269 💬
Opus 5 成为 Claude 5 家族的日常主力Opus 5 Becomes A Daily Driver
Thariq 将 Opus 5 定位为 Claude 5 家族的日常主力,并建议与 Fable 配合处理规划、头脑风暴和最难的 bug。产品分层正在从“单模型万能”转向多模型协作。
Thariq positions Opus 5 as a daily driver and pairs it with Fable for planning, brainstorming, and the hardest bugs. Product strategy is shifting toward coordinated model portfolios.
@trq212X 原文3.7K ❤ · 97 RT · 359 💬
03 / 15
推理产品 · Capability Must Become A Clear Tier03 / 15
更长推理、更高算力与更强模型,只有被包装成用户能理解的速度、价格和完成度,才会真正成为产品。More reasoning only becomes a product when users can feel it through speed, price, and task completion.
@samaX 原文3.1K ❤ · 89 RT · 436 💬
OpenAI 邀请用户测试更高难度推理模式OpenAI Tests A Harder Reasoning Mode
Sam Altman 邀请用户反馈新的高难度推理体验。名称带有实验色彩,但信号明确:前沿产品仍在探索如何把更高算力、更长推理与用户可感知的质量提升包装成清晰档位。
Sam Altman asks users to test a harder reasoning mode. The playful naming masks a serious product question: how to package more compute and longer reasoning into a clear, useful tier.
@samaX 原文3.1K ❤ · 89 RT · 436 💬
Opus 5 同价上线,并提供约 2.5 倍 Fast modeOpus 5 Ships At The Same Price With A 2.5x Fast Mode
Claude 宣布 Opus 5 已面向所有付费计划和 API 上线,价格与 Opus 4.8 相同;Fast mode 约为默认速度的 2.5 倍。前沿模型竞争正把能力、价格和响应速度同时推向产品前台。
Claude says Opus 5 is available on paid plans and the API at the same price as Opus 4.8, with a Fast mode around 2.5 times faster. Capability, price, and latency now compete together.
@claudeaiX 原文2.7K ❤ · 135 RT · 99 💬
04 / 15
产品实测 · Power, Friction, Deliverables, Safety04 / 15
Dan Shipper:Opus 5 很强,但并不容易驾驭A Powerful Model That Is Hard To Love
Every 团队测试 Opus 5 后给出更克制的评价:它在多类知识工作上能力突出,却可能与指令争辩或提前停下。新模型的真实价值仍需通过长期任务、遵循性与完成度来验证。
After testing Opus 5 across knowledge-work tasks, Every reports a more complicated picture: strong capability, but friction around instruction-following and task completion. Real value depends on sustained execution.
@danshipperX 原文2.2K ❤ · 128 RT · 151 💬
Opus 5 把表格与演示文稿推向顾问级产出Spreadsheets And Slides Approach Consultant-grade Output
Alex Albert 表示,短短半年多后,Opus 5 已能生成接近超人水平的表格和顾问风格演示文稿。文档与分析型工作正在从辅助草稿走向可直接交付的成品。
Alex Albert says Opus 5 can now produce near-superhuman spreadsheets and consultant-style slide decks. Document and analytical work is moving from draft assistance toward deliverable output.
@alexalbert__X 原文2.3K ❤ · 120 RT · 54 💬
Opus 5 提升网络安全能力,同时限制高风险利用Cyber Capability Rises With Safeguard Boundaries
Claude 表示 Opus 5 的网络安全能力强于 Opus 4.8,但在开发 exploit 上仍显著落后于 Mythos 5;其防护目标是在允许开发者发现和修复漏洞的同时阻断高风险用途。
Claude says Opus 5 improves on cybersecurity tasks while remaining well behind Mythos 5 at exploit development. Its safeguards aim to support defensive work while blocking high-risk use.
@claudeaiX 原文2.3K ❤ · 89 RT · 33 💬
Opus 5 在自动化行为审计中成为最对齐模型Opus 5 Leads Anthropic's Behavioral Alignment Audit
Claude 的自动化行为审计显示,Opus 5 在鲁莽或欺骗性行为上发生率更低,对 Claude Constitution 的遵循更强。能力增长与行为可靠性开始被放在同一张产品成绩单上。
Claude's automated audit reports lower reckless or deceptive behavior and stronger adherence to its constitution. Capability and behavioral reliability are increasingly evaluated together.
@claudeaiX 原文2.2K ❤ · 62 RT · 14 💬
05 / 15
深度博客 · Agent Containment05 / 15
更强的 agent 不可能只靠“更乖的模型”来保障安全;环境边界必须让最坏情况的损害仍然可控。Stronger agents cannot be secured by better model behavior alone; environment boundaries must keep worst-case damage bounded.
Anthropic:用 containment 限制 Agent 爆炸半径Containment Caps The Agent Blast Radius
Anthropic 回顾 claude.ai、Claude Code 与 Cowork 的隔离架构:仅靠逐次审批会造成 approval fatigue,遥测中用户批准约 93% 的提示;更可靠的方法是通过 sandbox、VM、文件系统边界与 egress control 限制 agent 实际能触达的环境。
Anthropic reviews containment across claude.ai, Claude Code, and Cowork. Per-action approval creates fatigue—users approved about 93% of prompts—so sandboxes, VMs, filesystem boundaries, and egress controls must cap what agents can actually reach.
06 / 15
播客深度 · Autonomous Delivery And Agentic Commerce06 / 15
PODCAST DEEP DIVE
DoorDash 的 AI 价值不是多一个聊天框,而是让发现、下单、补货与配送逐步变成由上下文驱动的闭环。DoorDash's AI value is not another chat box, but a context-driven loop across discovery, ordering, restocking, and delivery.
DoorDash:对话式商业正在改变发现与客单价Conversational Commerce Changes Discovery And Basket Size
No Priors 访谈 DoorDash 联合创始人 Andy Fang 与 Stanley Tang。Ask DoorDash 的餐厅使用路径中约 50% 导向用户从未下单过的商家,杂货场景客单量提升约 40%;更长期的方向是把拥有 30 亿次年配送规模的网络重构为 agent-first、自主交付系统。
No Priors interviews DoorDash co-founders Andy Fang and Stanley Tang. About 50% of Ask DoorDash restaurant journeys lead to a new merchant, while grocery baskets are roughly 40% larger. The longer-term bet is an agent-first autonomous delivery network operating at billions of deliveries.
07 / 15
播客理念 · Agentic Commerce In Production07 / 15
对话释放潜在需求Conversation Unlocks Latent Demand
用户可以直接描述饮食偏好、家庭计划或冰箱状态,减少传统关键词搜索与逐项点选的摩擦。
Users can describe preferences, meal plans, or fridge state directly, removing keyword-search and tap-through friction.
新商家发现率约 50%About 50% Discover A New Merchant
Ask DoorDash 的餐厅路径中,约一半用户最终向此前从未下单过的商家下单,说明 AI 正在改变消费发现。
Roughly half of Ask DoorDash restaurant journeys lead to a merchant the user has never ordered from, changing discovery behavior.
杂货客单量提升约 40%Grocery Baskets Rise About 40%
拍摄冰箱、制定餐食或一键补货,让 AI 从推荐层进入交易层,并显著扩大购物篮。
Fridge photos, meal planning, and restocking move AI from recommendation into transaction, materially expanding baskets.
下一代平台将 Agent-firstThe Next Platform Is Agent-first
当 agent 流量超过人工流量,平台要为机器客户设计身份、上下文、交易与履约接口,而不只是优化人类页面。
As agent traffic grows, platforms must design identity, context, transactions, and fulfillment for machine customers—not only human pages.
08 / 15
快讯速览 · Builder Signals08 / 15
Dan Shipper:Opus 5 很强,但并不容易驾驭A Powerful Model That Is Hard To Love
Every 团队测试 Opus 5 后给出更克制的评价:它在多类知识工作上能力突出,却可能与指令争辩或提前停下。新模型的真实价值仍需通过长期任务、遵循性与完成度来验证。
After testing Opus 5 across knowledge-work tasks, Every reports a more complicated picture: strong capability, but friction around instruction-following and task completion. Real value depends on sustained execution.
Opus 5 把表格与演示文稿推向顾问级产出Spreadsheets And Slides Approach Consultant-grade Output
Alex Albert 表示,短短半年多后,Opus 5 已能生成接近超人水平的表格和顾问风格演示文稿。文档与分析型工作正在从辅助草稿走向可直接交付的成品。
Alex Albert says Opus 5 can now produce near-superhuman spreadsheets and consultant-style slide decks. Document and analytical work is moving from draft assistance toward deliverable output.
Opus 5 提升网络安全能力,同时限制高风险利用Cyber Capability Rises With Safeguard Boundaries
Claude 表示 Opus 5 的网络安全能力强于 Opus 4.8,但在开发 exploit 上仍显著落后于 Mythos 5;其防护目标是在允许开发者发现和修复漏洞的同时阻断高风险用途。
Claude says Opus 5 improves on cybersecurity tasks while remaining well behind Mythos 5 at exploit development. Its safeguards aim to support defensive work while blocking high-risk use.
Opus 5 在自动化行为审计中成为最对齐模型Opus 5 Leads Anthropic's Behavioral Alignment Audit
Claude 的自动化行为审计显示,Opus 5 在鲁莽或欺骗性行为上发生率更低,对 Claude Constitution 的遵循更强。能力增长与行为可靠性开始被放在同一张产品成绩单上。
Claude's automated audit reports lower reckless or deceptive behavior and stronger adherence to its constitution. Capability and behavioral reliability are increasingly evaluated together.
ChatGPT Work 向全球付费用户全平台开放ChatGPT Work Goes Global Across Devices
Thibault 宣布 ChatGPT Work 已向全球付费用户开放,覆盖移动端、网页与桌面。AI 产品正在从对话窗口扩展为可跨设备承接任务与上下文的工作层。
Thibault says ChatGPT Work is now available globally to paid users across mobile, web, and desktop. AI is expanding from a chat window into a cross-device work layer.
Gemini 从学校日历 PDF 直接创建日程Gemini Turns A School Calendar PDF Into Calendar Actions
Josh Woodward 演示 Gemini 读取学校日历 PDF,并把所有停课日直接加入 Google Calendar。真正的 agent 体验不是多说几句,而是从非结构化材料提取约束并完成跨应用动作。
Josh Woodward shows Gemini reading a school-calendar PDF and adding every no-school day to Google Calendar. Agent value comes from extracting constraints and completing cross-app actions.
09 / 15
数据洞察 · Snapshot09 / 15
今日数据概览Today Stats
收录 Builder:21
总推文数:48
播客节目:1
博客文章:1

最高互动:Sam Altman:美国 AI 竞争必须同时赢得开放与专有模型 · 11.5K ❤
第二高互动:Claude Code 为新模型删掉约 80% 系统提示词 · 10.2K ❤
The snapshot includes 21 builders, 48 tweets, 1 podcast, and 1 blog post.
follow-buildersSnapshot follow-builders-2026-07-25-v2
5 条关键洞察5 Key Takeaways
开放与专有模型正在形成双轨竞争,而不是互斥路线。
Open and proprietary models are becoming parallel competitive tracks, not mutually exclusive paths.
更强模型让系统提示词变薄,但 skills 与上下文工程更重要。
Stronger models allow thinner system prompts while making skills and context engineering more important.
Opus 5 的竞争点同时覆盖能力、同价升级、2.5 倍速度与注入抗性。
Opus 5 competes on capability, same-price upgrades, 2.5x speed, and injection resistance at once.
Containment 把 agent 安全从行为劝导转成环境可达性的硬边界。
Containment turns agent safety from behavioral persuasion into hard reachability boundaries.
Work AI 与对话式商业都在把非结构化上下文直接转成跨应用动作。
Work AI and conversational commerce turn unstructured context directly into cross-app actions.
10 / 15
趋势拆解 · The New Model Product Stack10 / 15
能力层:更强通用工作Capability Layer
编码、分析、设计与文档产出正在汇入同一个高能力模型,但真实完成度仍需实测。
Coding, analysis, design, and documents converge in one capable model, but completion quality still needs real-world testing.
@bchernyX 原文5K ❤ · 383 RT · 269 💬
上下文层:更薄的系统提示词Context Layer
模型越强,越不需要堆叠指令;清晰目标、项目上下文与可复用 skill 成为关键。
Stronger models need less instruction stacking; clear goals, project context, and reusable skills matter more.
@trq212X 原文10.2K ❤ · 1.1K RT · 255 💬
交付层:同价、更快、可导出Delivery Layer
同价升级、Fast mode 与顾问级表格/演示,让模型能力变成用户可以直接感知的价值。
Same-price upgrades, Fast mode, and consultant-grade artifacts turn model progress into visible user value.
@claudeaiX 原文2.7K ❤ · 135 RT · 99 💬
治理层:对齐 + containmentGovernance Layer
注入抗性与行为对齐仍是概率防线,sandbox、VM 与 egress control 提供更硬的安全边界。
Injection resistance and alignment remain probabilistic; sandboxes, VMs, and egress controls provide harder boundaries.
@claudeaiX 原文2.2K ❤ · 62 RT · 14 💬
11 / 15
工作产品 · From Chat To Cross-app Execution11 / 15
ChatGPT Work 向全球付费用户全平台开放ChatGPT Work Goes Global Across Devices
Thibault 宣布 ChatGPT Work 已向全球付费用户开放,覆盖移动端、网页与桌面。AI 产品正在从对话窗口扩展为可跨设备承接任务与上下文的工作层。
Thibault says ChatGPT Work is now available globally to paid users across mobile, web, and desktop. AI is expanding from a chat window into a cross-device work layer.
@thsottiauxX 原文1.2K ❤ · 25 RT · 381 💬
Gemini 从学校日历 PDF 直接创建日程Gemini Turns A School Calendar PDF Into Calendar Actions
Josh Woodward 演示 Gemini 读取学校日历 PDF,并把所有停课日直接加入 Google Calendar。真正的 agent 体验不是多说几句,而是从非结构化材料提取约束并完成跨应用动作。
Josh Woodward shows Gemini reading a school-calendar PDF and adding every no-school day to Google Calendar. Agent value comes from extracting constraints and completing cross-app actions.
@joshwoodwardX 原文768 ❤ · 42 RT · 79 💬
12 / 15
安全策略 · Capability Needs Bounded Reach12 / 15
模型越能完成长任务,越不能把安全寄托在用户不断点击批准;能力增长必须与可达范围的硬限制同步。As models handle longer tasks, safety cannot depend on endless approval clicks; capability growth needs hard limits on reach.
@claudeaiX 原文2.3K ❤ · 89 RT · 33 💬
Opus 5 提升网络安全能力,同时限制高风险利用Cyber Capability Rises With Safeguard Boundaries
Claude 表示 Opus 5 的网络安全能力强于 Opus 4.8,但在开发 exploit 上仍显著落后于 Mythos 5;其防护目标是在允许开发者发现和修复漏洞的同时阻断高风险用途。
Claude says Opus 5 improves on cybersecurity tasks while remaining well behind Mythos 5 at exploit development. Its safeguards aim to support defensive work while blocking high-risk use.
@claudeaiX 原文2.3K ❤ · 89 RT · 33 💬
Opus 5 在自动化行为审计中成为最对齐模型Opus 5 Leads Anthropic's Behavioral Alignment Audit
Claude 的自动化行为审计显示,Opus 5 在鲁莽或欺骗性行为上发生率更低,对 Claude Constitution 的遵循更强。能力增长与行为可靠性开始被放在同一张产品成绩单上。
Claude's automated audit reports lower reckless or deceptive behavior and stronger adherence to its constitution. Capability and behavioral reliability are increasingly evaluated together.
@claudeaiX 原文2.2K ❤ · 62 RT · 14 💬
13 / 15
今日之声 VOICE OF THE DAY
更强模型让提示词变薄,却让上下文、skills 与 containment 变厚;真正的产品壁垒正在从“怎么问”转向“如何可靠完成”。
Stronger models make prompts thinner while context, skills, and containment grow thicker; the product moat is shifting from how to ask toward how to finish reliably.
@trq212X 原文10.2K ❤ · 1.1K RT · 255 💬
14 / 15
AI前沿每日脉动
AI Frontier Pulse · 2026.07.25
本期收录 21 位 Builder · 48 条推文 · 1 期播客 · 1 篇博客
Open Models · Opus 5 · Prompt Engineering · Agent Containment · Work AI
感谢阅读 · Thank You For Reading
Richard Liu · AI前沿每日脉动 · 2026 · Snapshot follow-builders-2026-07-25-v2