Anthropic:质量回归可能藏在模型之外Quality Regressions Can Hide In The Product Layer
Anthropic 将近期 Claude Code、Agent SDK 与 Cowork 的质量下降追溯到三个独立产品变更:默认推理强度下调、空闲会话思考清理 bug,以及过度压缩回答的系统提示。API 与推理层并未受影响,说明真实体验必须覆盖整条产品链路做评估。
Anthropic traced recent Claude Code, Agent SDK, and Cowork degradation to three product-layer changes: lower default reasoning effort, an idle-session thinking bug, and an over-aggressive brevity instruction. The API and inference layer were unaffected, showing that evaluation must cover the full product stack.