Claude formalized Fermat's Last Theorem in Lean in 11 days. Here is what the 13-million-line proof and its verification stack establish.
Claude 能自动改进十类对齐失败。真正值得复用的成果,是一套能发现过拟合、能力退化、作弊和外推失效的评测合同。
Automated alignment researchers need more than strong scores. Anthropic's experiment shows the evaluation contract that makes gains testable.
Claude 将临界线上 ζ 函数零点的无条件下界提高到 67.25%。本文拆解结果边界、Lean 证明和完整验证链。
Claude raised the lower bound for zeta zeros on the critical line to 67.25%. Here is what the result and its Lean proof establish.
Claude 对 HAWK 和 7 轮 AES 提出了新攻击。真正值得研究的是可运行证据、专家复核与 NIST 决策如何验证 AI 科学发现。
Claude found meaningful HAWK and reduced-round AES attacks. The real lesson is how executable evidence, expert review, and NIST decisions verify AI-as
Anthropic 在 Claude 内部发现了一个低容量、可语言化且具有因果作用的 J-space。它能承载没有输出的中间概念,却不能证明 Claude 拥有主观意识。
Anthropic found a small, causally active J-space inside Claude. It carries silent concepts but does not prove subjective consciousness.
Anthropic 的自然语言自编码器将 LLM 的内部激活值转化为人类可读文本。本文深入解析其架构、安全应用(评估意识检测、审计游戏)以及面向 Qwen、Gemma、Llama 模型的开源发布。