EvoUndo: A Recoverability Gate for Self-Modifying AI Agent Harnesses

EvoUndo shows why capability gains are insufficient for self-modifying Agent Harnesses. Require witnessed, counterfactual recovery before merge.

Administrator Administrator Published on 2026-09-02

Claude Fable 5.1 与 Mythos 5.1:怎样评测安全护栏成本

Fable 5.1 与 Mythos 5.1 使用相同权重、不同护栏。模型选型应联合评测能力、干预、可用结果成本与访问资格。

Administrator Administrator Published on 2026-09-02

Claude Fable 5.1 vs Mythos 5.1: How to Benchmark the Safeguard Tax

Fable 5.1 and Mythos 5.1 share model weights but use different safeguards. Measure capability, intervention, cost, and access together.

Administrator Administrator Published on 2026-09-02

ChatGPT Ads Need an Answer-Independence Audit

A sponsored label identifies the ad but cannot prove that advertising leaves ChatGPT answers unchanged. Here is a repeatable audit contract.

Administrator Administrator Published on 2026-09-01

LLM 表现出编码痕迹却仍答不出

WikiProfile 把被归为未编码的事实,与表现出行为编码证据却难以提取的事实分开,形成更准确的事实错误诊断。

Administrator Administrator Published on 2026-09-01

LLM Recall: Encoded Facts That Remain Hard to Retrieve

WikiProfile separates facts classified as not encoded from facts with behavioral encoding evidence that remain hard to recall.

Administrator Administrator Published on 2026-09-01

MCP Agent 身份:哪些已稳定,哪些仍是草案

MCP 已把 Agent 身份列为协议重点。本文区分现行授权、企业扩展、工作负载身份、委托与 DPoP 的真实成熟度。

Administrator Administrator Published on 2026-08-30

Thomson Reuters Thomson 1.0: What Its $40M Domain Model Actually Bought

Thomson Reuters spent about $40M on Thomson 1.0. Its reusable data, experts, evaluation, training, and deployment system matters most.

Administrator Administrator Published on 2026-08-30

MCP Agent Identity in 2026: What Is Stable, What Is Draft, and What Each Mechanism Proves

MCP agent identity is becoming a protocol priority. This guide separates stable authorization from draft workload identity, delegation, and DPoP work.

Administrator Administrator Published on 2026-08-30

自动对齐研究员:Claude 安全增益背后的评测合同

Claude 能自动改进十类对齐失败。真正值得复用的成果,是一套能发现过拟合、能力退化、作弊和外推失效的评测合同。

Administrator Administrator Published on 2026-08-29
Previous Next