Claude Fable 5.1 vs Mythos 5.1: How to Benchmark the Safeguard Tax

Fable 5.1 and Mythos 5.1 share model weights but use different safeguards. Measure capability, intervention, cost, and access together.

Administrator Administrator Published on 2026-09-02

LLM 表现出编码痕迹却仍答不出

WikiProfile 把被归为未编码的事实,与表现出行为编码证据却难以提取的事实分开,形成更准确的事实错误诊断。

Administrator Administrator Published on 2026-09-01

LLM Recall: Encoded Facts That Remain Hard to Retrieve

WikiProfile separates facts classified as not encoded from facts with behavioral encoding evidence that remain hard to recall.

Administrator Administrator Published on 2026-09-01

MCP Agent 身份:哪些已稳定,哪些仍是草案

MCP 已把 Agent 身份列为协议重点。本文区分现行授权、企业扩展、工作负载身份、委托与 DPoP 的真实成熟度。

Administrator Administrator Published on 2026-08-30

MCP Agent Identity in 2026: What Is Stable, What Is Draft, and What Each Mechanism Proves

MCP agent identity is becoming a protocol priority. This guide separates stable authorization from draft workload identity, delegation, and DPoP work.

Administrator Administrator Published on 2026-08-30

自动对齐研究员:Claude 安全增益背后的评测合同

Claude 能自动改进十类对齐失败。真正值得复用的成果,是一套能发现过拟合、能力退化、作弊和外推失效的评测合同。

Administrator Administrator Published on 2026-08-29

Automated Alignment Researchers Need an Evaluation Contract

Automated alignment researchers need more than strong scores. Anthropic's experiment shows the evaluation contract that makes gains testable.

Administrator Administrator Published on 2026-08-29

加州 AB 1651 AI 披露:高风险认证需要两套账

加州 AB 1651 提醒我们:日常生产需要结果核验账,高风险认证还需要独立的来源披露账。

Administrator Administrator Published on 2026-08-27

California AB 1651 AI Disclosure: Two Ledgers for High-Stakes Certification

California AB 1651 shows why AI governance needs one ledger for verified work and another for disclosure at certification gates.

Administrator Administrator Published on 2026-08-27

OpenAI Jalapeño 基准审计:全栈推理闭环

审计 OpenAI Jalapeño 首批基准:1.5 至 1.9 倍每瓦吞吐说明了什么,遗漏了什么,以及全栈推理为何才是真正的战略资产。

Administrator Administrator Published on 2026-08-26
Previous Next