Automated alignment researchers need more than strong scores. Anthropic's experiment shows the evaluation contract that makes gains testable.
加州 AB 1651 提醒我们:日常生产需要结果核验账,高风险认证还需要独立的来源披露账。
California AB 1651 shows why AI governance needs one ledger for verified work and another for disclosure at certification gates.
OpenAI Admin Plugin 把工作区管理变成权限感知的 Agent 闭环。真正的架构由身份、动作控制、审批、结果证据与审计组成。
OpenAI's Admin Plugin turns workspace administration into a permission-aware agent loop. The real design is identity, action control, approval, eviden
一套可复用的 Qwen3.8-27B 实战评测合同:比较 tok/s 之前,先控制量化、引擎、缓存、上下文、推理强度、质量与并发。
A practical Qwen3.8-27B benchmark contract that controls quantization, engine, cache, context, reasoning effort, quality, and concurrency before compa
审计 OpenAI Jalapeño 首批基准:1.5 至 1.9 倍每瓦吞吐说明了什么,遗漏了什么,以及全栈推理为何才是真正的战略资产。
An audit of OpenAI Jalapeño's first benchmarks: what the 1.5–1.9x performance-per-watt lead proves, what it omits, and why full-stack inference is the
MTIA 300 把网络放进加速器封装,MetaRoCE 把传输智能放进端点 NIC。本文拆解两条设计线的共同目标、机制差异与证据边界。