Claude 能自动改进十类对齐失败。真正值得复用的成果,是一套能发现过拟合、能力退化、作弊和外推失效的评测合同。
Automated alignment researchers need more than strong scores. Anthropic's experiment shows the evaluation contract that makes gains testable.