EvoUndo shows why capability gains are insufficient for self-modifying Agent Harnesses. Require witnessed, counterfactual recovery before merge.
Fable 5.1 and Mythos 5.1 share model weights but use different safeguards. Measure capability, intervention, cost, and access together.
Anthropic's conflict tests show why multi-agent coordination needs mandate invariants, escalation rules, and independent acceptance checks.
OpenAI is treating Astra as Critical for cyber safety. Here is what the threshold proves, what remains preliminary, and which controls must follow.
ChatGPT Health can read connected medical records across conversations. Evaluate its permissions, memory, deletion, HIPAA, and verification boundaries
Anthropic found a small, causally active J-space inside Claude. It carries silent concepts but does not prove subjective consciousness.
Anthropic's Natural Language Autoencoders convert opaque LLM activations into human-readable text. This deep dive covers the architecture, safety appl
Anthropic's latest interpretability research maps 171 emotion concepts inside Claude using sparse autoencoders. The findings reveal that emotion vecto
"OpenAI discovered that GPT-5 developed a 3,881% surge in 'goblin' references. The root cause traces to a personality feature and a reward signal that
"Anthropic assembled 12 tech giants and built a cybersecurity AI model too dangerous to release publicly. Project Glasswing found thousands of zero-da