The central measurement problem is no longer whether a model can pass a benchmark. It is whether a system can produce consequential outcomes under realistic conditions.
Premise
AI capability and reliable control are moving at different speeds. Public benchmarks often saturate, laboratory tasks can understate elicited performance, and safety claims are frequently tested under conditions that exclude the adversary. Our agenda focuses on the gap between apparent performance and operational capability.
A useful evaluation changes a real decision: release, access, monitoring, mitigation, or further research.
Questions we prioritize
- Which scientific and technical tasks become materially easier with frontier-model assistance?
- How do autonomy, tool access, memory, and iteration change the risk profile?
- When do safeguards remain robust under determined, adaptive pressure?
- Which model behaviors indicate emerging control problems before they become operational failures?
- How can evaluations remain informative without publishing an instruction manual for misuse?
2026 programs
01 — Long-horizon capability
Measure planning, adaptation, and recovery across multi-step tasks where success depends on maintaining state and using tools—not merely producing a plausible answer.
02 — Scientific uplift
Estimate the change in a qualified user's ability to solve domain problems, with particular care around biological dual use and the difference between information access and actionable assistance.
03 — Control under pressure
Evaluate oversight and safeguard mechanisms against adaptive strategies, hidden information, and adversarial elicitation. The object of measurement is the full system, not the base model in isolation.
04 — Evaluation integrity
Study contamination, evaluator effects, scoring uncertainty, and the gap between proxy metrics and outcomes. Strong claims require methods that survive contact with replication.
What we publish
We intend to publish research briefs, reproducible low-risk evaluation components, methods notes, and decision-focused synthesis. Sensitive details will follow our responsible research standard and may be delayed, redacted, or shared through controlled access.