The central measurement problem is no longer whether a model can pass a benchmark. It is whether a system can produce consequential outcomes under realistic conditions.

Premise

AI capability and reliable control are moving at different speeds. Public benchmarks often saturate, laboratory tasks can understate elicited performance, and safety claims are frequently tested under conditions that exclude the adversary. Our agenda focuses on the gap between apparent performance and operational capability.

A useful evaluation changes a real decision: release, access, monitoring, mitigation, or further research.

Questions we prioritize

  • Which scientific and technical tasks become materially easier with frontier-model assistance?
  • How do autonomy, tool access, memory, and iteration change the risk profile?
  • When do safeguards remain robust under determined, adaptive pressure?
  • Which model behaviors indicate emerging control problems before they become operational failures?
  • How can evaluations remain informative without publishing an instruction manual for misuse?

2026 programs

01 — Long-horizon capability

Measure planning, adaptation, and recovery across multi-step tasks where success depends on maintaining state and using tools—not merely producing a plausible answer.

02 — Scientific uplift

Estimate the change in a qualified user's ability to solve domain problems, with particular care around biological dual use and the difference between information access and actionable assistance.

03 — Control under pressure

Evaluate oversight and safeguard mechanisms against adaptive strategies, hidden information, and adversarial elicitation. The object of measurement is the full system, not the base model in isolation.

04 — Evaluation integrity

Study contamination, evaluator effects, scoring uncertainty, and the gap between proxy metrics and outcomes. Strong claims require methods that survive contact with replication.

What we publish

We intend to publish research briefs, reproducible low-risk evaluation components, methods notes, and decision-focused synthesis. Sensitive details will follow our responsible research standard and may be delayed, redacted, or shared through controlled access.