Capability evaluations
Task-grounded tests for autonomy, scientific reasoning, tool use, and long-horizon execution.
- Agentic behavior
- Scientific uplift
- Evaluation integrity
GPT-Rosalind CERTIFIED
Independent AI research
Boundary Signal evaluates frontier AI systems across capability, alignment, biosecurity, and adversarial robustness—then turns evidence into decisions.
The gap we work in
Frontier models are becoming capable in domains where failure is consequential and conventional benchmarks are thin. We build evaluations that expose the boundary: what a system can accomplish, where safeguards break, and which interventions hold.
Four connected programs
We pair empirical measurement with operational threat modeling so results are useful to model builders, auditors, and public-interest institutions.
Task-grounded tests for autonomy, scientific reasoning, tool use, and long-horizon execution.
Behavioral and mechanistic studies of control, oversight, deception, and goal-directed behavior.
Measurements of biological problem-solving uplift, safeguard reliability, and misuse pathways.
Structured red teaming that maps failure modes without turning findings into a misuse manual.
From concern to evidence
Every project begins with a decision someone needs to make—not a benchmark looking for a purpose. Our methods are designed around consequence, reproducibility, and clear uncertainty.
Define the actor, access, objective, constraints, and real-world consequence.
Build elicitation-aware tasks with controls, baselines, and independent review.
Stress the protocol and the safeguard—not only the model under ideal conditions.
Report findings, uncertainty, mitigations, and the evidence that would change the conclusion.
Open work
Useful, not enabling
Dual-use research creates an obligation beyond ordinary publication. We separate public evidence from operational detail, coordinate disclosure, and use controlled access when release could materially lower the barrier to harm.
Read our standard ↗Boundary Signal Research
Boundary Signal is an independent AI research company focused on the empirical gap between model capability and reliable control. We work across disciplines because consequential AI risk does not respect disciplinary borders.
We welcome research collaborations, evaluation partnerships, and careful criticism from people working on the same hard problems.
Start with the question