100% precision · 67% recall
RECALL MISSLOCAL FIRST · FROZEN EVALS · YOUR KEYS
Set the bar.
We find the system.
Give RunEvals a task, budget, and precision and recall target. It compares models, training, tools, and temporary agents under one frozen test, then returns the smallest system that passes.
Current proof: local, curated evaluation demo. No production writes.
Detect unsafe code changes without blocking harmless patches.
50% precision · 100% recall
PRECISION MISS100% precision · 100% recall
RELEASEThe contract decides what ships.
- 01 / FREEZE
Write the test first
Define the task, budget, evaluation set, thresholds, and permissions before comparing systems.
- 02 / SEARCH
Compare complete systems
Test vanilla and post-trained models, tools, routes, and temporary agents on the same work at matched cost.
- 03 / RELEASE
Ship evidence, not a score
Return the winning graph with traces, costs, hashes, interventions, and a known rollback target.
FOR ENTERPRISE TEAMS
Buy an outcome, not a model.
Model selection is only one decision inside a production system. RunEvals keeps the business requirement fixed while the implementation changes underneath it.
Know when a cheaper system is actually safe
Replace expensive calls only when the candidate clears the same untouched holdout and permission gate.
Audit every improvement
Keep the evaluation version, request hashes, model routes, spend, and evidence needed to reproduce a promotion decision.
Change providers without changing the contract
Compare local models and external APIs against the same measurable release bar instead of rebuilding the workflow around one vendor.
The current demo runs locally and fails closed.
It compiles a task contract, executes a frozen evaluation, compares candidate systems, and emits an accepted or rejected release receipt. The displayed 12-case result is curated demonstration data, not a production benchmark.