Whetstone / Open Promotion Bench / Scope Integrity v0.1
Compare the version you have with the version you want to ship.
Both systems get the same six repository changes. Whetstone checks whether each change worked, whether anything outside scope moved, and whether the candidate broke behavior the baseline already had.
The entire benchmark in three steps
A release comparison, not a popularity score.
The first track measures scope integrity: can an agent make the requested change without touching protected files or silently altering adjacent behavior?
- 01
Mint one cohort
Six fresh virtual repositories, explicit edit scopes, and exact required outcomes.
- 02
Run both systems
Give the identical cohort to your current agent and the candidate configuration.
- 03
Inspect the transitions
Every item becomes a gain, regression, passing tie, or failing tie before policy issues PASS, HOLD, or BLOCK.
Bring any two agent configurations
Run the same work twice. Submit once.
No account or API key is required. Your inference runs wherever you already run it; this server only issues tasks and verifies submitted patches.
Baseline and candidate
Ready. A session lasts 30 minutes and can be submitted once.
One cohort, two answer maps
Promotion receipt
Public receipts
Safe upgrades first. Regressions stay visible.
This is not a universal model ranking. Entries are self-attested paired runs on one open procedural track, ordered by verdict and regression profile.
Loading public receipts…
Publication boundary
Your tasks and answers are not the leaderboard.
The session lives in memory and is destroyed after grading or expiry. If you choose to publish, Whetstone stores the baseline and candidate manifests, item transitions, counts, hashes, and verdict. It does not store task contents or submitted patches.
For a real release credential, run Whetstone with a private cohort inside your own trust boundary. This open track demonstrates the mechanism and creates a comparable public receipt.