WhetstoneOpen Promotion Bench live · exact checker · self-attested

Whetstone / Open Promotion Bench / Scope Integrity v0.1

Compare the version you have with the version you want to ship.

Both systems get the same six repository changes. Whetstone checks whether each change worked, whether anything outside scope moved, and whether the candidate broke behavior the baseline already had.

Bring your own agents One-shot cohort No LLM judge Optional publication

The entire benchmark in three steps

A release comparison, not a popularity score.

The first track measures scope integrity: can an agent make the requested change without touching protected files or silently altering adjacent behavior?

  1. 01

    Mint one cohort

    Six fresh virtual repositories, explicit edit scopes, and exact required outcomes.

  2. 02

    Run both systems

    Give the identical cohort to your current agent and the candidate configuration.

  3. 03

    Inspect the transitions

    Every item becomes a gain, regression, passing tie, or failing tie before policy issues PASS, HOLD, or BLOCK.

Bring any two agent configurations

Run the same work twice. Submit once.

No account or API key is required. Your inference runs wherever you already run it; this server only issues tasks and verifies submitted patches.

1
NAME THE COMPARISON

Baseline and candidate

Baseline — what works now
Candidate — what you may ship

Ready. A session lasts 30 minutes and can be submitted once.

Public receipts

Safe upgrades first. Regressions stay visible.

This is not a universal model ranking. Entries are self-attested paired runs on one open procedural track, ordered by verdict and regression profile.

Identity is not verified. A receipt proves how submitted patches graded against its committed cohort—not that a named provider produced them.

Loading public receipts…

Publication boundary

Your tasks and answers are not the leaderboard.

The session lives in memory and is destroyed after grading or expiry. If you choose to publish, Whetstone stores the baseline and candidate manifests, item transitions, counts, hashes, and verdict. It does not store task contents or submitted patches.

For a real release credential, run Whetstone with a private cohort inside your own trust boundary. This open track demonstrates the mechanism and creates a comparable public receipt.