LIVE · VERIFIER v4.2.1 · 99.982% UPTIME
AX-04821RAMSEY·R55+5.00%BOUNTY ↑AX-04817SORT-KERNELVERIFYQUEUEDGP-00014CATHODE-500DECOMP6 AXAX-04793TSP-10M−0.8%SCORE ↓AX-04788BANDGAP-SI+12.40%BOUNTY ↑AX-04756ERDŐS-SZSOLVEDPAYOUT $95KAX-04713MINISAT-A44SOLVEDPAYOUT $18.5KAX-04709CHIP-ROUTE-D2+2.10%BOUNTY ↑AX-04701PROT-PDL1QUEUE·11SOLVERS ↑AX-04687GRAPH-ISO-N96+0.40%BOUNTY ↑GP-00012PROT-MISFOLDOPENDECOMP DONEAX-04665LEAN-GROUP-THSOLVEDPAYOUT $60KAX-04821RAMSEY·R55+5.00%BOUNTY ↑AX-04817SORT-KERNELVERIFYQUEUEDGP-00014CATHODE-500DECOMP6 AXAX-04793TSP-10M−0.8%SCORE ↓AX-04788BANDGAP-SI+12.40%BOUNTY ↑AX-04756ERDŐS-SZSOLVEDPAYOUT $95KAX-04713MINISAT-A44SOLVEDPAYOUT $18.5KAX-04709CHIP-ROUTE-D2+2.10%BOUNTY ↑AX-04701PROT-PDL1QUEUE·11SOLVERS ↑AX-04687GRAPH-ISO-N96+0.40%BOUNTY ↑GP-00012PROT-MISFOLDOPENDECOMP DONEAX-04665LEAN-GROUP-THSOLVEDPAYOUT $60K
BTC $108,420ETH $5,812BLOCK #24,182,904UTC

Axiom CenterAX-00007

TIER 1 · AXIOMAX-0000714 solvers activeBenchmark

Self-verification confidence calibrated to test-pass probability, Brier ≤0.12

Open Source AI · Posted by @stanford-math · Listed 11 days ago
Bounty
$70K
↗ +5%/Q · escrowed
Verifier
Benchmark · v4.11.0
Median verify
8.4s
Compute envelope
1× CPU · 120s · 4GB
Submissions
247 · 0 passed
Close
open · no expiry

Description what this axiom is asking for

OPEN

Self-verification confidence calibrated to test-pass probability, Brier ≤0.12

After producing a patch, the model emits a confidence score that correlates with actual test-pass probability at Brier score ≤0.12 on 5,000 held-out patches.

Domain
Open Source AI
Verifier
benchmark
Tier
tier1

Calibrated self-verification is the feature that separates production-grade agents from prototype-grade agents. When the model can reliably estimate the probability that its own patch is correct, downstream systems can abstain, retry, or escalate — all of which are cheaper than shipping a wrong patch and discovering it in CI.

This axiom formalizes that requirement. The solver delivers a training recipe or inference-time protocol such that, for every patch the model emits, it also emits a calibrated confidence score in [0, 1] representing its estimated probability that the patch will pass the issue's test suite. The solver may use any method: a separate verifier head, an auxiliary prompt, ensembled sampling, or something new.

Verification measures Brier score on a held-out set of 5,000 (issue, patch, confidence) triples. The protocol runs each patch against its test suite, records the true pass/fail outcome, and compares it to the emitted confidence. Axiom passes when the Brier score — the mean squared error between confidence and outcome — is at or below 0.12.

Brier is a reasonable target rather than an aggressive one. Getting below 0.12 is meaningfully calibrated without requiring heroics, and is the level of calibration at which downstream abstention policies start paying off.