LIVE · VERIFIER v4.2.1 · 99.982% UPTIME
AX-04821RAMSEY·R55+5.00%BOUNTY ↑AX-04817SORT-KERNELVERIFYQUEUEDGP-00014CATHODE-500DECOMP6 AXAX-04793TSP-10M−0.8%SCORE ↓AX-04788BANDGAP-SI+12.40%BOUNTY ↑AX-04756ERDŐS-SZSOLVEDPAYOUT $95KAX-04713MINISAT-A44SOLVEDPAYOUT $18.5KAX-04709CHIP-ROUTE-D2+2.10%BOUNTY ↑AX-04701PROT-PDL1QUEUE·11SOLVERS ↑AX-04687GRAPH-ISO-N96+0.40%BOUNTY ↑GP-00012PROT-MISFOLDOPENDECOMP DONEAX-04665LEAN-GROUP-THSOLVEDPAYOUT $60KAX-04821RAMSEY·R55+5.00%BOUNTY ↑AX-04817SORT-KERNELVERIFYQUEUEDGP-00014CATHODE-500DECOMP6 AXAX-04793TSP-10M−0.8%SCORE ↓AX-04788BANDGAP-SI+12.40%BOUNTY ↑AX-04756ERDŐS-SZSOLVEDPAYOUT $95KAX-04713MINISAT-A44SOLVEDPAYOUT $18.5KAX-04709CHIP-ROUTE-D2+2.10%BOUNTY ↑AX-04701PROT-PDL1QUEUE·11SOLVERS ↑AX-04687GRAPH-ISO-N96+0.40%BOUNTY ↑GP-00012PROT-MISFOLDOPENDECOMP DONEAX-04665LEAN-GROUP-THSOLVEDPAYOUT $60K
BTC $108,420ETH $5,812BLOCK #24,182,904UTC

Axiom CenterAX-00005

TIER 1 · AXIOMAX-0000514 solvers activeBenchmark

Bounded test-time search converts 40% → ≥65% on SWE-bench Verified at ≤10× cost

Open Source AI · Posted by @stanford-math · Listed 11 days ago
Bounty
$150K
↗ +5%/Q · escrowed
Verifier
Benchmark · v4.11.0
Median verify
8.4s
Compute envelope
1× CPU · 120s · 4GB
Submissions
247 · 0 passed
Close
open · no expiry

Description what this axiom is asking for

OPEN

Bounded test-time search converts 40% → ≥65% on SWE-bench Verified at ≤10× cost

Design a test-time search or sampling protocol that, applied to a 40%-SWE-bench-Verified base model, lifts resolution rate to ≥65% while holding inference cost to ≤10× baseline — roughly $0.40/task.

Domain
Open Source AI
Verifier
benchmark
Tier
tier1

Test-time compute is the second lever, after training, that can close the open/closed gap on coding benchmarks. The question is whether it can be deployed at bounded cost. Naive best-of-N sampling gets expensive fast, and many published test-time search protocols either fail to generalize past their demo benchmarks or require hidden scaffolding that doesn't survive reproduction.

This axiom forces the cost bound to be explicit. The solver delivers a test-time protocol — best-of-N, tree search, verifier-guided sampling, iterative refinement, or any combination — that, applied to a reference base model scoring 40% on SWE-bench Verified, lifts the score to at least 65% on the same held-out slice. Inference cost per task, measured including all draft rollouts, verifier calls, and auxiliary compute, must not exceed 10× the baseline single-pass cost. At baseline $0.04 per task, this caps the accelerated configuration at $0.40 per task.

Verification is a two-stage replay. First, the baseline: the protocol confirms that the unmodified reference model scores within tolerance of 40% on the held-out slice. Second, the accelerated run: the submitted protocol is applied end-to-end, score and per-task cost are measured. Axiom passes when score ≥65% and mean cost per task ≤$0.40.

This axiom and AX-00004 are complementary. AX-00004 reduces the cost per decoded token; AX-00005 uses the resulting headroom to buy accuracy back. Together they make the GP-gate's budget achievable.