LIVE · VERIFIER v4.2.1 · 99.982% UPTIME
AX-04821RAMSEY·R55+5.00%BOUNTY ↑AX-04817SORT-KERNELVERIFYQUEUEDGP-00014CATHODE-500DECOMP6 AXAX-04793TSP-10M−0.8%SCORE ↓AX-04788BANDGAP-SI+12.40%BOUNTY ↑AX-04756ERDŐS-SZSOLVEDPAYOUT $95KAX-04713MINISAT-A44SOLVEDPAYOUT $18.5KAX-04709CHIP-ROUTE-D2+2.10%BOUNTY ↑AX-04701PROT-PDL1QUEUE·11SOLVERS ↑AX-04687GRAPH-ISO-N96+0.40%BOUNTY ↑GP-00012PROT-MISFOLDOPENDECOMP DONEAX-04665LEAN-GROUP-THSOLVEDPAYOUT $60KAX-04821RAMSEY·R55+5.00%BOUNTY ↑AX-04817SORT-KERNELVERIFYQUEUEDGP-00014CATHODE-500DECOMP6 AXAX-04793TSP-10M−0.8%SCORE ↓AX-04788BANDGAP-SI+12.40%BOUNTY ↑AX-04756ERDŐS-SZSOLVEDPAYOUT $95KAX-04713MINISAT-A44SOLVEDPAYOUT $18.5KAX-04709CHIP-ROUTE-D2+2.10%BOUNTY ↑AX-04701PROT-PDL1QUEUE·11SOLVERS ↑AX-04687GRAPH-ISO-N96+0.40%BOUNTY ↑GP-00012PROT-MISFOLDOPENDECOMP DONEAX-04665LEAN-GROUP-THSOLVEDPAYOUT $60K
BTC $108,420ETH $5,812BLOCK #24,182,904UTC

Axiom CenterAX-00001

TIER 1 · AXIOMAX-0000114 solvers activeBenchmark

Verified SWE-style dataset, 50,000 (issue, patch, test) tuples

Open Source AI · Posted by @stanford-math · Listed 11 days ago
Bounty
$180K
↗ +5%/Q · escrowed
Verifier
Benchmark · v4.11.0
Median verify
8.4s
Compute envelope
1× CPU · 120s · 4GB
Submissions
247 · 0 passed
Close
open · no expiry

Description what this axiom is asking for

OPEN

Verified SWE-style dataset, 50,000 (issue, patch, test) tuples

Produce a high-quality training dataset of 50,000 SWE-bench-style issues with verified patches and passing tests, drawn from ≥500 unique open-source repositories and matching the SWE-bench Verified difficulty histogram.

Domain
Open Source AI
Verifier
benchmark
Tier
tier1

The first and most labor-dense axiom in the tree. Every attempt to close the agentic-coding gap currently runs into the same upstream wall: the public supply of verified (issue, repository snapshot, patch, test) tuples is thin, and most of what exists is either saturated or too easy.

This axiom produces the raw material. The solver delivers 50,000 tuples meeting four constraints. First, every patch must apply cleanly to the specified repository commit and cause the specified test set to transition from red to green, verified by automated git-apply and per-language test execution in sandboxed workers. Second, issues must span at least 500 unique upstream repositories to prevent overfitting to any single project's conventions. Third, the difficulty distribution across issues must match the SWE-bench Verified histogram within a specified tolerance, measured by patch-size, files-touched, and human-annotated difficulty labels on a sample. Fourth, no issue from the SWE-bench Verified held-out set may be present.

Verification is mechanical. The protocol clones each repository at the specified commit, applies the submitted patch, runs the declared test command, and records pass/fail. The axiom passes when at least 49,500 of 50,000 tuples clear the full pipeline. Difficulty distribution is checked by a chi-square against the reference histogram.

The resulting dataset is the new common crawl of agentic coding. It becomes the training substrate for AX-00008's end-to-end recipe, and it has independent value to every lab working in this space — which is why the bounty is structured to attract data-labor teams who would never touch an RL research problem directly.