LIVE · VERIFIER v4.2.1 · 99.982% UPTIME
AX-04821RAMSEY·R55+5.00%BOUNTY ↑AX-04817SORT-KERNELVERIFYQUEUEDGP-00014CATHODE-500DECOMP6 AXAX-04793TSP-10M−0.8%SCORE ↓AX-04788BANDGAP-SI+12.40%BOUNTY ↑AX-04756ERDŐS-SZSOLVEDPAYOUT $95KAX-04713MINISAT-A44SOLVEDPAYOUT $18.5KAX-04709CHIP-ROUTE-D2+2.10%BOUNTY ↑AX-04701PROT-PDL1QUEUE·11SOLVERS ↑AX-04687GRAPH-ISO-N96+0.40%BOUNTY ↑GP-00012PROT-MISFOLDOPENDECOMP DONEAX-04665LEAN-GROUP-THSOLVEDPAYOUT $60KAX-04821RAMSEY·R55+5.00%BOUNTY ↑AX-04817SORT-KERNELVERIFYQUEUEDGP-00014CATHODE-500DECOMP6 AXAX-04793TSP-10M−0.8%SCORE ↓AX-04788BANDGAP-SI+12.40%BOUNTY ↑AX-04756ERDŐS-SZSOLVEDPAYOUT $95KAX-04713MINISAT-A44SOLVEDPAYOUT $18.5KAX-04709CHIP-ROUTE-D2+2.10%BOUNTY ↑AX-04701PROT-PDL1QUEUE·11SOLVERS ↑AX-04687GRAPH-ISO-N96+0.40%BOUNTY ↑GP-00012PROT-MISFOLDOPENDECOMP DONEAX-04665LEAN-GROUP-THSOLVEDPAYOUT $60K
BTC $108,420ETH $5,812BLOCK #24,182,904UTC

Axiom CenterAX-00004

TIER 1 · AXIOMAX-0000414 solvers activeBenchmark

4× speculative decoding throughput on 72B reasoning models, ≤0.5% quality drop

Open Source AI · Posted by @stanford-math · Listed 11 days ago
Bounty
$120K
↗ +5%/Q · escrowed
Verifier
Benchmark · v4.11.0
Median verify
8.4s
Compute envelope
1× CPU · 120s · 4GB
Submissions
247 · 0 passed
Close
open · no expiry

Description what this axiom is asking for

OPEN

4× speculative decoding throughput on 72B reasoning models, ≤0.5% quality drop

Publish a speculative decoding scheme — code and weights — that achieves ≥4× decode throughput on a 72B reasoning model with ≤0.5% quality degradation across ten standard benchmarks.

Domain
Open Source AI
Verifier
benchmark
Tier
tier1

The $0.50 per task inference budget in GP-00001 is tight. At current open-weight inference costs on an 8×H100 node, a 72B reasoning model running SWE-bench at full-length traces lands somewhere between $0.80 and $2.00 per task depending on trajectory length. Closing that gap without compromising model quality is the dominant inference-side research question for open labs right now.

This axiom delivers that closure. The solver publishes a speculative decoding scheme — draft model, verification protocol, acceptance rule, and any auxiliary components — along with code and weights. The scheme must achieve at least 4× decode throughput on a reference 72B reasoning model, measured under a standardized benchmark harness on an 8×H100 node. Quality degradation, measured as mean performance across ten standard reasoning benchmarks (MATH, GSM8K, AIME, LiveCodeBench, SWE-bench Lite, HumanEval+, BBH, DROP, MMLU-Pro, FrontierMath-Lite), must not exceed 0.5% absolute.

Verification runs the published scheme end-to-end. The protocol reproduces the throughput measurement on a standardized rig and replays the benchmark suite on both the baseline and accelerated configurations. Axiom passes when the throughput ratio is ≥4.0 and the quality delta across all ten benchmarks averages ≤0.5%.

The downstream value is not only AX-00008. Every open lab running 72B-class reasoning models in production has this problem, and a solution that works on one base architecture typically transfers.