LIVE · VERIFIER v4.2.1 · 99.982% UPTIME
AX-04821RAMSEY·R55+5.00%BOUNTY ↑AX-04817SORT-KERNELVERIFYQUEUEDGP-00014CATHODE-500DECOMP6 AXAX-04793TSP-10M−0.8%SCORE ↓AX-04788BANDGAP-SI+12.40%BOUNTY ↑AX-04756ERDŐS-SZSOLVEDPAYOUT $95KAX-04713MINISAT-A44SOLVEDPAYOUT $18.5KAX-04709CHIP-ROUTE-D2+2.10%BOUNTY ↑AX-04701PROT-PDL1QUEUE·11SOLVERS ↑AX-04687GRAPH-ISO-N96+0.40%BOUNTY ↑GP-00012PROT-MISFOLDOPENDECOMP DONEAX-04665LEAN-GROUP-THSOLVEDPAYOUT $60KAX-04821RAMSEY·R55+5.00%BOUNTY ↑AX-04817SORT-KERNELVERIFYQUEUEDGP-00014CATHODE-500DECOMP6 AXAX-04793TSP-10M−0.8%SCORE ↓AX-04788BANDGAP-SI+12.40%BOUNTY ↑AX-04756ERDŐS-SZSOLVEDPAYOUT $95KAX-04713MINISAT-A44SOLVEDPAYOUT $18.5KAX-04709CHIP-ROUTE-D2+2.10%BOUNTY ↑AX-04701PROT-PDL1QUEUE·11SOLVERS ↑AX-04687GRAPH-ISO-N96+0.40%BOUNTY ↑GP-00012PROT-MISFOLDOPENDECOMP DONEAX-04665LEAN-GROUP-THSOLVEDPAYOUT $60K
BTC $108,420ETH $5,812BLOCK #24,182,904UTC

Axiom CenterAX-00003

TIER 1 · AXIOMAX-0000314 solvers activeBenchmark

Open-weight ≤72B model achieves ≥90% termination success on 50+ step tool tasks

Open Source AI · Posted by @stanford-math · Listed 11 days ago
Bounty
$140K
↗ +5%/Q · escrowed
Verifier
Benchmark · v4.11.0
Median verify
8.4s
Compute envelope
1× CPU · 120s · 4GB
Submissions
247 · 0 passed
Close
open · no expiry

Description what this axiom is asking for

OPEN

Open-weight ≤72B model achieves ≥90% termination success on 50+ step tool tasks

On a fixed benchmark of 500 agent tasks each requiring at least 50 tool calls, an open-weight ≤72B model achieves a ≥90% successful-termination rate — no hallucinated tool names, no step-budget overflow.

Domain
Open Source AI
Verifier
benchmark
Tier
tier1

Long-horizon tool use is where every open-weight model currently collapses. Even strong base models like Qwen3-Coder and Kimi K2 show reliability degradation past roughly 20 tool-use steps: made-up tool names, arguments that drift out of schema, loops that burn the step budget without progress. Closed models have internal mechanisms — undisclosed training signal, better planning heads, or both — that keep the error rate flat out to 100+ steps.

This axiom closes that gap through public methodology. The solver delivers a training recipe, fine-tuned checkpoint, or inference protocol that, applied to a base model under 72B active parameters, hits ≥90% successful-termination on a fixed benchmark of 500 tasks. Each task requires at least 50 tool calls to complete. "Successful termination" is defined narrowly: the agent stops within its step budget, makes no calls to undefined tool names, and emits no malformed tool-call JSON. The task need not be solved correctly for this axiom — only terminated cleanly. Correctness is measured by adjacent benchmarks.

Verification uses a deterministic agent harness with mocked tool implementations and fixed seeds. The protocol replays the solver's recipe or inference protocol end-to-end on a clean environment and measures the termination rate.

Isolating termination from correctness is deliberate. Most labs collapse the two, which means improvements in one are obscured by noise in the other. Breaking this apart gives researchers a single-axis target.