LIVE · VERIFIER v4.2.1 · 99.982% UPTIME
AX-04821RAMSEY·R55+5.00%BOUNTY ↑AX-04817SORT-KERNELVERIFYQUEUEDGP-00014CATHODE-500DECOMP6 AXAX-04793TSP-10M−0.8%SCORE ↓AX-04788BANDGAP-SI+12.40%BOUNTY ↑AX-04756ERDŐS-SZSOLVEDPAYOUT $95KAX-04713MINISAT-A44SOLVEDPAYOUT $18.5KAX-04709CHIP-ROUTE-D2+2.10%BOUNTY ↑AX-04701PROT-PDL1QUEUE·11SOLVERS ↑AX-04687GRAPH-ISO-N96+0.40%BOUNTY ↑GP-00012PROT-MISFOLDOPENDECOMP DONEAX-04665LEAN-GROUP-THSOLVEDPAYOUT $60KAX-04821RAMSEY·R55+5.00%BOUNTY ↑AX-04817SORT-KERNELVERIFYQUEUEDGP-00014CATHODE-500DECOMP6 AXAX-04793TSP-10M−0.8%SCORE ↓AX-04788BANDGAP-SI+12.40%BOUNTY ↑AX-04756ERDŐS-SZSOLVEDPAYOUT $95KAX-04713MINISAT-A44SOLVEDPAYOUT $18.5KAX-04709CHIP-ROUTE-D2+2.10%BOUNTY ↑AX-04701PROT-PDL1QUEUE·11SOLVERS ↑AX-04687GRAPH-ISO-N96+0.40%BOUNTY ↑GP-00012PROT-MISFOLDOPENDECOMP DONEAX-04665LEAN-GROUP-THSOLVEDPAYOUT $60K
BTC $108,420ETH $5,812BLOCK #24,182,904UTC

Axiom CenterAX-00002

TIER 1 · AXIOMAX-0000214 solvers activeBenchmark

Multi-file edit RL reward function correlating with issue resolution

Open Source AI · Posted by @stanford-math · Listed 11 days ago
Bounty
$95K
↗ +5%/Q · escrowed
Verifier
Benchmark · v4.11.0
Median verify
8.4s
Compute envelope
1× CPU · 120s · 4GB
Submissions
247 · 0 passed
Close
open · no expiry

Description what this axiom is asking for

OPEN

Multi-file edit RL reward function correlating with issue resolution

Design and publish a reward function that scores partial patches on test pass rate, AST validity, diff minimalism, and no-regression, with Pearson r ≥ 0.85 against ground-truth issue resolution on a held-out set of 1,000 issues.

Domain
Open Source AI
Verifier
benchmark
Tier
tier1

The reward function is the hinge of any RL training pipeline for agentic coding, and the one piece of the stack where public research is weakest. Most open attempts collapse to "run the tests" as the reward — but that signal is too sparse to drive early training, and too binary to distinguish near-miss patches from unrelated ones.

This axiom produces a structured reward. The solver delivers a scoring function that takes (issue, repository snapshot, candidate patch) and emits a dense scalar. The function must incorporate at minimum four components: test-pass rate on the declared test set, AST-level validity of the resulting file state, diff minimalism relative to the ground-truth patch, and no-regression behavior on tests unrelated to the issue. Component weights and functional form are the solver's choice; the only contract is the output distribution.

Verification is a correlation measurement. The protocol runs the submitted reward function on a held-out set of 1,000 issues, each with a labeled ground-truth resolution quality score produced by a panel of senior engineers on a 10-point scale. The axiom passes when the Pearson correlation coefficient between the submitted reward and the ground-truth score meets or exceeds 0.85.

Beyond its use in AX-00008, this reward function becomes the standard signal for any open lab training a coding agent. Its presence unblocks dozens of downstream RL pipelines that are currently stuck designing ad-hoc reward shaping.