Codex
Research: a different coding agent as reviewer beats a bigger budget (Meta's RankEvolve study)
In a Meta study, having Codex review and repair Claude Code's patches raised correctness from 45.8% to 62.5% on 96 ML-code tasks at about the same budget, while more Claude Code passes barely helped.
What changed
Meta researchers published RankEvolve (arXiv 2609.39551, September 30, 2026), a system where coding agents edit real ML training code, launch experiments, and keep iterating. Inside it is a benchmark of 96 tasks graded by hidden executable tests that check whether a patch is actually correct, not just whether it runs.
At roughly the same observable budget, a chain where Claude Code wrote the change, Codex reviewed and repaired it, and Claude Code finished it got 62.5% of tasks right. The best setup that used only Claude Code (three Claude Code passes in a row) got 45.8%, best-of-N sampling got 43.8%, and a single Claude Code run with a bigger budget got 33.3%.
The mixed chain also cut silent critical defects, meaning code that runs and prints a believable number but is wrong, to 10.4%, compared with 16.7% for same-product review and 27.1% for the bigger single budget. On a second codebase, LitGPT, the mixed setup beat the best same-product setup by 12.5 points.
Youssef Hosni's summary on X adds the likely reason: mistakes made by a Claude Code reviewer lined up with the author's mistakes far more often (correlation 0.58) than a Codex reviewer's did (0.21), so a different product has more to catch.
Who is affected
Developers who already have both Codex and Claude Code, and anyone tempted to fix a shaky agent run by just giving the same agent more turns or more reruns. It matters most for code where bugs don't crash: training loops, evaluation scripts, metrics, and data pipelines.
What to do now
- For risky diffs, make the reviewer a different product than the author. Let Claude Code write and Codex review and repair, or the reverse, then give the author one final pass. Our earlier piece on cross-model code review shows the commands.
- Don't treat "it ran" as a pass. Keep a test for the behavior that matters, such as no validation data leaking into training, eval flags actually wired, and gradients still flowing, and run it yourself.
- Break long jobs into fixed steps with checkpoints (plan, implement, review, test) instead of one long prompt. The paper's runtime enforces the workflow step by step rather than trusting the prompt alone.
- Before paying for more budget on the same agent, try a second product as reviewer. In this study, more same-agent budget was the weakest option.
What is not confirmed
The results come from two Python and PyTorch ML codebases (HSTU and LitGPT), so they may not carry over to web apps or other languages. Budgets were matched on what the researchers could see; compute on the providers' side couldn't be matched. The paper is an arXiv preprint and hasn't been peer reviewed. The 0.58 and 0.21 correlation figures come from the X summary and weren't checked against the paper's tables.
Sources
- OfficialRankEvolve paper on arXiv · Sep 30, 2026
- ReportYoussef Hosni on X · Oct 7, 2026
Get updates like this every morning
- ① Email
- ② Card on Stripe
- ③ 7 days free
Then $2/month · cancel anytime in one click