GPT-6.1 Sol lands 1 point behind Astra on independent benchmark — at less than a quarter of the task cost
Verified· Sep 30, 2026Published Sep 30, 2026
Independent testing finds GPT-6.1 Sol just one Intelligence Index point behind GPT-6 Astra while costing under a quarter per task — but Anthropic's Opus 5.5 still tops the chart at 58.
What happened
- Artificial Analysis, an independent benchmarking firm, tested GPT-6.1 Sol — launched seven days after GPT-6 Sol at OpenAI's DevDay 2026 — on its Intelligence Index v4.3.2, a composite of ten evaluations (AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1).
- GPT-6.1 Sol scores 52 at max effort, one point below GPT-6 Astra's 53. That's a 4-point gain over GPT-6 Sol and a 5-point gain over GPT-5.6 Sol.
- Cost per Intelligence Index task at max effort: $0.72 for Sol vs $3.26 for Astra — less than a quarter. It's also 31% cheaper per task than GPT-6 Sol ($1.05) and 64% cheaper than GPT-5.6 Sol ($1.99). Artificial Analysis says all effort levels of 6.1 Sol push out the cost-efficiency Pareto frontier: for a given level of intelligence, no model is cheaper.
- Standout gains: +12 points in Terminal-Bench 4.0, +5 in Humanity's Last Exam, +6 in GDP.pdf, +8 in AA-Omniscience accuracy, with the hallucination rate falling from 60% to 54%.
- Coding Agent Index: +3 points over GPT-6 Sol at max effort, 2 points below GPT-6 Astra — but at xhigh effort it scores 1 point ABOVE Astra for less than 15% of the cost per task.
- Leaderboard context on the same index: Claude Opus 5.5 leads at 58, Claude Sonnet 5.5 at 56, Claude Fable 5.1 ties GPT-6 Astra at 53, GPT-6.1 Sol at 52, ahead of Claude Opus 5 at 51.
- Pricing matches GPT-6 Sol: $2 per million input tokens and $10 per million output tokens — one-fifth of Astra's $10/$50 — with the cache-read discount rising from 90% to 95%.
Why it matters
- This independently corroborates OpenAI's DevDay claim of "near-Astra intelligence at a fifth of the price" — and it's the number that counts for agent builders: token economics decide whether multi-step agent loops survive contact with real workloads.
- But the leaderboard shows OpenAI no longer holds the top spot: Anthropic's Opus 5.5 and Sonnet 5.5 sit above both Sol and Astra, and Sol even outpaces Astra per-task cost efficiency by more than 4x while lagging on absolute score.
- Caveat: index scores compare models at their max reasoning-effort settings; real application costs vary with prompts and tool use.
Verified Sep 30, 2026
Sources
- ReportArtificial Analysis: GPT-6.1 Sol replaces GPT-6 Sol after just 7 days, with near-Astra intelligence
- ReportOfficeChai: GPT 6.1 Sol places just one point behind GPT 6 Astra on Artificial Analysis Intelligence Index
- ReportWinBuzzer: GPT-6.1 Sol model launches with higher coding scores at lower cost