ARC-AGI-3: GPT-5.6 Sol aur Claude Opus 5 ne benchmark ko kaise hilaya
> ARC-AGI-3 benchmark par GPT-5.6 Sol 38.3% aur Claude Opus 5 30.16% score le aaye. Samjho ke yeh score kya measure karta hai aur kyun dono labs ki taraf se yeh ek tactics-driven race ban gayi hai.
🎧 Listen — ~7 min
Ready · ARC-AGI-3: GPT-5.6 Sol aur Claud
Introduction
ARC-AGI-3 March 2026 mein launch hua tha ek interactive reasoning benchmark ke tor par. Iska maqsad yeh check karna tha keh AI agents naye, unfamiliar 2D game environments mein rules infer kar sakte hain, goals discover kar sakte hain, aur progressively harder levels complete kar sakte hain ya nahi. July 2026 mein do baray announcements ne is benchmark ko dobara headline banaya: Anthropic ne Claude Opus 5 release kiya aur usne 30.16% score hasil ki, jabke OpenAI ne reveal kiya keh GPT-5.6 Sol proper API harness ke saath 38.3% tak pahunch sakta hai.

Courtesy: OpenAI. Generated illustration based on data from "How enabling two settings tripled our scores on the ARC-AGI-3 benchmark". Accessed: 2026-07-30.
Yeh screenshot OpenAI ki official post se hai aur directly yeh dikhata hai keh API harness design score par kitna bara farq daal sakti hai. Left side official harness par 13.3% hai, right side retained reasoning + compaction enabled hone par 38.3%.
ARC-AGI-3 actually kya test karta hai
ARC-AGI-2 pattern-matching par based tha — grids mein transformation rules infer karna. ARC-AGI-3 usse ek qadam aage jaata hai: yahan models ko interactive games di jati hain jinke rules explicit nahi hote. Agent har turn par ek action bhejta hai, game state ka text/vision representation milta hai, aur usay levels ke through progress karna hota hai.
Metric Relative Human Action Efficiency (RHAE) hai. Matlab sirf yeh nahi keh level solve hua, balkay kitne actions mein solve hua compared to median human. Har level ke liye action count capped hoti hai — typically five times the median human action count. Is wajah se brute-force exploration heavily penalize hoti hai. Efficient, adaptive reasoning ko reward milti hai.
ARC Prize ka argument yeh hai keh traditional benchmarks saturation ki taraf ja rahe hain; ARC-AGI-3 models ko force karta hai keh wo unfamiliar domains mein generalize karein bina extensive fine-tuning ke.
Claude Opus 5: Anthropic ka verified record
Anthropic ne Claude Opus 5 ko July 24, 2026 ko publicly release kiya. Pricing standard API par $5/million input tokens aur $25/million output tokens hai — half the price of Claude Fable 5. Opus 5 default model hai Claude Max par aur strongest model hai Claude Pro par.
ARC-AGI-3 semi-private set par, high reasoning effort par, Opus 5 ne 30.16% score ki. Previous official high GPT-5.6 Sol ka tha 7.78% at maximum effort. ARC Prize ne is result ko verify kiya. Anthropic ki system card mein details hain keh evaluation standard setup mein run hua.

Courtesy: Anthropic. Generated illustration based on data from "Introducing Claude Opus 5". Accessed: 2026-07-30.
Is screenshot mein announcement date, model positioning, aur ARC-AGI-3 chart sab visible hain. Yeh prove karta hai keh Anthropic ne officially is benchmark par claim kiya hai verified ARC Prize result ke saath.
GPT-5.6 Sol: OpenAI ka harness argument
OpenAI ne July 2026 mein ek detailed post publish ki: “How enabling two settings tripled our scores on the ARC-AGI-3 benchmark.” Unhon ne bataya keh jab unhon ne pehli baar GPT-5.6 Sol ka score dekha — 7.8% official harness par — toh woh confuse hue kyunke yeh model Pokemon FireRed aur complex math problems solve kar chuka tha.
Investigation mein pata chala keh do API settings jo ChatGPT aur Codex mein standard hain, ARC-AGI-3 official harness mein off theen:
- Retained reasoning: model apni chain-of-thought ko multiple tool calls aur turns ke darmiyan retain karta hai.
- Compaction: previous reasoning ko compress karke carry forward kiya jaata hai, jis se output tokens kam hote hain aur context manageable rehta hai.
Jab in settings ko enable karke Responses API harness use ki gayi, toh public set par score 13.3% se 38.3% ho gaya. OpenAI ne official gameplay logs analyze karke estimate kiya keh average human tester 48% par tha, toh GPT-5.6 Sol human baseline ke kareeb aa gaya.

Courtesy: OpenAI. Generated illustration based on data from "How enabling two settings tripled our scores on the ARC-AGI-3 benchmark". Accessed: 2026-07-30.
Yeh visual comparison direct evidence hai keh harness design capability measurement par kis qadar asar andaz hoti hai. Same model, same game, different harness — completely different outcome.
Score controversy: kya yeh AGI proof hai
Nahi. ARC-AGI-3 score ek useful signal hai lekin AGI ka proof nahi. Important caveats:
- RHAE score human efficiency se compare hoti hai, na keh pure solve rate se.
- Models ko scoring function nahi pata hoti; unhe infer karna padta hai, lekin yeh bhi ek skill hai jo specific architectures ke liye optimize ho sakti hai.
- Harness differences, prompting, aur API features score ko significantly move kar sakte hain, jaisay GPT-5.6 Sol ne dikhaya.
- Dono results vendor-reported ya vendor-optimized harness par hain; independent third-party replication abhi public mein limited hai.
Jo important hai woh yeh keh ARC-AGI-3 ab clearly frontier model comparison ka battlefield ban gaya hai, aur is battlefield par model capability se zyada model harnessing bhi matter karti hai.
Practical implications for builders
Agar aap AI agents, coding agents, ya autonomous systems bana rahe hain, toh yahan se kuch lessons hain:
- Reasoning persistence matters. Multi-turn tasks mein agar model apni soch ko carry forward nahi kar sakta, toh wo har turn par scratch se shuru hota hai. Retained reasoning ki tarah features production agents ke liye critical hain.
- Compaction = cost + context. Lambi reasoning chains ko compress karna na sirf tokens bachata hai, balkay context window ko bhi efficient use karti hai.
- Benchmarks incomplete hain. Official harness standardization ke liye achi hai, lekin apne use case ke liye custom harness aur evaluation banana zyada informative ho sakta hai.
- Specialized vs general tension. Opus 5 aur GPT-5.6 Sol dono general frontier models hain lekin specific benchmark par inki performance unki harnessing par depend karti hai.
Frequently asked questions
ARC-AGI-3 ka score percentage kis cheez ka percentage hai?
Yeh Relative Human Action Efficiency (RHAE) percentage hai. Matlab model ne human-level efficient actions mein kitna progress kiya. Yeh simply "percentage of games solved" nahi hai.
Kya GPT-5.6 Sol Claude Opus 5 se behtar hai?
Direct comparison mushkil hai kyunke dono ne alag-alag harness par report kiya. GPT-5.6 Sol ka 38.3% OpenAI ki optimized Responses API harness par hai, jabke Claude Opus 5 ka 30.16% ARC Prize verified semi-private set par hai. Dono impressive hain lekin apples-to-apples comparison abhi nahi hai.
Claude Opus 5 kab release hua?
July 24, 2026 ko Anthropic ne publicly release kiya across its products aur API.
GPT-5.6 Sol retained reasoning kaise use karta hai ARC-AGI-3 par?
Retained reasoning allow karti hai keh model apni chain-of-thought ko multiple turns ke darmiyan bachaye rakhe, jis se har turn par game state scratch se reconstruct nahi karni padti. Compaction is reasoning chain ko compress karke context manageable rakhti hai.
Kya yeh score AGI ki taraf ek qadam hai?
ARC-AGI-3 intentionally aisi capabilities test karta hai jo AGI ke liye zaroori samjhi jati hain: exploration, goal inference, aur efficient adaptation. Lekin ek benchmark score AGI proof nahi; yeh sirf ek directional signal hai.
Conclusion: benchmark race ab harness race bhi ban gayi
ARC-AGI-3 ne July 2026 mein do important data points diye: Claude Opus 5 ne verified 30.16% score set ki, aur GPT-5.6 Sol ne dikhaya keh sahi API harness ke saath 38.3% tak ja sakta hai. Dono announcements ek saath yeh message dete hain keh frontier model capability sirf weights mein nahi, unhe run karne ke tareeqay mein bhi hai.
Builders ke liye takeaway clear hai: agar aapke agents multi-turn complex tasks handle karte hain, toh reasoning persistence, compaction, aur harness design par invest karo. Aur jab bhi benchmark numbers dekho, yeh zaroor check karo keh wo number kis harness, kis setting, aur kis evaluation protocol ne produce kiya.
Tags: AI benchmarks, frontier models, agentic AI, OpenAI, Anthropic, ARC Prize, reasoning
Sources
- OpenAI, "How enabling two settings tripled our scores on the ARC-AGI-3 benchmark." https://openai.com/index/how-two-settings-tripled-our-arc-agi-3-scores/
- Anthropic, "Introducing Claude Opus 5." https://www.anthropic.com/news/claude-opus-5
- ARC Prize, "Claude Opus 5" results page. https://arcprize.org/results/anthropic-claude-opus-5
- ARC Prize, ARC-AGI-3 technical report. https://arcprize.org/media/ARC_AGI_3_Technical_Report.pdf
- ARC Prize, ARC-AGI-3 overview. https://arcprize.org/arc-agi/3
- MLQ.ai, "Claude Opus 5 sets a verified ARC-AGI-3 record, with important caveats." https://mlq.ai/news/claude-opus-5-sets-a-verified-arc-agi-3-record-with-important-caveats/
Keep reading
Related reading
⚡ Daily AI Model Drop — Get Kimi K3 benchmarks before Twitter
Join 2,400+ AI engineers. 1 email/day, no spam, unsubscribe anytime