Date: September 4th, 2026 2:26 AM Author: The Penis
99.9%. gpt-6-astra
30.2%. claude-opus-5
7.8%. gpt 5.6 sol
ARC-AGI-3 tests how well agents learn as they solve unfamiliar interactive tasks. GPT‑6 Astra saturates the eval, scoring 99.9%. The average human tester scored 48%. GPT‑6 Astra was measured with our responses API harness, which better reflects real-world performance than the original benchmark harness, which discards past reasoning and past messages. With this harness, we estimate Sol would score in the ballpark of ~30%
(http://www.autoadmit.com/thread.php?thread_id=5900459&forum_id=2\u0026mark_id=5310486#50114904) |