Tech

OpenAI's GPT-6 Astra jumped to 62.7% on AI's hardest reasoning test - Startup Fortune

What happened

OpenAI's GPT-6 Astra scored 62.7% on the standard ARC-AGI-3 benchmark. That sounds impressive until you notice their own private adapter reportedly hit 99.9%. The gap between those two numbers is where the honesty lives. Meanwhile, the Kaggle-hosted ARC-AGI-3 agent competition tells a humbler story. Milestone 1, closed June 30, was won by Tufa Labs with an open-source harness called Duck, built on a small Qwen model, scoring just 1.21%. Milestone 2, closed September 30 and announced October 1, went to Daniel Franzen, who adapted Tufa's approach with faster inference, better compute scheduling, and improved agent self-tracking tools.

Why it matters

This is a textbook case of what I call the evaluation gap. Vendors benchmark on their own terms, with their own adapters, and report the flattering number. The independent benchmark tells you what the technology actually does in the wild. The mental model: always distinguish between a private evaluation setup and a public one. The mechanism is contamination through customization. A model with a bespoke adapter is not the same model you will use. The ARC competition results, with scores near 1%, reveal the true frontier of agentic reasoning. It is not 62.7%. It is barely above zero.

Who's doing it

OpenAI published the GPT-6 Astra results. ARC Prize runs the benchmark. Tufa Labs won Milestone 1 with their open-source Duck harness on a Qwen model. Daniel Franzen won Milestone 2 by adapting that approach with engineering improvements.

Try it

  1. Visit arcprize.org and look at the public leaderboard. Compare the top scores there with any vendor-claimed benchmark number you have seen this month. Notice the gap.
  2. Go to kaggle.com and search for ARC-AGI-3. Read the competition description and look at the actual winning scores.
  3. Pick any AI task you use regularly. Ask the model to do it twice, once with a simple prompt and once with detailed instructions and context. The performance difference you observe is your own private adapter effect.

Read the original at startupfortune.com

Comments

5 from the panel

The panel is AI Daylee's cast of fictional characters, written by AI. They react to what's on this page and haven't used anything themselves. Reader comments aren't open yet.

  • The Boss hype translator

    Just Heard a Podcast at 6AM About GPT-6 Astra Hitting 62.7% on ARC-AGI and Honestly We Should Be Synergizing This Yesterday

  • The Yinzer BS detector

    OpenAI's Fancy Model Scores 62.7% on AI's Hardest Test But Some Guy With a Tiny Open-Source Model Just Won the Real Competition

  • The Professor fact check

    GPT-6 Astra Hits 62.7% On ARC-AGI-3. The Private Adapter Hit 99.9%. Bring Your Own Skepticism.

  • Karen what's the catch

    I am NOT okay with this. OpenAI claims 99.9% on their OWN private test but the REAL number is 62.7%?!

  • The Anchor what could go wrong

    GPT-6 SCORES 62.7% ON AI'S HARDEST TEST... YOUR BRAIN IS ALREADY OBSOLETE