58: Astra, Day Four: The Threshold Was the Mouse
Four days after GPT-6 Astra shipped, the aggregate benchmarks say almost nothing moved. Artificial Analysis scores it 61, identical to GPT-5.6 Sol. Epoch calls its record score "within the uncertainty range" of the existing trend. The people using it say somet...
Show Notes
Episode 0058: Astra, Day Four: The Threshold Was the Mouse
Episode 0058 | DTF:FTL | September 2026
Why it matters. Four days after GPT-6 Astra shipped, the aggregate benchmarks say almost nothing moved. Artificial Analysis scores it 61, identical to GPT-5.6 Sol. Epoch calls its record score "within the uncertainty range" of the existing trend. The people using it say something else: a developer routing $330,000 a month of inference reports Astra navigated 150 pages of medical dashboards in 15 minutes and that "Codex now uses my computer more than I do." The threshold everyone watched was benchmarks. The threshold that mattered was the mouse. Once a model can operate software it has never seen, every program becomes an actuator with no API required. This follow-up to episode 0057 reads the first four days of fallout: the split benchmark profile, the pricing whiplash, the pull request Astra fixed but did not push, the Sanders/Casar bill, and the open-weights clock. The recalibration: jagged is not small, task-level labor reprices before job-level, verification is the scarce skill, and leverage goes to whoever can direct and check the work, so access has to stay open or it concentrates fast.
Computer Use: The Threshold
- WIRED, "OpenAI Says GPT-6 Can Use a Computer Better Than a Human": https://www.wired.com/story/openai-says-gpt-6-can-use-a-computer-better-than-a-human/
- OpenAI announcement (computer use section): https://openai.com/index/gpt-6-astra/
- Business Insider, "Welcome to the AGI era" (Brockman press call): https://www.businessinsider.com/astra-model-launch-agi-milestone-openai-greg-brockman-2026-9
- Axios: https://www.axios.com/2026/09/03/openai-astra-gpt-6-agi-brockman
- Theo (t3.gg) field report, summarized by BigGo: https://finance.biggo.com/news/8091413d579e4abf
- Decrypt, "Shockingly Good at Almost Everything" (Clip Studio Paint colorist, 3D demos, writing results): https://decrypt.co/377514/openai-gpt-6-astra-review-shockingly-good
- Latent Space / AINews launch roundup (544 accounts, 12 subreddits): https://www.latent.space/p/ainews-gpt-6-astra-openais-biggest
- GPT-6 Astra computer-use guide summary (code-execution vs structured computer tool): https://www.elser.ai/news/gpt-6-astra-computer-use-guide
Jagged, Measured
- Artificial Analysis, "Benchmarking GPT-6 Astra": https://artificialanalysis.ai/articles/benchmarking-gpt-6-astra
- Coding Agent Index 67 (Opus 5 and Fable 5 parity; Fable 5.1 at 70); 70% more token-efficient than Sol
- Intelligence Index 61, equal to Sol, 5 below Fable 5.1, behind Muse Spark 1.3
- AA-Briefcase +80 Elo; GDPval-AA v2 -80 Elo; hallucination rate 92% to 51%
- Artificial Analysis model page: https://artificialanalysis.ai/models/gpt-6-astra
- Epoch AI (ECI 169 vs 163, on trend; 2 of 68 Erdős problems): https://epoch.ai/models/gpt-6-astra and https://x.com/EpochAIResearch/status/2095602754282783108
- ARC Prize (62.7% standard harness at $26,098/run; 99.9% provider adapter): https://arcprize.org/arc-agi/3 and https://x.com/arcprize/status/2095597602545025138
- François Chollet ("saturated ~2x faster than I expected"; ARC-AGI-4 Q1 2027): https://x.com/fchollet/status/2095598451115614371
- Editorial-voice writing Elo (Astra 11th, Sol 6th): https://x.com/Whats_AI/status/2096050974037082380
- Bach Benchmark (first model to write passing tones): https://x.com/aug5thmusic/status/2096030719156089029
- Giuseppe Paleologo on portfolio ideas ("Actual creativity is still far, far away"): https://x.com/__paleologo/status/2096058430284787812
- r/LocalLLaMA, "What the Artificial Analysis / GPT-6 Astra mess actually teaches us": https://www.reddit.com/r/LocalLLaMA/comments/1w89mdu/what_the_artificial_analysis_gpt6_astra_mess/
Pricing, Access, Tenancy
- Pricing: $10 / $50 per million tokens standard; $20 / $100 fast. 2.5x GPT-5.6 Sol's current promotional rate.
- Codex usage limits by plan: https://www.codexusage.dev/limits/astra
- Banked reset (Sep 5): https://explainx.ai/blog/openai-astra-banked-reset-rollout-september-2026
- Reported 4x usage-limit cut (Sep 6-7, unconfirmed by OpenAI): https://www.explainx.ai/blog/openai-gpt-6-astra-usage-limits-cut-4x-september-2026
- Post-launch benchmark revisions (Fortune, Sep 4; summarized): https://www.explainx.ai/blog/openai-astra-benchmark-numbers-changed-post-launch-2026
- r/codex, "gpt 6 astra is GENUINELY unusable on plus": https://www.reddit.com/r/codex/comments/1w7uj5i/gpt_6_astra_is_genuinely_unusuable_on_plus/
- OpenAI Developer Community, Plus access clarification request: https://community.openai.com/t/clarification-needed-gpt-6-astra-was-announced-for-all-chatgpt-plus-users-but-plus-access-is-currently-limited-to-work-codex/1395038
- r/OpenAI, "GPT-6" ("$25 secretary or $100 an hour Astra rig?"): https://www.reddit.com/r/OpenAI/comments/1w40n1m/gpt6/
- Azure: GPT-6 Astra in Microsoft Foundry: https://azure.microsoft.com/en-us/blog/gpt-6-astra-frontier-intelligence-for-work-now-generally-available-in-microsoft-foundry/
- 9to5Mac (Pro/Business rollout, Sep 4): https://9to5mac.com/2026/09/04/openai-releasing-major-upgrade-to-chatgpt-and-codex-with-gpt-6-astra-details-here/
Verification and Agency
- Theo's scrolling-bug PR (fixed, did not push) and Lakebed latency overhaul (800 ms to under 30 ms, verified by Fable 5.1): https://finance.biggo.com/news/8091413d579e4abf
- Matt Shumer's overnight survival world ("they'd started talking to each other"): https://x.com/mattshumer_/status/2095596175705399482
- r/singularity, "After trying Astra" (MTG compiler; unprompted Google Drive search): https://www.reddit.com/r/singularity/comments/1w7gwe4/after_trying_astra/
- System card (computer-use safety, prompt injection robustness): https://deploymentsafety.openai.com/gpt-6-astra
- METR investigation of the Hugging Face incident (context): https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/
The Ban Artificial Superintelligence Act
- Sanders/Casar press release (Sep 3): https://www.sanders.senate.gov/press-releases/news-sanders-casar-introduce-legislation-to-ban-artificial-superintelligence-and-temporarily-pause-advanced-ai-development/
- The Hill: https://thehill.com/policy/technology/6069131-sanders-casar-ai-superintelligence-ban/
- Gary Marcus, "why I oppose it": https://garymarcus.substack.com/p/the-new-sanders-casar-ban-artificial
- Science, "experts can't agree on what the term means": https://www.science.org/content/article/bernie-sanders-aims-ban-ai-superintelligence-experts-can-t-agree-what-term-means
- c4573.org, "Chicken Little Goes to Washington" (Sep 4): https://c4573.org/blog/
- c4573.org, "You Can't Ban Math": https://c4573.org/blog/
The Open-Weights Clock
- Local AI Zone, September 2026 model updates (open-weights wave, price ledger, DevDay Sep 29): https://local-ai-zone.github.io/blog/September_2026_AI_Model_Updates.html
- Kimi K3 (Moonshot), Qwen3.8 family (Alibaba), Hy4 preview (Tencent, Apache 2.0), GLM-5.3 / GLM-5.3-Flash (Z.ai), Muse Spark 1.3 (Meta), Nemotron 3.5 Lightning (NVIDIA), MiniMax H3
- OSWorld 2.0 benchmark: https://os-world.github.io/
- ScreenSpot-Pro: https://github.com/likaixin2000/ScreenSpot-Pro-GUI-Grounding
- What to watch: the best OSWorld score on hardware you own; GLM-5.3 weights (mid-to-late September); Muse open weights (~Q4); OpenAI DevDay, September 29
Recalibration Threads
- r/singularity, "What are your thoughts? I still believe AGI is a long way off" (L4 geofencing analogy; "AGI will arrive many more times now"): https://www.reddit.com/r/singularity/comments/1w9ag49/what_are_your_thoughts_i_still_believe_agi_is_a/
- r/singularity, "Have we reached AGI?": https://www.reddit.com/r/singularity/comments/1w6s5og/have_we_reached_agi/
- Pivot to AI, "GPT-6 is totally Artificial General Intelligence, guys" (the "nothing changed" case): https://pivot-to-ai.com/2026/09/04/openai-gpt-6-is-totally-artificial-general-intelligence-guys/
- Chubby on X (launch timing and the two IPOs): https://x.com/kimmonismus/status/2095613117904347260
- Ethan Mollick, Zork to 3D: https://x.com/emollick/status/2096047660662722620
Previous Episode
- Episode 0057, "GPT-6 Astra: The Capability Jump That Shipped With a Blind Spot" (launch, benchmark footnotes, Critical cyber rating, Section 9 monitorability, recurrent depth): https://pod.c457.org/dtfftl/
Disclosure
Scripts for this show are written by a Claude model. Astra is Claude's direct competitor and every comparison table in this story has a Claude column. We named the conflict on air. Judge the argument on the sources above.
This podcast is entirely AI generated. Not affiliated with OpenAI, Anthropic, or any organization discussed.