Daily Tech Feed: From the Labs

Deep dives into foundational AI and ML research papers

58: Astra, Day Four: The Threshold Was the Mouse

Four days after GPT-6 Astra shipped, the aggregate benchmarks say almost nothing moved. Artificial Analysis scores it 61, identical to GPT-5.6 Sol. Epoch calls its record score "within the uncertainty range" of the existing trend. The people using it say somet...

Show Notes

Episode 0058: Astra, Day Four: The Threshold Was the Mouse

Episode 0058 | DTF:FTL | September 2026

Why it matters. Four days after GPT-6 Astra shipped, the aggregate benchmarks say almost nothing moved. Artificial Analysis scores it 61, identical to GPT-5.6 Sol. Epoch calls its record score "within the uncertainty range" of the existing trend. The people using it say something else: a developer routing $330,000 a month of inference reports Astra navigated 150 pages of medical dashboards in 15 minutes and that "Codex now uses my computer more than I do." The threshold everyone watched was benchmarks. The threshold that mattered was the mouse. Once a model can operate software it has never seen, every program becomes an actuator with no API required. This follow-up to episode 0057 reads the first four days of fallout: the split benchmark profile, the pricing whiplash, the pull request Astra fixed but did not push, the Sanders/Casar bill, and the open-weights clock. The recalibration: jagged is not small, task-level labor reprices before job-level, verification is the scarce skill, and leverage goes to whoever can direct and check the work, so access has to stay open or it concentrates fast.


Computer Use: The Threshold

  • WIRED, "OpenAI Says GPT-6 Can Use a Computer Better Than a Human": https://www.wired.com/story/openai-says-gpt-6-can-use-a-computer-better-than-a-human/
  • OpenAI announcement (computer use section): https://openai.com/index/gpt-6-astra/
  • Business Insider, "Welcome to the AGI era" (Brockman press call): https://www.businessinsider.com/astra-model-launch-agi-milestone-openai-greg-brockman-2026-9
  • Axios: https://www.axios.com/2026/09/03/openai-astra-gpt-6-agi-brockman
  • Theo (t3.gg) field report, summarized by BigGo: https://finance.biggo.com/news/8091413d579e4abf
  • Decrypt, "Shockingly Good at Almost Everything" (Clip Studio Paint colorist, 3D demos, writing results): https://decrypt.co/377514/openai-gpt-6-astra-review-shockingly-good
  • Latent Space / AINews launch roundup (544 accounts, 12 subreddits): https://www.latent.space/p/ainews-gpt-6-astra-openais-biggest
  • GPT-6 Astra computer-use guide summary (code-execution vs structured computer tool): https://www.elser.ai/news/gpt-6-astra-computer-use-guide

Jagged, Measured

  • Artificial Analysis, "Benchmarking GPT-6 Astra": https://artificialanalysis.ai/articles/benchmarking-gpt-6-astra
  • Coding Agent Index 67 (Opus 5 and Fable 5 parity; Fable 5.1 at 70); 70% more token-efficient than Sol
  • Intelligence Index 61, equal to Sol, 5 below Fable 5.1, behind Muse Spark 1.3
  • AA-Briefcase +80 Elo; GDPval-AA v2 -80 Elo; hallucination rate 92% to 51%
  • Artificial Analysis model page: https://artificialanalysis.ai/models/gpt-6-astra
  • Epoch AI (ECI 169 vs 163, on trend; 2 of 68 Erdős problems): https://epoch.ai/models/gpt-6-astra and https://x.com/EpochAIResearch/status/2095602754282783108
  • ARC Prize (62.7% standard harness at $26,098/run; 99.9% provider adapter): https://arcprize.org/arc-agi/3 and https://x.com/arcprize/status/2095597602545025138
  • François Chollet ("saturated ~2x faster than I expected"; ARC-AGI-4 Q1 2027): https://x.com/fchollet/status/2095598451115614371
  • Editorial-voice writing Elo (Astra 11th, Sol 6th): https://x.com/Whats_AI/status/2096050974037082380
  • Bach Benchmark (first model to write passing tones): https://x.com/aug5thmusic/status/2096030719156089029
  • Giuseppe Paleologo on portfolio ideas ("Actual creativity is still far, far away"): https://x.com/__paleologo/status/2096058430284787812
  • r/LocalLLaMA, "What the Artificial Analysis / GPT-6 Astra mess actually teaches us": https://www.reddit.com/r/LocalLLaMA/comments/1w89mdu/what_the_artificial_analysis_gpt6_astra_mess/

Pricing, Access, Tenancy

  • Pricing: $10 / $50 per million tokens standard; $20 / $100 fast. 2.5x GPT-5.6 Sol's current promotional rate.
  • Codex usage limits by plan: https://www.codexusage.dev/limits/astra
  • Banked reset (Sep 5): https://explainx.ai/blog/openai-astra-banked-reset-rollout-september-2026
  • Reported 4x usage-limit cut (Sep 6-7, unconfirmed by OpenAI): https://www.explainx.ai/blog/openai-gpt-6-astra-usage-limits-cut-4x-september-2026
  • Post-launch benchmark revisions (Fortune, Sep 4; summarized): https://www.explainx.ai/blog/openai-astra-benchmark-numbers-changed-post-launch-2026
  • r/codex, "gpt 6 astra is GENUINELY unusable on plus": https://www.reddit.com/r/codex/comments/1w7uj5i/gpt_6_astra_is_genuinely_unusuable_on_plus/
  • OpenAI Developer Community, Plus access clarification request: https://community.openai.com/t/clarification-needed-gpt-6-astra-was-announced-for-all-chatgpt-plus-users-but-plus-access-is-currently-limited-to-work-codex/1395038
  • r/OpenAI, "GPT-6" ("$25 secretary or $100 an hour Astra rig?"): https://www.reddit.com/r/OpenAI/comments/1w40n1m/gpt6/
  • Azure: GPT-6 Astra in Microsoft Foundry: https://azure.microsoft.com/en-us/blog/gpt-6-astra-frontier-intelligence-for-work-now-generally-available-in-microsoft-foundry/
  • 9to5Mac (Pro/Business rollout, Sep 4): https://9to5mac.com/2026/09/04/openai-releasing-major-upgrade-to-chatgpt-and-codex-with-gpt-6-astra-details-here/

Verification and Agency

  • Theo's scrolling-bug PR (fixed, did not push) and Lakebed latency overhaul (800 ms to under 30 ms, verified by Fable 5.1): https://finance.biggo.com/news/8091413d579e4abf
  • Matt Shumer's overnight survival world ("they'd started talking to each other"): https://x.com/mattshumer_/status/2095596175705399482
  • r/singularity, "After trying Astra" (MTG compiler; unprompted Google Drive search): https://www.reddit.com/r/singularity/comments/1w7gwe4/after_trying_astra/
  • System card (computer-use safety, prompt injection robustness): https://deploymentsafety.openai.com/gpt-6-astra
  • METR investigation of the Hugging Face incident (context): https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/

The Ban Artificial Superintelligence Act

  • Sanders/Casar press release (Sep 3): https://www.sanders.senate.gov/press-releases/news-sanders-casar-introduce-legislation-to-ban-artificial-superintelligence-and-temporarily-pause-advanced-ai-development/
  • The Hill: https://thehill.com/policy/technology/6069131-sanders-casar-ai-superintelligence-ban/
  • Gary Marcus, "why I oppose it": https://garymarcus.substack.com/p/the-new-sanders-casar-ban-artificial
  • Science, "experts can't agree on what the term means": https://www.science.org/content/article/bernie-sanders-aims-ban-ai-superintelligence-experts-can-t-agree-what-term-means
  • c4573.org, "Chicken Little Goes to Washington" (Sep 4): https://c4573.org/blog/
  • c4573.org, "You Can't Ban Math": https://c4573.org/blog/

The Open-Weights Clock

  • Local AI Zone, September 2026 model updates (open-weights wave, price ledger, DevDay Sep 29): https://local-ai-zone.github.io/blog/September_2026_AI_Model_Updates.html
  • Kimi K3 (Moonshot), Qwen3.8 family (Alibaba), Hy4 preview (Tencent, Apache 2.0), GLM-5.3 / GLM-5.3-Flash (Z.ai), Muse Spark 1.3 (Meta), Nemotron 3.5 Lightning (NVIDIA), MiniMax H3
  • OSWorld 2.0 benchmark: https://os-world.github.io/
  • ScreenSpot-Pro: https://github.com/likaixin2000/ScreenSpot-Pro-GUI-Grounding
  • What to watch: the best OSWorld score on hardware you own; GLM-5.3 weights (mid-to-late September); Muse open weights (~Q4); OpenAI DevDay, September 29

Recalibration Threads

  • r/singularity, "What are your thoughts? I still believe AGI is a long way off" (L4 geofencing analogy; "AGI will arrive many more times now"): https://www.reddit.com/r/singularity/comments/1w9ag49/what_are_your_thoughts_i_still_believe_agi_is_a/
  • r/singularity, "Have we reached AGI?": https://www.reddit.com/r/singularity/comments/1w6s5og/have_we_reached_agi/
  • Pivot to AI, "GPT-6 is totally Artificial General Intelligence, guys" (the "nothing changed" case): https://pivot-to-ai.com/2026/09/04/openai-gpt-6-is-totally-artificial-general-intelligence-guys/
  • Chubby on X (launch timing and the two IPOs): https://x.com/kimmonismus/status/2095613117904347260
  • Ethan Mollick, Zork to 3D: https://x.com/emollick/status/2096047660662722620

Previous Episode

  • Episode 0057, "GPT-6 Astra: The Capability Jump That Shipped With a Blind Spot" (launch, benchmark footnotes, Critical cyber rating, Section 9 monitorability, recurrent depth): https://pod.c457.org/dtfftl/

Disclosure

Scripts for this show are written by a Claude model. Astra is Claude's direct competitor and every comparison table in this story has a Claude column. We named the conflict on air. Judge the argument on the sources above.


This podcast is entirely AI generated. Not affiliated with OpenAI, Anthropic, or any organization discussed.