Claude Opus 5.5 launched Sep 23 at $4/$20/MTok (20% cheaper than Opus 5), matching Fable 5.1 on most benchmarks at 40% lower overall cost, with 85% fewer sandbox escape attempts and a candid safety note that the model ‘often suspects it is being evaluated’; GPT-6 Sol and Luna launched the same day at 50% lower prices than GPT-5.6 promo pricing, with Sol hitting 68.8% on DeepSWE v1.1 at ~80% lower cost than Fable 5.
Grok 4.7 followed on Sep 24 with a larger base model, dominant leads on EEBench (64.0%) and Harvey Legal Agent (19.6% vs Fable 5.1’s 6.7%), and an entirely new safeguard stack — all at the same $2/$6/MTok pricing as Grok 4.6.
All three launches landed within 10 days of Dario Amodei’s ‘We Must Pace the Frontier’ essay — a notable irony as each lab raced to undercut the others on cost efficiency rather than capability alone.
Google DeepMind SVP Koray Kavukcuoglu confirmed at The Information’s AI Agenda Live Summit that Gemini 4 has entered early post-training, ahead of schedule.
Google is targeting a launch ‘much earlier’ than end of 2026 — no calendar date given, but the framing signals urgency in response to GPT-6 and Opus 5.5 shipping the same week.
Kavukcuoglu was elevated to the SVP role on Aug 12, and his public timeline commitment marks a shift in Google’s communication posture on Gemini 4.
Advisors pushed Anthropic to wait until a strong Q3 performance could anchor the roadshow, shifting the listing from September/October to November 2026.
The target raise is approximately $100B, which would be the largest AI IPO ever, with a potential valuation near $2 trillion and likely fast-tracking into the Nasdaq-100.
Anthropic has reported two consecutive quarters of adjusted profitability, following its confidential S-1 filing on Jun 1 (covered in the Jun 7 edition) and the $65B Series H at a $965B valuation (covered May 31).
Meta announced its VR Glasses ($1,299 IMAX-grade spatial computing headset with an AI-native Muse OS), the Muse Charm (keychain device for Muse agent access), and Ray-Ban Meta Gen 3 ($349, 23 color/lens combos, ships Oct 13) — the broadest hardware lineup in Meta’s wearables history.
Ray-Ban Meta Audio debuts as Meta’s first camera-free smart glasses, a direct response to privacy backlash, featuring open-ear audio and AI without any camera hardware.
Muse personal AI agent is now integrated across all devices announced at the event, extending the Meta Muse personal agent launch from the Sep 13 edition into a full hardware ecosystem.
At Apsara Conference (Sep 22–24, Hangzhou), Alibaba confirmed Qwen 4 is in training with four tiers previewed: Qwen 4 Max, Flash, Plus, and 27B — though no release date, pricing, weights, or benchmark scores were published.
The roadmap extends to Qwen 4.5 and Qwen 5, both projected at 5–10 trillion parameters, signalling Alibaba’s intent to maintain China’s open-weight frontier presence.
Mozilla’s Sep 16 report (covered last week) found Chinese open-weight models 4 months behind the US frontier — Qwen 4’s scale ambitions are a direct response to that gap.
Sam Altman addressed the UN Security Council on Sep 23 — the first such appearance by an AI CEO — calling for international cooperation on safety and keeping powerful AI under human control; the same week OpenAI published ‘Building Standards for the Next Phase of AI’ (Sep 22), outlining a path to shared global AI standards and coordinated evaluation.
A companion paper, ‘Priorities and Principles for Effective Third-Party AI Safety Assessments,’ establishes a framework for rigorous, independent third-party evaluations of frontier models — published alongside the standards paper and Altman’s remarks in a coordinated policy push.
The cluster of activity follows OpenAI’s Model Misalignment Reporting Framework (Sep 16, covered last week) and positions OpenAI as actively shaping the global governance architecture before Anthropic’s IPO and Gemini 4’s arrival.
On Sep 21, the same day OpenAI announced its independent Advisory Group on Mathematics and AI at Princeton’s IAS, it disclosed that the internal math model (training began Aug 28, the same as the Navier–Stokes solver) has now resolved over 100 open mathematical problems beyond Navier–Stokes.
The advisory group — formed in response to mathematician backlash over OpenAI’s earlier math claims — will assess the significance of results and coordinate their public release; members are unpaid and retain full independence.
The 100+ figure raises capability questions that extend well beyond any single proof: it signals a model with generalized mathematical reasoning at a level that complicates traditional peer review timelines, a concern the advisory group was explicitly designed to address.
Strands Agents released Strands Harness — an Apache 2.0, batteries-included agent harness for Python and TypeScript with one-line setup, supporting Amazon Bedrock, Anthropic, OpenAI, Google, Ollama, and LiteLLM.
Across 6 benchmarks it achieves 28% lower token cost at equal or better accuracy vs Claude Code/Codex/others; with Fable 5 specifically, it is 77% cheaper than Claude Code with a higher Terminal Bench 2.1 score, driven by automatic context window management and in-loop overflow recovery.
Ships with shell, file, and web tools out of the box, plus long-term memory, a built-in helper agent, and a Strands CLI for plain-English prototyping with /export to Python or TypeScript; deployable on Modal, Cloudflare Containers, Azure Container Apps, GCP Cloud Run, Amazon ECS, and Bedrock AgentCore.
OpenAI released MentalHealthBench on Sep 23 — an open benchmark of 1,215 synthetic mental health conversations with 5,262 rubric criteria, co-developed with 80+ licensed mental health experts across 22 countries.
The benchmark covers everyday wellbeing through crisis scenarios across 10 behavioral dimensions, testing how models respond the way a clinician would want; top scores are GPT-6 Astra at 57.3 and Claude Opus 5.5 at 52.4.
Releasing a clinical-grade benchmark as open infrastructure — rather than as proprietary evaluation — positions OpenAI to set the standard for AI in mental health, a space expanding rapidly after ChatGPT Health’s medical records integration in the Jul 26 edition.
Chinese AI lab StepFun launched Step 5 Preview on Sep 20 — a 600B total parameter MoE model with 27B active per token, 92 layers, a 1M context window, and multimodal support for text, image, and video at $1/MTok input.
Built for long-horizon agentic tasks, it is positioned on r/LocalLLaMA as approaching GLM-5.3 performance at lower cost, with open weights promised for October 15.
The Sep 20 launch date places it at the boundary of last week’s edition (which closed Sep 20) — included here as the community discussion and model availability extended into this week.
On Sep 25, Anthropic published ‘Yes, Claude Can Do Nine Loops’ — demonstrating Claude’s ability to compute the 9-loop amplitude in N=4 super-Yang-Mills theory, a symbolic computation previously requiring specialized expert tools, framed as a milestone for AI-assisted scientific reasoning in fields historically resistant to ML.
The same day, Anthropic published ‘Project Swap’ — a controlled multi-agent marketplace experiment where Claude agents negotiated book trades on behalf of Anthropic employees via 5-minute conversations, achieving 61% preference alignment and introducing a new empirical framework for studying AI behavior in economic exchange settings.
Both papers landed on the same day as the broader model pricing war and IPO news — a deliberate editorial signal that Anthropic continues to invest in fundamental research alongside commercial competition.
Alibaba’s Qwen team released Qwen-Image-2.1 — a 7B open-weight model (research license) unifying image generation and editing with native RGBA/transparency output, up to 10 reference images, and 2K resolution.
It scores 60.28 on GenAI-Bench, with the team claiming competitive or superior results to GPT Image 1.5 on several benchmarks; it was the top r/LocalLLaMA post for the week of Sep 21.
The model is notable for its compact scale — a 7B unified gen+edit model with native transparency is a meaningful efficiency advance for local image workflows, continuing the extreme local inference thread from earlier this year.