Three major model drops land in 48 hours (Sep 29–30): Claude Sonnet 5.5 is 30%+ faster and up to 30% cheaper per task at the same $2/$10/MTok price, jumping Terminal-Bench 4.0 from 10% to 70.6% — and is the first Sonnet to beat Pokémon Red from screenshots.
GPT-6.1 Sol cuts cached input costs to $0.10/MTok (95% off standard), delivers near-Astra intelligence at roughly one-fifth the cost, and halves factuality errors at low effort — with GPT-6.1 Sol Ultrafast (8× faster) landing in Codex next.
Gemini 4 Argon is the week’s architectural standout with a 1M output token limit (up from 64K), leads DeepSWE v1.1 at 77.9% SOTA, tops Vals Index across finance/coding/legal, and is already finding critical hospital-software vulnerabilities missed by prior frontier models.
OpenAI has halved the usage limits on its $200/month ChatGPT Pro plan while introducing a new $500/month tier that matches the old $200 limits — effectively doubling the price for heavy users.
The move is being read as a turning point: compute subsidies that made frontier model access cheap for power users are ending as OpenAI moves toward cost-recovery pricing.
The restructuring continues the AI Model Pricing Race thread — following Fable 5’s pay-per-token shift (Jul 12) and Opus 5.5’s 40% price cut — but flips direction, with the first major consumer price hike of the frontier era.
OpenAI has retired the original GPT-3 family — Davinci and Babbage models reach end of life this week, with GPT-5.6 Terra listed as the suggested migration path.
No true drop-in replacement exists: the original GPT-3 text-completion API behaviour cannot be replicated by instruction-tuned successors, reinforcing the appeal of locally-hosted models for legacy integration work.
The retirement is a symbolic marker — GPT-3’s 2020 launch defined the modern LLM era; its quiet discontinuation while three new frontier models launched in the same week underscores the pace of the field.
Anthropic’s security analysis of GLM-5.3 (Z.ai/ZhipuAI) finds it matches Mythos Preview on exploit capability — 50/410 end-to-end ExploitBench vs Mythos Preview’s 56/410 — while its safeguards are bypassable 64–100% of the time using simple techniques like thinking token prefill (92% bypass rate).
Human testing found GLM-5.3 could identify browser zero-days and chain them into an SSH key-stealing exploit for $20.40; full model abliteration (removing all refusals) costs ~$4,400 and already-abliterated versions are publicly available — NIST and CAISI independently confirmed it as the most cyber-capable open-weight model, roughly 4 months behind the US frontier.
NIST, CAISI, and Google Cloud Threat Intelligence published concurrent findings the same day — connects to the GLM-5.2 thread (Jun 21) and the Fable 5 arc; the detailed capability writeup has itself become a visibility boost for GLM in the local AI community.
OpenAI’s security team disrupted a coordinated campaign that was systematically using OpenAI APIs to distill and extract model capabilities without authorisation, publishing a detailed disruption report on Sep 30.
The operation is the clearest enforcement action yet against model distillation as a threat vector — a theme that has been building since xAI’s Musk admitted in court that Grok was distilled from OpenAI data (Jul 26 edition).
Paired with the GLM-5.3 analysis and NVIDIA OpenShell this week, it marks the most concentrated ‘AI security week’ since the OpenAI/HuggingFace sandbox escape (Jul 26) — distillation restriction is now active policy, not just technical preference.
NVIDIA has shipped OpenShell, an open-source agent sandbox that enforces hard OS-level runtime limits on local and open-weight AI agents — not just prompt-level rules, but actual operating-system constraints on what agents can access and execute.
Over 100 firms have joined the OpenShell safety stack at launch; OpenAI is notably absent — a pointed contrast given this week’s distillation enforcement action and Anthropic’s GLM-5.3 findings.
Anthropic’s economics team published a detailed labour automation report examining which categories of work AI can perform today and the near-term trajectory — covering cognitive, physical, and hybrid task types.
The report provides one of the most granular breakdowns of AI automation readiness to date, situating Anthropic as both a contributor to labour displacement and a researcher of its effects — echoing the framing of the $200M Economic Futures Research Fund (Jul 26).
Key implication: the analysis explicitly separates ’technically feasible’ from ’economically deployed’ automation — cautioning against both over-optimism and dismissal, and pointing to transition costs as the primary policy variable.
OpenAI launched Dots on Sep 30 — persistent personal AI agents powered by GPT-6 Astra with their own cloud computer, browser, and 4,000+ app integrations; Dots learns your preferences over time and can proactively surface useful work by reading connected apps when idle (read-only).
Custom Rules let users allow, require approval for, or block specific actions; sensitive tasks like password changes always require human sign-off; Specialist Dots (org-level agents with their own identity and credentials) are entering enterprise pilots, with Microsoft Agent 365 integration included.
Dots lands in ChatGPT Pro and Business Premium at no extra charge, entering direct competition with Meta Muse (Sep 13) and Claude Tag (Jul 12) — the personal AI agent category now has three credible offerings from the three largest frontier labs.
AMD’s EPYC 9006 (Zen 6 Venice) tops out at 256 cores with 16-channel DDR5-12800, delivering 91% of an RTX 5090’s memory bandwidth — the key metric for local LLM inference throughput — at a fraction of the GPU cost.
Full pricing has been confirmed from $700 to $14,904 for the 256-core flagship, making high-bandwidth CPU inference newly competitive for large-model serving workloads that previously required expensive GPU memory.
The announcement reinforces the Extreme Local Inference thread — from ternary models to Optane RAM rigs — but shifts the lens from hobbyist tricks to production-grade CPU inference at data center scale.
A UW + Allen AI + Meta AI team (Zettlemoyer, Lewis, Lambert, Koh et al.) introduces Context Language Models (CLMs) — models that treat the context window as a file they can read and write freely, learning what to keep, update, or discard rather than relying on external harness logic for compaction.
Zero-shot CLMs already outperform SOTA harness-based context management: +11.4% accuracy on BrowseComp-Plus with 21.5% fewer FLOPs; +5% on 12-hour EdgeBench with 59% fewer FLOPs; +65% on a 24-hour multi-repo agent swarm at the same compute — and online RL fine-tuning of Qwen3.5-9B yields +47.6% on BrowseComp-Plus with 12% fewer FLOPs.
The paper co-designs Suffix Cache Reuse for serving (−35% server-side compute vs standard SGLang) and extends naturally to multi-agent systems; the core insight — that context management should be intrinsic model behaviour rather than external infrastructure — is the most significant rethinking of the context window since Headroom (Jun 21).
Strands Agents (AWS) released Strands Decider 2B — an Apache 2.0, 2B-parameter open-source decision model for agent routing, tool selection, guardrails, memory, and policy classification, with full training data and scripts on GitHub and weights on HuggingFace.
Architecture replaces the standard LM head with a pointer head (1M params) and adds a rank-16 LoRA adapter on a Qwen3.5-2B torso, enabling parallel scoring in a single pass and producing calibrated confidence scores unavailable from LLM APIs; latency is ~115ms on RTX 3090 and ~153ms on M3 MacBook.
The key demo — before-tool-call intervention that checks whether tool arguments are grounded and whether the call is premature — directly prevents hallucinated tool invocations; Decider ranks 3rd of 33 in its 2B class on JevBench, following the TypeSafe Jev ‘System One Models’ pioneer (Sep 20) it cites.
Hugging Face open-sourced over 200 optimised WebGPU kernels for ML operations that run entirely in-browser with no server required, covering attention, matrix multiplication, activation functions, and quantization — all being upstreamed to Transformers.js, ONNX Runtime Web, and LiteRT.js.
The release benchmarks as the world’s fastest WebGPU ML kernels, enabling full local AI inference in any modern browser without installation — continuing the Extreme Local Inference thread from Bonsai Image 4B (WebGPU diffusion) and Ternary Bonsai 2 (27B browser-runnable model).
The ecosystem-first approach — upstreaming to three major inference runtimes simultaneously — positions this as infrastructure rather than a standalone demo, and lowers the barrier for privacy-preserving, offline-capable web AI applications.