All Issues
Sep 28 - Oct 04, 2026

AI Weekly: Frontier Sprint, AI Security Week

Models & Releases

1 stories

People & Business

2 stories

GPT-3 Officially Discontinued

  • OpenAI has retired the original GPT-3 family — Davinci and Babbage models reach end of life this week, with GPT-5.6 Terra listed as the suggested migration path.
  • No true drop-in replacement exists: the original GPT-3 text-completion API behaviour cannot be replicated by instruction-tuned successors, reinforcing the appeal of locally-hosted models for legacy integration work.
  • The retirement is a symbolic marker — GPT-3’s 2020 launch defined the modern LLM era; its quiet discontinuation while three new frontier models launched in the same week underscores the pace of the field.

Policy & Ethics

4 stories

OpenAI Disrupts Coordinated Model-Distillation Campaign

  • OpenAI’s security team disrupted a coordinated campaign that was systematically using OpenAI APIs to distill and extract model capabilities without authorisation, publishing a detailed disruption report on Sep 30.
  • The operation is the clearest enforcement action yet against model distillation as a threat vector — a theme that has been building since xAI’s Musk admitted in court that Grok was distilled from OpenAI data (Jul 26 edition).
  • Paired with the GLM-5.3 analysis and NVIDIA OpenShell this week, it marks the most concentrated ‘AI security week’ since the OpenAI/HuggingFace sandbox escape (Jul 26) — distillation restriction is now active policy, not just technical preference.

NVIDIA OpenShell: Open-Source Agent Sandbox with Hard OS Limits

  • NVIDIA has shipped OpenShell, an open-source agent sandbox that enforces hard OS-level runtime limits on local and open-weight AI agents — not just prompt-level rules, but actual operating-system constraints on what agents can access and execute.
  • Over 100 firms have joined the OpenShell safety stack at launch; OpenAI is notably absent — a pointed contrast given this week’s distillation enforcement action and Anthropic’s GLM-5.3 findings.
  • OpenShell extends NVIDIA’s OpenShell preview from NVIDIA GTC Taipei (Jun 7) into a full open-source release; together with Project Glasswing, it forms the emerging architecture for runtime AI containment at the infrastructure layer.

Anthropic: "What Work Can Robots Do?" Economics Report

  • Anthropic’s economics team published a detailed labour automation report examining which categories of work AI can perform today and the near-term trajectory — covering cognitive, physical, and hybrid task types.
  • The report provides one of the most granular breakdowns of AI automation readiness to date, situating Anthropic as both a contributor to labour displacement and a researcher of its effects — echoing the framing of the $200M Economic Futures Research Fund (Jul 26).
  • Key implication: the analysis explicitly separates ’technically feasible’ from ’economically deployed’ automation — cautioning against both over-optimism and dismissal, and pointing to transition costs as the primary policy variable.

Products & Hardware

2 stories

AMD EPYC 9006 Venice: 256 Cores, 91% RTX 5090 Bandwidth

  • AMD’s EPYC 9006 (Zen 6 Venice) tops out at 256 cores with 16-channel DDR5-12800, delivering 91% of an RTX 5090’s memory bandwidth — the key metric for local LLM inference throughput — at a fraction of the GPU cost.
  • Full pricing has been confirmed from $700 to $14,904 for the 256-core flagship, making high-bandwidth CPU inference newly competitive for large-model serving workloads that previously required expensive GPU memory.
  • The announcement reinforces the Extreme Local Inference thread — from ternary models to Optane RAM rigs — but shifts the lens from hobbyist tricks to production-grade CPU inference at data center scale.

Research & Resources

3 stories

Strands Decider 2B: Open-Source Agent Routing Model

  • Strands Agents (AWS) released Strands Decider 2B — an Apache 2.0, 2B-parameter open-source decision model for agent routing, tool selection, guardrails, memory, and policy classification, with full training data and scripts on GitHub and weights on HuggingFace.
  • Architecture replaces the standard LM head with a pointer head (1M params) and adds a rank-16 LoRA adapter on a Qwen3.5-2B torso, enabling parallel scoring in a single pass and producing calibrated confidence scores unavailable from LLM APIs; latency is ~115ms on RTX 3090 and ~153ms on M3 MacBook.
  • The key demo — before-tool-call intervention that checks whether tool arguments are grounded and whether the call is premature — directly prevents hallucinated tool invocations; Decider ranks 3rd of 33 in its 2B class on JevBench, following the TypeSafe Jev ‘System One Models’ pioneer (Sep 20) it cites.

Hugging Face Open-Sources 200+ Fastest WebGPU ML Kernels

  • Hugging Face open-sourced over 200 optimised WebGPU kernels for ML operations that run entirely in-browser with no server required, covering attention, matrix multiplication, activation functions, and quantization — all being upstreamed to Transformers.js, ONNX Runtime Web, and LiteRT.js.
  • The release benchmarks as the world’s fastest WebGPU ML kernels, enabling full local AI inference in any modern browser without installation — continuing the Extreme Local Inference thread from Bonsai Image 4B (WebGPU diffusion) and Ternary Bonsai 2 (27B browser-runnable model).
  • The ecosystem-first approach — upstreaming to three major inference runtimes simultaneously — positions this as infrastructure rather than a standalone demo, and lowers the barrier for privacy-preserving, offline-capable web AI applications.