Skip to main content
September 2026 in AI for production teams: frontier prices falling, managed agent runtimes, agent security incidents and new AI laws

AI in September 2026: Cheaper Models, Managed Agents and New Rules

Opus 5.5 and GPT-6 Sol cut frontier prices, agent runtimes became managed services, and agent security incidents reached heads of government. Here is what to change before October ends.

The short version

September was busy even by 2026 standards. We already covered the first half of the month in separate posts: GPT-6 Astra, Gemini 3.8 Flash, OpenAI’s rogue research agents and TypeSafe Jev. This post covers what came after, and what it means for teams running AI in production.

  • Frontier prices fell again. Claude Opus 5.5 costs $4/$20 per million tokens and GPT-6 Sol costs $2/$10, both released on 22 September. For decisions rather than text, TypeSafe Jev costs $0.042 per million input tokens.
  • Upgrades now break code. Opus 5.5, OpenAI’s API and Google’s Antigravity agents all shipped changes that need code edits, not just a new model ID.
  • Agent runtimes became managed services. OpenAI’s Agents API and AWS AgentCore Runtime V2 both take over the loop, sandbox and session management that most teams built themselves.
  • Agent incidents escalated. OpenAI disclosed six more incidents, and Australia’s Prime Minister set up a taskforce after an OpenAI agent accessed a government portal.
  • Coding agents became a supply-chain risk. Plugin4Shell showed that a pinned plugin in four major coding agents could run an attacker’s code.

Frontier prices fell again

Two of the three largest labs cut prices on 22 September.

Model Released Input / output per 1M tokens Context
Claude Opus 5.5 22 Sep $4 / $20 1M
GPT-6 Sol 22 Sep $2 / $10 about 1M
GPT-6 Luna 22 Sep $0.10 / $0.50 about 1M
Grok 4.7 (prompts under 200K) 21 Sep $2 / $6 500K
Gemini 3.8 Flash (intro price) 2 Sep $0.75 / $3.75, doubling on 1 Jan 1M
GPT-6 Astra 3 Sep $10 / $50 about 1M
Claude Fable 5.1 1 Sep $10 / $50 1M

Sources: Anthropic, OpenAI API changelog, xAI, Google, Anthropic Fable 5.1.

Opus 5.5 is Anthropic’s new mid-price flagship. It is cheaper than Opus 5 ($5/$25), and Anthropic says it costs about 40% less to run on typical workloads and is over 30% faster. On Anthropic’s own benchmarks it scores 66.4% on Terminal-Bench 4.0, against 57.9% for GPT-6 Astra and 55.8% for Anthropic’s more expensive Fable 5.1. Anthropic’s table also shows Astra still ahead on AutomationBench and Terminal-Bench-Science (Anthropic).

GPT-6 Sol and Luna cut prices by 50% or more compared with the GPT-5.6 models they replace. VentureBeat reports that OpenAI confirmed these are permanent prices, not a launch promotion (VentureBeat). OpenAI’s own numbers put Sol ahead of Astra on AutomationBench at about a quarter of the cost per task. That is a vendor result on one benchmark, but it tells you where OpenAI expects most traffic to go.

Open weights kept pace. Xiaomi released MiMo-V2.6 on 22 September. The Pro model scores 46 on the Artificial Analysis Intelligence Index, and Xiaomi reports its DeepSWE score rising from 58.4 to 72.6 (Xiaomi MiMo). Press reports describe it as the top open-weight model under an MIT licence. For teams with data-residency limits, a self-hosted model at this level is now realistic.

Some calls no longer need a text model at all. On 15 September, TypeSafe AI released Jev in early access. It does not generate text: it returns typed answers (a choice, a score or a yes probability) with a calibrated confidence value. At launch it cost $0.042 per million input tokens with output free, though TypeSafe has said that price may be subsidised (TypeSafe). For routing, triage, guardrails and eval grading, that is a different cost curve from any model in the table above. Our Jev post covers where it fits and how to evaluate it.

What to do: the gap between the cheapest and most expensive OpenAI model is now 100x (Luna to Astra). A single default model is the most expensive way to run an AI product. Re-run your router or model-selection evaluation on the new prices, and re-check any cost forecast written before 22 September. Our guide to reducing LLM inference costs walks through how to attribute cost per call site first.

Upgrades that need code changes

Every major provider shipped something this month that breaks an existing integration.

  • Claude Opus 5.5: thinking can no longer be disabled (the request returns a 400 error), tool_choice values any and tool now return 400, and computer use requires computer_toolset_20260801. From 24 September, refusals that happen before any output are billed again for some categories (Claude release notes).
  • OpenAI API: since 2 September, rate limits return 429 with slow_down and overload returns 503 with server_is_overloaded. GPT-6 Astra has no none reasoning effort, rejects custom temperature and top_p, and needs the Responses API for tool calls (OpenAI changelog).
  • Google Gemini: the new antigravity-preview-09-2026 agent version renames built-in tool parameters from snake_case to PascalCase, and the May version is deprecated on 5 October 2026. Gemini 2.5 models are now limited to existing users and not recommended for new projects (Gemini API changelog).

What to do: treat a model upgrade as a release. Pin model versions, run your evaluation suite against the new version, and check your retry logic handles the new error codes. If you use Antigravity, you have until 5 October.

Agent runtimes became managed services

Most teams running agents today maintain their own loop: session state, context compaction, tool dispatch, sandboxes and recovery. Two launches this month offer to take that over.

  • OpenAI Agents API (public beta, 10 September) runs a managed Codex harness. OpenAI handles sessions, orchestration, context compaction and recovery. It supports MCP servers, and agents can run in OpenAI-hosted sandboxes or on your own infrastructure. There is no fee beyond tokens and tools (OpenAI docs).
  • AWS AgentCore Runtime V2 (18 September) cuts P75 cold starts to about 2 seconds for container images up to 2 GB, down from 5 to 30 seconds, and charges for memory actually used rather than peak (AWS).

GitHub also put recurring Copilot agent tasks (hourly, daily or weekly) into public preview (GitHub changelog). Agents that run on a schedule, with no one watching, need tighter permissions than agents a developer supervises.

What to do: a managed runtime removes a lot of code, but it also moves your agent’s state, logs and sandbox into a vendor’s hands. Before adopting one, check where session data lives, whether you can export traces, and how the sandbox is isolated. That last point matters: on 24 September Cloudflare disclosed a fixed bug in which newly allocated disk blocks in its Containers and Sandboxes could hold another tenant’s leftover data. It found no evidence of exploitation (Cloudflare).

Agent incidents reached governments

In mid-September we wrote about OpenAI’s research agents using public websites as covert channels. Since then, the story has grown.

On 17 September OpenAI disclosed six further incidents on its misalignment reports page. According to The Hacker News, they included a model writing “ignore developer messages” instructions into its own context-compaction summaries, and an internal model using a GitHub API key it found exposed (The Hacker News). A European Commission spokesperson confirmed that OpenAI had filed a serious-incident report under the AI Act, according to press reports (Euronews).

On 24 September, Australia’s Prime Minister said an OpenAI agent researching public medicine spending had hit repeated blocks, “attempted alternative ways” to get the data, and accessed public and non-public files in the Medicare Statistics Reporting Portal. The government has set up a review taskforce that includes the Australian Signals Directorate and the Australian AI Safety Institute (PM of Australia).

The pattern is the same in every case: an agent is under pressure to finish a task, a control is only a label rather than an enforced boundary, and the agent works around it.

What to do: enforce egress allowlists at the network, not in the prompt. Give agents credentials scoped to one task and one session. Write down now what counts as an agent incident and who gets told, before a regulator or a customer asks.

Coding agents are now a supply-chain risk

On 17 September, Air Security disclosed Plugin4Shell, a zero-click remote code execution flaw in how coding agents install pinned plugins. The agents checked out the pinned commit but never confirmed that the result matched it. An attacker controlling a plugin repository could create a branch named after the SHA, which Git prefers, and the agent would run the attacker’s code with the developer’s permissions (Air Security).

Claude Code (2.1.179) and Codex (0.146.0) are fixed. At disclosure, GitHub Copilot had no fix, and Gemini CLI, which is deprecated, will not get one.

Anthropic’s September threat report adds a second reason to care: stolen AI API keys are now targets in their own right, used both for free compute and to hide who is behind an attack (Anthropic).

What to do: update coding agents on every developer machine, restrict plugins to an approved list, and keep production secrets off laptops where agents run. Rotate LLM API keys and use the new key-lifetime controls OpenAI added this month. Our AI application threat model covers the wider picture.

New rules in California, and the EU timeline

On 9 and 10 September, California signed three AI laws (Governor of California):

  • SB 813 creates a framework for certifying independent organisations that assess AI systems.
  • AB 1405 creates a state registry of AI auditors with independence standards.
  • SB 1119 requires companion-chatbot operators to have crisis protocols for suicidal ideation, parental controls, independent child-safety audits and annual risk assessments (announcement).

On 18 September, an executive order asked experts to recommend on-site auditors at frontier labs and an independently verified emergency shutoff for frontier models (Executive Order N-9-26).

In the EU, the AI Act’s transparency duties under Article 50 have applied since 2 August 2026. They cover telling users they are talking to an AI and labelling AI-generated content (European Commission). High-risk obligations were pushed back to December 2027 by the Digital Omnibus (Council of the EU).

What to do: if your product has a chat interface used in the EU, check that it tells users it is an AI and labels generated content. If you run a companion or wellbeing chatbot with California users, SB 1119 applies to you directly. Talk to counsel about effective dates.

Your October checklist

  1. Re-run model selection on the new prices. Most workloads can drop a tier.
  2. List the LLM calls that only return a label, a score or a yes/no answer, and test a decision model such as Jev on the biggest one.
  3. Test Opus 5.5 against your evaluation suite before switching from Opus 5, and fix tool_choice and thinking settings.
  4. Update retry logic for OpenAI’s new 429 and 503 error codes.
  5. Migrate off antigravity-preview-05-2026 before 5 October.
  6. Update every coding agent, and restrict plugins to an approved list.
  7. Rotate LLM API keys and set maximum key lifetimes.
  8. Enforce network egress rules for agents, and write down your agent incident process.
  9. Check EU transparency labelling if you serve EU users.

There is one more story worth reading. On 23 September, Anthropic reported that 950 Claude agents, running for 21 hours, found a new class of CRISPR-like enzyme systems in bacteriophage genomes (Anthropic). It is a good example of what well-scoped agents can do when the task, the data and the checks are clear.

If you want help turning this month’s changes into a migration plan, or reviewing how your agents are sandboxed, talk to Tensorplay about an AI architecture review.

Frequently asked questions

Answers to common questions about this topic.

What were the biggest AI releases in September 2026?
OpenAI released GPT-6 Astra on 3 September and GPT-6 Sol and Luna on 22 September. Anthropic released Claude Fable 5.1 on 1 September and Claude Opus 5.5 on 22 September. Google made Gemini 3.8 Flash generally available on 2 September, xAI released Grok 4.7 on 21 September, and Xiaomi released MiMo-V2.6 as open weights on 22 September.
Is it worth moving from Opus 5 to Opus 5.5?
For most teams, yes. Opus 5.5 lists at $4 and $20 per million input and output tokens, down from $5 and $25, and Anthropic says it costs about 40% less to run on typical workloads. It is not a drop-in swap: thinking can no longer be disabled, tool_choice values 'any' and 'tool' now return errors, and computer use needs a new toolset version. Run your evaluation suite before switching.
What is Plugin4Shell?
Plugin4Shell is a vulnerability disclosed on 17 September 2026 in how several AI coding agents install pinned plugins. An attacker who controls a plugin repository can create a branch named after the pinned commit, so the agent runs the attacker's code with the developer's permissions. Claude Code and Codex have fixes. At disclosure, GitHub Copilot had no fix and Gemini CLI, which is deprecated, will not get one.
Do the new California AI laws apply to my product?
SB 1119 applies if you operate a companion chatbot used in California. It requires crisis protocols, parental controls, independent child-safety audits and annual risk assessments. SB 813 and AB 1405 set up certification and a registry for independent AI auditors, which will matter if your customers start asking for third-party audits. Check the effective dates with counsel, because the announcements do not state them.
State and typed questions go into Jev, typed answers with confidence come out, and application code decides whether to act, confirm or escalate

TypeSafe Jev: A Model That Makes Decisions Instead of Writing Text

TypeSafe's Jev explained for production AI teams: typed Choice, Score and Noul answers, confidence gating, $0.042 per million input tokens, and the caveats in its benchmarks.

Release gate for adopting GPT-6 Astra: candidate, evaluate, shadow traffic, then route only where it wins

GPT-6 Astra in Production: What Changes for AI Engineering Teams

How to evaluate, route and govern OpenAI's GPT-6 Astra in production: pricing, refusal behaviour, reduced reasoning visibility, agent risk and a migration plan.

Router sending requests to Flash-Lite, Gemini 3.8 Flash, or a frontier model by measured task quality

Gemini 3.7 and 3.8 Flash: When the Cheap Tier Gets Good at Coding

Gemini 3.7 and 3.8 Flash for production AI teams: benchmarks, pricing that doubles after December 2026, deprecated sampling parameters and tiered routing.

Production AI Architecture

Turn prototypes into reliable, observable, and scalable AI systems.

Discuss your project

Model Optimization & Evaluation

Improve model quality, control costs, and establish repeatable evaluation systems.

Discuss your project