Sep 30, 2026 | Gemini 4 Argon delivers frontier performance in complex workflows across real-world software engineering, enterprise knowledge work like legal and finance, and cybersecurity defense. Today, we’re announcing our new frontier model, Gemini 4 Argon, which is rolling out to a set of trusted cyber defenders through our Fairwind Program.
GPT-6.1 Sol is now generally available on Amazon Bedrock, offering stronger reasoning for agentic coding, computer use and everyday professional workloads. OpenAI says it matches GPT-6 Astra on DeepSWE v1.1 at roughly one-fifth the cost per task, and exceeds GPT-6 Sol’s best result by 6.4 percentage points.
OpenAI's new GPT-6.1 Sol is supposed to come close to GPT-6 Astra for a fraction of the cost. The planned flagship model Astra is staying under wraps because of safety problems. OpenAI is betting on cost efficiency over peak performance.
OpenAI unveiled GPT-6.1 Sol at DevDay, claiming near-GPT-6 Astra intelligence in agentic coding, computer use and professional work, with standard input and output token prices just one-fifth as much.
OpenAI announced extensions to Codex and its API at DevDay 2026 in San Francisco. Codex adds reusable cloud environments, repository security scanning through Codex Security Cloud, a desktop code-review view and a redesigned CLI, while the Agents API gains Computer Use, Tool Search and Context Compaction.
OpenAI has released GPT-6.1 Sol, which it says offers intelligence close to Astra and is aimed at coding, computer use and professional work. The model's API input and output token prices are one fifth of Astra's standard price.
OpenAI has published a DevDay 2026 recap rounding up more than 20 announcements spanning GPT-6 Astra, ChatGPT, Codex, the API, safety and new tools for developers. The original gives only an overview and does not detail each update.
Latent Space's AINews rounds up developments from 9/24–9/25, noting positive community feedback since Opus 5.5 launched this week, particularly on generating explainer videos with code.
xAI’s Grok 4.7 is available on Amazon Bedrock with a 500K-token context window, four configurable reasoning levels—low, medium, high and xhigh—and support for the Responses, Chat Completions and Converse APIs.
AWS announced Claude Sonnet 5.5 availability on Amazon Bedrock and Claude Platform on AWS, positioning it as a more efficient Sonnet for coding and knowledge work, with lower cost per task and faster speeds for most tasks.
Anthropic released Claude Sonnet 5.5, the second model in the Claude 5.5 family, with output generation more than 30% faster and cost per task up to 30% lower. It approaches Opus 5.5 on several benchmarks.
Anthropic released Claude Opus 5.5, calling it the first model in the new Claude 5.5 family, matching Claude Fable 5.1 on most tasks at 40% lower running costs than Opus 5. It becomes the default model in Claude Code and the Claude app.
GitHub used the Copilot app and Copilot CLI to completely rewrite the Copilot agent runtime from TypeScript into over 800,000 lines of production Rust. AI agents wrote most of the code, delivered incrementally across 128 PRs, improving performance by several orders of magnitude.
Mistral helped a European energy operator migrate 40,000 lines of Fortran 77 to C++, targeting a physics-heavy reservoir simulator without a test suite or centralised documentation. The team first built a numerical-alignment testing framework, used Skill.md to guide agents in exporting state snapshots and validating migrated modules, then launched over a hundred agents with Vibe CLI to analyse call trees and used Mistral OCR to organise scattered documents.
GitHub released a research preview of Project HydraFusion, using runtime multi-model orchestration to deliver frontier-level coding through Copilot. Users can enable it via /experimental in GitHub Copilot CLI, paying each model's standard rates for tokens actually consumed.
Google DeepMind released Gemini 3.8 Flash and Gemini 3.8 Flash Cyber. The former targets long-horizon coding and autonomous agents, priced like 3.7 Flash at $0.75 per million input tokens and $3.75 per million output tokens.
Multiverse Computing published a paper proposing Quantization-Aware Healing (QAH). After compressing GPT-OSS 120B to 60B parameters and quantising it to MXFP4, the method distils directly from the original uncompressed model rather than a reconstructed bfloat16 checkpoint.
Google DeepMind released Gemini 3.7 Flash, positioning it as its strongest workhorse model for coding and agents, just three weeks after Gemini 3.6 Flash.
Mistral released an AI stack for industrial engineering at AI Now Summit 2026, partnering with Airbus, BMW and ASML to optimise design, simulation and production while retaining control over proprietary data and IP.