Introducing the Agents API
OpenAI released the Agents API, a managed service powered by the Codex harness for building and launching cloud agents. It supports orchestration, long-running sessions and tool calls.
New features, redesigns and business moves in AI products and apps: which got better, pricier or finally usable.
OpenAI released the Agents API, a managed service powered by the Codex harness for building and launching cloud agents. It supports orchestration, long-running sessions and tool calls.
Google DeepMind launched AlphaGenome Atlas, a platform containing predicted effects for 9 billion single-nucleotide variants across the human genome. At 1 PB, it is over 30 times the size of the AlphaFold Database.
OpenAI released ChatGPT Images 2.5, turning ideas, sketches and reference photos into more personalised, polished images. The company says the results better reflect users’ ideas.
GitHub released a research preview of Project HydraFusion, using runtime multi-model orchestration to deliver frontier-level coding through Copilot. Users can enable it via /experimental in GitHub Copilot CLI, paying each model's standard rates for tokens actually consumed.
Hugging Face released funes, an open-source persistent memory layer for Claude Code, Codex, pi, Hermes and other coding agents. It indexes existing local sessions and performs embedding and reranking locally by default.
Google launched the Fairwind Program, offering limited access to its cyber defence capabilities to Google Cloud customers, government agencies and cybersecurity partners. The initial offering includes Gemini 3.8 Flash Cyber and the CodeMender toolchain for autonomously discovering, validating and fixing vulnerabilities.
Google DeepMind introduced agentic video understanding for Gemini 3.7 Flash, 3.6 Flash and 3.5 Flash-Lite, reducing video analysis token use by up to 88% and costs by up to 66%, while improving accuracy by up to 7%.
Hugging Face’s WebAI team released @huggingface/kernels, a lightweight library for loading and running optimised WebGPU kernels from the Hugging Face Hub. An initial collection of 207 kernels is also available as separate repositories under Apache-2.0.
Voice Arena and Hugging Face added Monsoon en-IN and Monsoon hi-IN evaluation datasets to Open ASR Leaderboard, making Hindi its first Indian language.
Hugging Face introduced gr.Workflow in Gradio to describe multi-step AI pipelines as graphs of typed nodes. Gradio provides a draggable canvas where each node can run and every intermediate result is visible.
Liquid AI released DSpark draft model checkpoints for LFM2.5-1.2B-Instruct, LFM2.5-2.6B and LFM2.5-8B-A1B. Speculative decoding accelerates decoding without changing output quality, increasing throughput by up to 3.18 times on GPUs and 2.87 times on devices.
Mistral released Agentic Search, a retrieval layer that lets models find, examine and verify information in multi-step retrieval loops. It is provided through Mistral Search Toolkit and built into Libraries in Studio and Vibe.
Sentence Transformers v6.0 adds a fourth model type, MultiVectorEncoder, for ColBERT-style late-interaction retrieval. It directly loads PyLate, Stanford-NLP ColBERT checkpoints and visual document retrieval models from colpali-engine.
NVIDIA released Magpie TTS Multilingual, a 364M-parameter open-weight speech synthesis model supporting English, Spanish, French, German, Italian, Vietnamese, Chinese, Hindi and Japanese, plus newly added Modern Standard Arabic, Korean and Brazilian Portuguese, for 12 languages in total.
Microsoft Research released the open-source Orchard framework, centred on the Kubernetes environment service Orchard Env. It reuses environments, data pipelines and evaluation workflows across tasks and supports training agents directly within real deployment frameworks such as Codex, OpenClaw and ZeroClaw.
Hugging Face introduced Nunchaku Lite into Diffusers, allowing Nunchaku quantised checkpoints to load directly with from_pretrained(), without custom pipelines or local CUDA compilation.
Hugging Face released Grabette, an open-source system recording manipulation demonstrations with a handheld gripper and two cameras, producing robot-ready datasets without robots or teleoperation equipment. The handheld hardware costs about €490 in materials, with the accompanying motorised Gripette gripper around €120. Hardware CAD, Raspberry Pi collection software and browser-based processing are all open-source.
Hugging Face announced that vLLM’s transformers modelling backend now matches or exceeds the throughput of vLLM’s handwritten native implementations across several LLM architectures. It uses torch.fx for static graph analysis and ast to rewrite source code, dynamically applying inference-related layer fusion at runtime to match custom-code performance.
At Build 2026, Microsoft announced Foundry Managed Compute and a Hugging Face model collection, with weekly updates to open-weight models and one-click deployment to Foundry's managed GPU platform.
Hugging Face released LeRobot v0.6.0, introducing world model policies VLA-JEPA, FastWAM and LingBot-VA, new VLAs GR00T N1.7, MolmoAct2, EO-1, EVO1 and Multitask DiT, and a unified reward model API with Robometer and TOPReward for the robotics learning loop.