The official AWS blog demonstrates deploying a three-agent music production pipeline on Amazon Bedrock AgentCore Runtime Instances. The composition agent uses Claude Sonnet 4.6 to generate a music brief and runs ACE-Step on the instance's NVIDIA L4 to render audio. The delivery agent reads .wav files from the shared volume for measurement and DSP processing, while the compliance agent independently remeasures them and checks harmonic similarity against a music catalogue.
CoreWeave announced that NVIDIA Vera Rubin NVL72 systems with Spectrum-X 102.4T Ethernet were available on CoreWeave Cloud, with Cognition becoming the first customer to run them in production.
GPT-6.1 Sol is now generally available on Amazon Bedrock, offering stronger reasoning for agentic coding, computer use and everyday professional workloads. OpenAI says it matches GPT-6 Astra on DeepSWE v1.1 at roughly one-fifth the cost per task, and exceeds GPT-6 Sol’s best result by 6.4 percentage points.
OpenAI announced extensions to Codex and its API at DevDay 2026 in San Francisco. Codex adds reusable cloud environments, repository security scanning through Codex Security Cloud, a desktop code-review view and a redesigned CLI, while the Agents API gains Computer Use, Tool Search and Context Compaction.
AWS announced Claude Sonnet 5.5 availability on Amazon Bedrock and Claude Platform on AWS, positioning it as a more efficient Sonnet for coding and knowledge work, with lower cost per task and faster speeds for most tasks.
Microsoft Research systematically measured mobile manipulation robot workloads and found that offloading physical AI inference from onboard GPUs to edge or cloud GPUs can improve task success, support larger models and extend battery life.
NVIDIA released Isaac ROS 5.0 at ROSCon in Toronto, Canada. The collection of ROS-based GPU-accelerated packages adds agentic workflows and platform support to help developers build, customise and deploy robot applications faster.
TypeSafe AI founder and CEO Diogo Almeida introduced Jev on the Latent Space podcast, describing it as a System One large model designed for consumption by software.
GitHub used the Copilot app and Copilot CLI to completely rewrite the Copilot agent runtime from TypeScript into over 800,000 lines of production Rust. AI agents wrote most of the code, delivered incrementally across 128 PRs, improving performance by several orders of magnitude.
NVIDIA submitted preview results for Vera Rubin NVL72 to MLPerf Inference v6.1 for the first time, achieving throughput up to 3.7 times that of GB300 NVL72 on Qwen3-VL and up to 2.5 times on DeepSeek-R1.
NVIDIA introduced DSX at AI Infra Summit to improve AI-factory token throughput within a fixed power budget. Lambda validated DSX MaxLPS on a five-rack, 19-node HGX B200 cluster, running 19 nodes within the same budget as 16 at full power. Cluster token throughput rose 24%, from around four million to five million tokens per second, with performance per watt up 23%.
NVIDIA announced several advances for Vera Rubin and the DSX platform at AI Infra Summit, focusing on energy-efficiency improvements in token throughput per megawatt.
GitHub released a research preview of Project HydraFusion, using runtime multi-model orchestration to deliver frontier-level coding through Copilot. Users can enable it via /experimental in GitHub Copilot CLI, paying each model's standard rates for tokens actually consumed.
Hugging Face’s WebAI team released @huggingface/kernels, a lightweight library for loading and running optimised WebGPU kernels from the Hugging Face Hub. An initial collection of 207 kernels is also available as separate repositories under Apache-2.0.
Liquid AI released DSpark draft model checkpoints for LFM2.5-1.2B-Instruct, LFM2.5-2.6B and LFM2.5-8B-A1B. Speculative decoding accelerates decoding without changing output quality, increasing throughput by up to 3.18 times on GPUs and 2.87 times on devices.
A Hugging Face blog tutorial demonstrates a streaming data loop for Strands Robots. The same Robot() object records demonstrations, syncs them to a Storage Bucket, streams training data from the Hub and deploys the checkpoint back to hardware, keeping the LeRobot disk format unchanged throughout.
NVIDIA released Magpie TTS Multilingual, a 364M-parameter open-weight speech synthesis model supporting English, Spanish, French, German, Italian, Vietnamese, Chinese, Hindi and Japanese, plus newly added Modern Standard Arabic, Korean and Brazilian Portuguese, for 12 languages in total.
Hugging Face announced that vLLM’s transformers modelling backend now matches or exceeds the throughput of vLLM’s handwritten native implementations across several LLM architectures. It uses torch.fx for static graph analysis and ast to rewrite source code, dynamically applying inference-related layer fusion at runtime to match custom-code performance.
At Build 2026, Microsoft announced Foundry Managed Compute and a Hugging Face model collection, with weekly updates to open-weight models and one-click deployment to Foundry's managed GPU platform.
Hugging Face and SkyPilot released an integration allowing a Hugging Face Bucket or any model, dataset or Space repository to be mounted into SkyPilot tasks with an hf:// URL and an existing HF_TOKEN, running across more than 20 clouds, Kubernetes, Slurm and local environments.