NVIDIA AI 工厂如何最大化投资回报:高产出、长寿命、通用性
NVIDIA 提出 AI 工厂投资回报取决于三个因素:每兆瓦产出能力、硬件使用寿命和 token 需求,并以 Vera Rubin NVL72 为例称其每兆瓦吞吐量比 GB300 NVL72 高 30 倍以上,在 DeepSeek V4 Pro 上每百万 token 成本最多低 45 倍。
NVIDIA 提出 AI 工厂投资回报取决于三个因素:每兆瓦产出能力、硬件使用寿命和 token 需求,并以 Vera Rubin NVL72 为例称其每兆瓦吞吐量比 GB300 NVL72 高 30 倍以上,在 DeepSeek V4 Pro 上每百万 token 成本最多低 45 倍。
Amazon Bedrock Knowledge Bases adds the AgenticRetrieveStream API, which breaks multi-part insurance claims questions into subqueries and retrieves evidence over multiple rounds until there is enough to generate a cited answer.
Microsoft Research has developed an end-to-end machine learning pipeline that uses solar wind observations at L1, forecasts of the AE and Dst indices, and local geological conductivity data to estimate local space weather risks 30 to 60 minutes ahead for 66,935 substations in the contiguous United States.
The official AWS blog demonstrates deploying a three-agent music production pipeline on Amazon Bedrock AgentCore Runtime Instances. The composition agent uses Claude Sonnet 4.6 to generate a music brief and runs ACE-Step on the instance's NVIDIA L4 to render audio. The delivery agent reads .wav files from the shared volume for measurement and DSP processing, while the compliance agent independently remeasures them and checks harmonic similarity against a music catalogue.
CoreWeave announced that NVIDIA Vera Rubin NVL72 systems with Spectrum-X 102.4T Ethernet were available on CoreWeave Cloud, with Cognition becoming the first customer to run them in production.
Amazon Bedrock has introduced Anthropic's Claude Opus 5, Claude Sonnet 5 and Claude Haiku 4.5 in India, using geographic cross-region inference to process data within the country.
Amazon Bedrock now supports in-region inference for Anthropic Claude models in Seoul and Singapore. Seoul supports Claude Opus 5 and Claude Sonnet 5, while Singapore supports Claude Sonnet 5, all invoked through the bedrock-runtime endpoint.
GPT-6.1 Sol is now generally available on Amazon Bedrock, offering stronger reasoning for agentic coding, computer use and everyday professional workloads. OpenAI says it matches GPT-6 Astra on DeepSWE v1.1 at roughly one-fifth the cost per task, and exceeds GPT-6 Sol’s best result by 6.4 percentage points.
A contract intelligence platform built on Amazon Quick and Amazon Bedrock AgentCore uses AI agents to extract eight key fields from contract PDFs and independently verify them. Amazon Textract adjudicates disagreements over signature recognition, and results are stored in Amazon Aurora PostgreSQL.
Condé Nast partnered with the AWS Generative AI Innovation Center (GenAIIC) to build multimodal video retrieval on Amazon Bedrock and Amazon OpenSearch Service, using the TwelveLabs Marengo embedding model to jointly encode visuals, audio and transcripts.
AWS announced Claude Sonnet 5.5 availability on Amazon Bedrock and Claude Platform on AWS, positioning it as a more efficient Sonnet for coding and knowledge work, with lower cost per task and faster speeds for most tasks.
AWS released a vLLM-Omni DLC for deploying Qwen3-TTS on SageMaker AI. A single persistent bidirectional connection streams text in and audio out, with playback beginning before generation finishes.
AWS published a tutorial deploying two endpoints from the same vLLM-Omni DLC image on SageMaker AI: a real-time endpoint running FLUX.2-klein-4B for images and an asynchronous endpoint running Wan2.1-VACE-1.3B for image-conditioned video. A text prompt first generates a PNG, which is sent with a motion prompt through Amazon S3 to the video endpoint; the resulting MP4 is retrieved from S3. The example includes a command-line workflow and an optional Streamlit interface.
Mistral opened a hub in Munich with a research team focused on Physics AI and industrial AI, and plans to build one gigawatt of European computing capacity by 2030. It will work with BMW on crash simulation and engineering AI, Siemens Energy on industrial AI applications, and the Technical University of Munich (TUM) on automotive aerodynamics digital twins using wind-tunnel facilities.
An agent-driven synthetic monitoring solution built on Amazon Nova Act and Amazon Bedrock AgentCore replaces Selenium and Playwright DOM-selector scripts with natural language actions. A multimodal LLM reads UI screenshots to execute critical flows such as login, ordering and checkout.
AWS released a solution for automating Amazon Textract adapter lifecycle management across accounts. By externalising adapter IDs to AWS Systems Manager Parameter Store, production adapter references can be updated without downtime or application redeployment, reducing coordination that previously took hours or days to seconds.
AWS presents an architecture combining EFA and DeepEP on Amazon EKS to scale MoE reinforcement learning training, increasing throughput by 40%.
Using the open-source SkyRL framework on Amazon SageMaker HyperPod, multimodal reinforcement learning post-training with GRPO increased the Qwen3-VL-8B vision-language model’s maze-solving success rate from 43.75% to over 95% on a fixed evaluation set of 64 mazes.
NarrateAI uses five techniques on Amazon Bedrock to achieve around 99% numerical accuracy with real-time streaming responses for over 4,000 AWS executives: adaptive pipeline orchestration, cross-account multi-model failover, real-time streaming evaluation, a composite evaluation framework and data-accuracy validation. Around 90% of queries take a single-pass fast path; only about 10% use parallel batch processing.
Qwen3-TTS-12Hz-1.7B-Base, open-sourced by Alibaba Cloud's Qwen team, is available on Amazon SageMaker JumpStart. It deploys to fully managed real-time inference endpoints and clones voices from a few seconds of reference audio and matching text without retraining.
Datacor embedded Amazon Quick Sight in TrackAbout to provide industrial gas and welding distributors with interactive dashboards and natural-language search, allowing business users to query rental data without raising IT tickets.
Amazon SageMaker HyperPod with Cloud Native Qumulo (CNQ) and Cloud Data Fabric (CDF) lets training clusters read remote datasets directly across AWS Regions without copying data or changing code.
The GitHub Copilot app introduced canvas, full-stack mini-apps running inside the app without browser chrome. They communicate bidirectionally with Copilot agents and can call third-party APIs or execute code locally.
AWS released the WhisperX Deep Learning Container, adding batch inference, word-level timestamps through wav2vec2 forced alignment and speaker diarisation to OpenAI Whisper. It can be deployed to Amazon SageMaker AI real-time or asynchronous endpoints without building a custom image.
Enterprises can build multi-account AI agent architectures with Amazon Bedrock AgentCore Gateway and MCP, keeping each business unit's data in its own AWS account and releasing it only as needed during queries.
Aderant used Amazon Nova Lite through Amazon Bedrock to build an intelligent ticket analyser for its 38-person SierraOps team, automating ticket context gathering, classification and routing.
Liquid AI released experimental draft model LFM2.5-VL-DSpark for vision-language model LFM2.5-VL-3B. Speculative decoding speeds it up without changing output quality: decoding improves by up to 3.13x on devices and 2.66x on H100, with end-to-end gains of up to 2.62x and 2.27x respectively.
Remedy Entertainment's CONTROL Resonant joined GeForce NOW on launch day. Ultimate members can stream at GeForce RTX 5080-class performance with NVIDIA DLSS 4, ray tracing, NVIDIA Reflex and up to 5K HDR, without a 100 GB local installation.
Dutch retailer HEMA built internal AI assistant HAL with MCP and Amazon Bedrock AgentCore, consolidating knowledge scattered across portals and wikis and delivering it directly into everyday tools such as Kiro and Claude.
GitHub rebuilt its Copilot app's pull request view to handle rendering pressure from huge diffs and comments, testing an open-source PR with 2,200 files, over 1 million changed lines and more than 400 inline comments. It separates document height into deterministic code geometry and dynamic block geometry: code line heights are computed precisely in advance, while comments and other dynamic blocks use bounded heights and deferred measurements, with corrections anchored to the user's current position.
Microsoft Research systematically measured mobile manipulation robot workloads and found that offloading physical AI inference from onboard GPUs to edge or cloud GPUs can improve task success, support larger models and extend battery life.
Airbnb is expanding use of GPT-6 Astra and OpenAI frontier models to help engineering teams investigate bugs, design systems and deliver faster.
NVIDIA released Isaac ROS 5.0 at ROSCon in Toronto, Canada. The collection of ROS-based GPU-accelerated packages adds agentic workflows and platform support to help developers build, customise and deploy robot applications faster.
NVIDIA launched DSX Ready certification to validate partner products against its DSX AI factory reference design, initially covering battery energy storage systems (BESS) and coolant distribution units (CDU).
NVIDIA explains that deploying physical AI at scale requires safety covering hardware, software, AI behaviour and operating environments, and introduces its full-stack safety system NVIDIA Halos.
NVIDIA held an AI ecosystem event at the Grand Egyptian Museum, with Egyptian Deep Learning Institute participation growing over tenfold in a year. Hassan Allam received a data centre licence from Egypt's National Telecom Regulatory Authority and is developing a new centre with A15 with expected investment of $400 million. NVIDIA says four African AI factories have been announced or launched, with another 656 MW under construction.
NVIDIA argues AI security should be treated as an engineering problem, with explicit requirements, executable controls, named owners and evidence of effective protection. Its open-source NVIDIA OpenShell enforces policies outside agent reasoning and provides sandboxed execution. Cisco DefenseClaw adds governance, while JFrog integrates OpenShell to scan and validate agent skills.
NVIDIA highlighted five AI clean-energy companies at New York Climate Week. ThinkLabs AI's NVIDIA CUDA-based digital twins and agents help Southern California Edison cut grid connection application assessments from 30–45 days to two minutes.
Hugging Face released tokenizers v1, preserving token IDs, APIs, vocabularies and merge ranks from v0.23 while increasing speed by up to tens of times.
Google Research released MilleMiglia, a C++ generator of realistic, privacy-preserving benchmark instances for middle-mile logistics delivery. Source and documentation are public on GitHub.