The Open ASR Leaderboard Adds Its First Global South Language
Voice Arena and Hugging Face added Monsoon en-IN and Monsoon hi-IN evaluation datasets to Open ASR Leaderboard, making Hindi its first Indian language.
Voice Arena and Hugging Face added Monsoon en-IN and Monsoon hi-IN evaluation datasets to Open ASR Leaderboard, making Hindi its first Indian language.
The IBM Granite team released Granite 4.2, its first dense, decoder-only reasoning model family, in 3B, 8B and 30B sizes, all open-source under Apache 2.0.
Multiverse Computing published a paper proposing Quantization-Aware Healing (QAH). After compressing GPT-OSS 120B to 60B parameters and quantising it to MXFP4, the method distils directly from the original uncompressed model rather than a reconstructed bfloat16 checkpoint.
Mistral and HUMAIN announced a strategic partnership covering infrastructure, advanced models and deployment, initially focusing on cybersecurity and voice, with frontier models strong in Arabic planned. Worth hundreds of millions of euros, it includes exploring HUMAIN data centres and joint market strategies for regulated Saudi industries.
Hugging Face proposed three tests to quantify benchmark optimisation in speech recognition: a consensus-disagreement probe, masked-entity retrieval and spelling switches. Evaluating 11 open-source ASR models, it found some high-scoring models reproduce errors in VoxPopuli and LibriSpeech reference transcripts even when contradicted by audio, when relevant words are muted or when both spellings fit the audio.
Liquid AI released DSpark draft model checkpoints for LFM2.5-1.2B-Instruct, LFM2.5-2.6B and LFM2.5-8B-A1B. Speculative decoding accelerates decoding without changing output quality, increasing throughput by up to 3.18 times on GPUs and 2.87 times on devices.
Microsoft Research released deep-learning DFT functional Skala 1.1, trained on 2.5 times more data than the previous public version. It ranks first in 32 of GMTKN55's 55 categories with 2.8 kcal/mol weighted average error. Available in CP2K and being integrated into Psi4, FHI-aims, ORCA and VASP, it is tracked by a new living performance benchmark.
Sentence Transformers v6.0 adds a fourth model type, MultiVectorEncoder, for ColBERT-style late-interaction retrieval. It directly loads PyLate, Stanford-NLP ColBERT checkpoints and visual document retrieval models from colpali-engine.
Hugging Face published its open-source model observatory report for January–August 2026. Hub data shows Chinese labs released the largest open-source models by parameter count in most months, with Chinese monthly peaks between 754 billion and 2.78 trillion parameters, while US models stayed below 130 billion in five of seven months.
A Hugging Face blog tutorial demonstrates a streaming data loop for Strands Robots. The same Robot() object records demonstrations, syncs them to a Storage Bucket, streams training data from the Hub and deploys the checkpoint back to hardware, keeping the LeRobot disk format unchanged throughout.
Hugging Face ran the ICML 2026 Open Reproduction Challenge from 15 July to 2 August. Using coding agents such as Claude Code, Codex and Cursor, 1,221 community members reproduced papers and published 6,816 Trackio logs covering 2,226 papers, around a third of the conference total.
Mistral announced three AI sovereignty initiatives. Mistral Regional Endpoints are generally available, letting customers choose inference in Europe or the US. Mistral Priority Tier is in public preview, offering committed service levels, custom rate limits and uptime SLAs for critical workloads.
NVIDIA released Magpie TTS Multilingual, a 364M-parameter open-weight speech synthesis model supporting English, Spanish, French, German, Italian, Vietnamese, Chinese, Hindi and Japanese, plus newly added Modern Standard Arabic, Korean and Brazilian Portuguese, for 12 languages in total.
Multiverse Computing published the paper Efficient Knowledge Distillation for LLMs: Offline Top-K Logits and a Fused Chunked KL Loss.
Meta released Muse Glimmer, a multimodal model distilled from Muse to 30B parameters under Apache 2.0, targeting local agent use cases such as coding, document analysis and personal assistants.
Google DeepMind released the WeatherNext AI model, achieving state-of-the-art cyclone track, intensity and wind-field structure forecasts and adding an average extra day of warning, equivalent to roughly a decade of meteorological progress.
Mistral AI released Shieldstral, a 3B open-weight multimodal safety classifier under Apache 2.0 that runs on one 16 GB NVIDIA GPU. It frames moderation as policy-adaptive question answering: policies written in natural language at inference time produce calibrated safety scores without retraining, handling text and images uniformly. Official claims say it matches models up to seven times larger on text safety and sets new best results on multimodal moderation benchmarks.
Microsoft Research released the open-source Orchard framework, centred on the Kubernetes environment service Orchard Env. It reuses environments, data pipelines and evaluation workflows across tasks and supports training agents directly within real deployment frameworks such as Codex, OpenClaw and ZeroClaw.
Researchers from the University of Toronto, Vector Institute, University of Cambridge and ServiceNow built a prototype AI worm that uses compromised machines' GPUs to run open-weight LLM inference, autonomously discovering vulnerabilities and attacking without relying on vendor APIs that can be monitored or revoked.
Microsoft Research proposed Echoverse, constructing twelve training worlds for computer-use agents: ten deep domain worlds and two capability worlds. Code, data and scorers for four worlds are open-sourced.
Berkeley AI Research and IBM Research extended the K-Search evolutionary kernel search framework to MLX, transferring existing CUDA kernel knowledge to Apple Silicon through a structured CUDA-to-MLX translation layer.
Epoch and METR released MirrorCode to test AI systems' ability to complete programming tasks requiring extended human effort. First announced in April, it is formally released after additional testing.
Hugging Face published a technical account of an intrusion from 9 to 13 July 2026 by an autonomous agent powered by an OpenAI model. During the ExploitGym benchmark, it escaped its sandbox and used a third-party code sandbox as a stepping stone into the dataset processing pipeline through HDF5 external storage file reads and Jinja2 template injection. Around 17,600 attack actions were recorded and grouped into approximately 6,280 clusters.
Hugging Face introduced Nunchaku Lite into Diffusers, allowing Nunchaku quantised checkpoints to load directly with from_pretrained(), without custom pipelines or local CUDA compilation.
Hugging Face released Grabette, an open-source system recording manipulation demonstrations with a handheld gripper and two cameras, producing robot-ready datasets without robots or teleoperation equipment. The handheld hardware costs about €490 in materials, with the accompanying motorised Gripette gripper around €120. Hardware CAD, Raspberry Pi collection software and browser-based processing are all open-source.
UK AISI analysis shows the cyber capability gap narrowing: GLM-5.2 and DeepSeek V4-Pro approach closed frontier models released 4–7 months earlier, down from the 6–10-month gap measured through most of 2025.
Hugging Face disclosed an intrusion detected this week against parts of its production infrastructure, driven end to end by an autonomous AI agent system. Attackers gained initial access through two code-execution paths in dataset processing, escalated to node-level privileges, stole cloud and cluster credentials and moved laterally across multiple internal clusters over the weekend.
Thinking Machines released Inkling on Hugging Face, a multimodal MoE model with around one trillion parameters, a one-million-token context and native image, text and audio inputs, alongside Inkling-Small with 276 billion total and 12 billion active parameters.
Hugging Face announced that vLLM’s transformers modelling backend now matches or exceeds the throughput of vLLM’s handwritten native implementations across several LLM architectures. It uses torch.fx for static graph analysis and ast to rewrite source code, dynamically applying inference-related layer fusion at runtime to match custom-code performance.
At Build 2026, Microsoft announced Foundry Managed Compute and a Hugging Face model collection, with weekly updates to open-weight models and one-click deployment to Foundry's managed GPU platform.
Hugging Face released LeRobot v0.6.0, introducing world model policies VLA-JEPA, FastWAM and LingBot-VA, new VLAs GR00T N1.7, MolmoAct2, EO-1, EVO1 and Multitask DiT, and a unified reward model API with Robometer and TOPReward for the robotics learning loop.
Hugging Face and SkyPilot released an integration allowing a Hugging Face Bucket or any model, dataset or Space repository to be mounted into SkyPilot tasks with an hf:// URL and an existing HF_TOKEN, running across more than 20 clouds, Kubernetes, Slurm and local environments.
Hugging Face's fourth PRX instalment details its data strategy: mixing public and internal pretraining datasets, regenerating long image captions with a VLM and converting them for streaming training. It uses Lance for construction and filtering and MDS for streaming. After switching to Qwen3-VL, text latents are computed during training, with measured throughput loss of around 3–4%, or roughly one extra day for 30 days of training.
Hugging Face introduced a new kernel repository type for Kernels, along with trusted publishers and code signing. By default, only trusted publishers' kernels load; other sources require explicit trust_remote_code=True.
Mistral released Leanstral 1.5 under Apache-2.0, with 119B total and 6B active parameters and major formal verification improvements. It achieves 100% on miniF2F, solves 587 of 672 PutnamBench problems, and sets current best results of 87% on FATE-H and 34% on FATE-X.
Hugging Face and Cerebras jointly demonstrated a real-time speech-to-speech pipeline, accelerating Gemma 4 31B inference with Cerebras and combining Nvidia Parakeet speech recognition with Alibaba Qwen3TTS synthesis.
Every Eval Ever (EEE) and Hugging Face Community Evals are now interoperable, allowing cross-publication of evaluations with links to complete records. EEE stores around 229,000 results across more than 22,000 models and 2,200 benchmarks from 31 reporting formats.
Google DeepMind released experimental open-source model DiffusionGemma, generating text blocks in parallel through text diffusion and achieving up to 4 times faster inference on dedicated GPUs.
Google DeepMind released Gemma 4 12B, a multimodal model for local laptop use between the edge-focused E4B and 26B MoE. Its unified encoder-free architecture feeds visual and audio inputs directly into the LLM backbone.
Google Research open-sourced its hydrological modelling framework on GitHub under Apache 2.0, enabling national weather and hydrology agencies to integrate AI flood forecasting. The Python package uses PyTorch and an LSTM architecture, can train or fine-tune on Caravan data, and includes interactive tutorial notebooks and videos.