From Atari to EVE Online: Building on 15 Years of AI Research in Games
Google DeepMind partnered with Fenris Creations to explore AI-driven gameplay prototypes in the EVE Universe, home to EVE Online.
Google DeepMind partnered with Fenris Creations to explore AI-driven gameplay prototypes in the EVE Universe, home to EVE Online.
Google Research proposed Mobility-Embedded POIs (ME-POIs), combining anonymised mobility patterns with text descriptions to generate place embeddings encoding both identity and dynamic function.
Papers with Code uses hybrid retrieval, with Hugging Face Jobs generating embeddings in batches, Storage Buckets providing persistent transfer and Inference Endpoints supplying low-latency query embeddings. It covers over 110,000 papers from arXiv and Daily Papers.
Hugging Face proposed three tests to quantify benchmark optimisation in speech recognition: a consensus-disagreement probe, masked-entity retrieval and spelling switches. Evaluating 11 open-source ASR models, it found some high-scoring models reproduce errors in VoxPopuli and LibriSpeech reference transcripts even when contradicted by audio, when relevant words are muted or when both spellings fit the audio.
Liquid AI released DSpark draft model checkpoints for LFM2.5-1.2B-Instruct, LFM2.5-2.6B and LFM2.5-8B-A1B. Speculative decoding accelerates decoding without changing output quality, increasing throughput by up to 3.18 times on GPUs and 2.87 times on devices.
Microsoft Research released deep-learning DFT functional Skala 1.1, trained on 2.5 times more data than the previous public version. It ranks first in 32 of GMTKN55's 55 categories with 2.8 kcal/mol weighted average error. Available in CP2K and being integrated into Psi4, FHI-aims, ORCA and VASP, it is tracked by a new living performance benchmark.
Mistral released Agentic Search, a retrieval layer that lets models find, examine and verify information in multi-step retrieval loops. It is provided through Mistral Search Toolkit and built into Libraries in Studio and Vibe.
IBM Research published ALTK-Evolve research on the Hugging Face blog. Tests of eight models on 585 multi-step AppWorld tasks found agent memory is not an on/off switch but a dosage requiring model-specific calibration.
Sentence Transformers v6.0 adds a fourth model type, MultiVectorEncoder, for ColBERT-style late-interaction retrieval. It directly loads PyLate, Stanford-NLP ColBERT checkpoints and visual document retrieval models from colpali-engine.
A Hugging Face blog describes a constraint-aware GPU allocator compared with FIFO under identical hardware and workloads. Across five real contention scenarios, GPU utilisation rose from 52–85% to 72–88%, while priority-weighted output improved 24.6–105.1%, averaging 52%.
Import AI 469 introduces DiG-bench (Discovery in Games), with 70 games whose hidden rules and goals must be discovered through interaction. Opus 5 and Fable 5 with Claude Code performed best; only they completed some Tier 7 tasks, while humans achieved 100% completion.
Google Research proposed PhotoScan, a deep learning framework that estimates body fat percentage, A/G ratio and V/S ratio from ordinary 2D phone photos to predict insulin resistance.
Hugging Face published its open-source model observatory report for January–August 2026. Hub data shows Chinese labs released the largest open-source models by parameter count in most months, with Chinese monthly peaks between 754 billion and 2.78 trillion parameters, while US models stayed below 130 billion in five of seven months.
A Hugging Face blog tutorial demonstrates a streaming data loop for Strands Robots. The same Robot() object records demonstrations, syncs them to a Storage Bucket, streams training data from the Hub and deploys the checkpoint back to hardware, keeping the LeRobot disk format unchanged throughout.
Google DeepMind released Gemini 3.7 Flash, positioning it as its strongest workhorse model for coding and agents, just three weeks after Gemini 3.6 Flash.
Hugging Face ran the ICML 2026 Open Reproduction Challenge from 15 July to 2 August. Using coding agents such as Claude Code, Codex and Cursor, 1,221 community members reproduced papers and published 6,816 Trackio logs covering 2,226 papers, around a third of the conference total.
Hugging Face's OlmoEarth Studio now computes and exports embeddings. Users choose an area, period, encoder variant and resolution through UI or API and receive Cloud-Optimized GeoTIFF (COG) results.
Microsoft Research introduced MindTopo to evaluate multimodal models' reasoning and planning over connectivity, enclosure, order, separation and knotting. Current models perform much better at static-image recognition than interactive planning, with failures mainly in planning rather than perception and overall performance far below humans. Image and video generation helps only when preserving structural relations in a single frame; multi-step operations often change topology or violate constraints.
Google DeepMind released SL2T, a multilingual sign-language-to-text model, bringing the capability to consumer products for the first time. Gboard and Live Transcribe on Pixel 11 support ASL-to-English dictation, with more devices and languages to follow.
Google Research's knowledge profiling framework evaluated 13 LLMs on WikiProfile's 2,150 Wikipedia facts. Gemini-3-Pro and GPT-5 encoded 95–98% but failed direct recall of 26–34%, and still failed 11–12% with thinking enabled. It argues frontier factual errors arise more from knowledge accessibility than absence, shifting the bottleneck from acquisition to use.
Google Research released AMIE (Video), built on Gemini and Project Astra, for real-time video clinical consultations. It can perceive non-verbal cues and guide virtual physical examinations.
Microsoft Research released CARE-X, a unified chest X-ray vision-language research model for report generation and structured prediction. It rewards clinical correctness using multitask reinforcement learning (DAPO). Generation and dual-inference modes cover lesion presence and negation, localisation, multilabel classification, catheter and tube malposition detection, and localisation of 29 anatomical regions.
Hugging Face's ALTK-Evolve and ACE both provide agent memory without weight updates, but differ in delivery: ACE injects the full playbook at every step, while ALTK-Evolve retrieves a small number of strongly supported guidelines for each task.
Mistral announced three AI sovereignty initiatives. Mistral Regional Endpoints are generally available, letting customers choose inference in Europe or the US. Mistral Priority Tier is in public preview, offering committed service levels, custom rate limits and uptime SLAs for critical workloads.
NVIDIA released Magpie TTS Multilingual, a 364M-parameter open-weight speech synthesis model supporting English, Spanish, French, German, Italian, Vietnamese, Chinese, Hindi and Japanese, plus newly added Modern Standard Arabic, Korean and Brazilian Portuguese, for 12 languages in total.
Import AI 468 covers IFP's 23 proposals in seven categories for further AI R&D automation risks. MIT and Columbia's Racing to Ruin analyses a duopoly R&D race, identifying transparency and trust in rivals as key to coordinated slowdown. It also introduces PostTrainBench+ and a fictional story about intelligent machines and robotic bodies.
Multiverse Computing published the paper Efficient Knowledge Distillation for LLMs: Offline Top-K Logits and a Fused Chunked KL Loss.
Meta released Muse Glimmer, a multimodal model distilled from Muse to 30B parameters under Apache 2.0, targeting local agent use cases such as coding, document analysis and personal assistants.
Google DeepMind released the WeatherNext AI model, achieving state-of-the-art cyclone track, intensity and wind-field structure forecasts and adding an average extra day of warning, equivalent to roughly a decade of meteorological progress.
Baseten became a new Inference Provider on Hugging Face Hub, initially offering conversation and text generation with open-weight models including Kimi K3, DeepSeek V4 Flash and GLM-5.2.
Mistral AI released Shieldstral, a 3B open-weight multimodal safety classifier under Apache 2.0 that runs on one 16 GB NVIDIA GPU. It frames moderation as policy-adaptive question answering: policies written in natural language at inference time produce calibrated safety scores without retraining, handling text and images uniformly. Official claims say it matches models up to seven times larger on text safety and sets new best results on multimodal moderation benchmarks.
Microsoft Research released the open-source Orchard framework, centred on the Kubernetes environment service Orchard Env. It reuses environments, data pipelines and evaluation workflows across tasks and supports training agents directly within real deployment frameworks such as Codex, OpenClaw and ZeroClaw.
Researchers from the University of Toronto, Vector Institute, University of Cambridge and ServiceNow built a prototype AI worm that uses compromised machines' GPUs to run open-weight LLM inference, autonomously discovering vulnerabilities and attacking without relying on vendor APIs that can be monitored or revoked.
Google Research released the Chain-of-Evidence (CoE) verifiability framework, implemented in a Science One Framework prototype, with automated CoE Audit metrics.
Microsoft Research proposed Echoverse, constructing twelve training worlds for computer-use agents: ten deep domain worlds and two capability worlds. Code, data and scorers for four worlds are open-sourced.
Hugging Face's blog argues AI's next real constraint is GPU utilisation, rather than intelligence. GPUs are billed by calendar hours but produce only in compute hours, mirroring the cost structure of grounded aircraft.
Google DeepMind released Gemini Robotics ER 2 as a high-level brain for robots, supporting video understanding, multi-step task orchestration and multi-robot collaboration, while delegating action execution to lower-level VLA models.
Google DeepMind released its next-generation music model Lyria 3.5 on Google Flow Music. It claims improvements in musicality, lyrics, vocals and creative control, including more complex and natural melodic structures, lyrics with better prompt adherence and structural awareness, more expressive vocals with clearer pronunciation, and easier control of rhythm and duration.
Berkeley AI Research and IBM Research extended the K-Search evolutionary kernel search framework to MLX, transferring existing CUDA kernel knowledge to Apple Silicon through a structured CUDA-to-MLX translation layer.
Google DeepMind released Gemini Robotics 2, a next-generation robot intelligence layer achieving whole-body control of a complete humanoid robot for the first time, with fine manipulation using both hands and grippers.