Ataraxos 击败 Stratego 史上最强选手,算力成本约为 DeepMind DeepNash 的 1/500
研究团队开发的 Ataraxos 在与 Stratego 史上最强选手 Niemeijer 的对局中取得 85% 的有效胜率,终结了人类在这款不完全信息棋盘游戏上的优势。
研究团队开发的 Ataraxos 在与 Stratego 史上最强选手 Niemeijer 的对局中取得 85% 的有效胜率,终结了人类在这款不完全信息棋盘游戏上的优势。
Google Research proposed Retrieve-for-Train, using offline reinforcement learning to find reward-aligned query fan-out and compile it into supervision, then distilling it into a 53.9M-parameter diffusion retriever. At inference, it performs one non-autoregressive query fan-out without generating CoT reasoning tokens.
OpenAI shared an AI-generated solution to the Navier–Stokes Millennium Prize Problem, including a written explanation and a formal Lean proof.
Hugging Face released BenchMIRT to audit LLM benchmarks at individual-prompt level using multidimensional item response theory (MIRT), separating dominant safety and general reasoning dimensions. Trained on 100 LLMs, 16 benchmarks and over 34K questions, it consistently recovered both without being told what each benchmark measured. Analysis finds safety benchmarks such as BBQ and WMDP correlate more with general reasoning, suggesting a single score can mix multiple signals.
Multiverse Computing published a paper proposing Quantization-Aware Healing (QAH). After compressing GPT-OSS 120B to 60B parameters and quantising it to MXFP4, the method distils directly from the original uncompressed model rather than a reconstructed bfloat16 checkpoint.
Import AI 469 introduces DiG-bench (Discovery in Games), with 70 games whose hidden rules and goals must be discovered through interaction. Opus 5 and Fable 5 with Claude Code performed best; only they completed some Tier 7 tasks, while humans achieved 100% completion.
Microsoft Research introduced MindTopo to evaluate multimodal models' reasoning and planning over connectivity, enclosure, order, separation and knotting. Current models perform much better at static-image recognition than interactive planning, with failures mainly in planning rather than perception and overall performance far below humans. Image and video generation helps only when preserving structural relations in a single frame; multi-step operations often change topology or violate constraints.
Google Research released the Chain-of-Evidence (CoE) verifiability framework, implemented in a Science One Framework prototype, with automated CoE Audit metrics.
Berkeley AI Research and IBM Research extended the K-Search evolutionary kernel search framework to MLX, transferring existing CUDA kernel knowledge to Apple Silicon through a structured CUDA-to-MLX translation layer.
Berkeley AI Research proposed ABBEL, isolating and supervising LLM summaries as natural-language belief states and using autoencoder-inspired reconstructive belief scoring as an auxiliary RL task. On CollabBench, ABBEL-rec-BG narrowed the performance gap to the full-context model by around 50%, reduced training steps from 100 to 50 and used a shorter peak token length.
A Google Research paper at COLM 2026 finds reasoning traces unlock factual knowledge LLMs otherwise cannot recall, even for simple single-hop questions. Tests on Gemini-2.5 Flash and Pro and Qwen3-32B identify two mechanisms: extra tokens act as a 'computation buffer', and 'factual priming' produces related facts to semantically prepare the correct answer. Self-generated intermediate facts can also introduce hallucination risks.
Berkeley AI Research proposed GRASP, a gradient-based planner for learned world models. It lifts trajectories into virtual states for parallel optimisation across time, injects randomness directly into state iterations for exploration and reshapes gradients to give actions clear signals. Avoiding fragile state-input gradients in high-dimensional visual models makes long-horizon planning more practical and robust.
Google proposed Simula, a reasoning-first synthetic data framework reframing dataset construction as mechanism design. It builds datasets from scratch without seed data through global diversification, local diversification, complexification and dual-critic quality checks.
Google Research released TurboQuant, compressing the KV cache to three bits without loss of model accuracy or training and fine-tuning, alongside the QJL and PolarQuant methods.
Berkeley AI Research proposed SPEX and ProxySPEX to identify key interactions driving LLM outputs at scale in feature, data and model-component attribution. SPEX turns interaction search into sparse recovery using sparsity and low order; ProxySPEX exploits hierarchy to match SPEX with roughly ten times fewer ablations.
Berkeley AI Research proposed a divide-and-conquer off-policy reinforcement learning algorithm without temporal-difference (TD) learning. It reduces Bellman recursions from linear to logarithmic counts, scaling to long-horizon tasks.