Skip to content

#Reasoning

1 today

Oct 1

ThursdayToday1 items
  1. Gemini 4 Argon: our next era of frontier intelligence

    Sep 30, 2026 | Gemini 4 Argon delivers frontier performance in complex workflows across real-world software engineering, enterprise knowledge work like legal and finance, and cybersecurity defense. Today, we’re announcing our new frontier model, Gemini 4 Argon, which is rolling out to a set of trusted cyber defenders through our Fairwind Program.

Sep 24

Thursday

Sep 16

Wednesday

Sep 8

Tuesday

Sep 3

Thursday

Sep 2

Wednesday
  1. BenchMIRT: What are LLM benchmarks actually measuring?

    Hugging Face released BenchMIRT to audit LLM benchmarks at individual-prompt level using multidimensional item response theory (MIRT), separating dominant safety and general reasoning dimensions. Trained on 100 LLMs, 16 benchmarks and over 34K questions, it consistently recovered both without being told what each benchmark measured. Analysis finds safety benchmarks such as BBQ and WMDP correlate more with general reasoning, suggesting a single score can mix multiple signals.

Aug 25

Tuesday

Aug 21

Friday

Aug 13

Thursday
  1. MindTopo reveals VLMs’ spatial reasoning abilities

    Microsoft Research introduced MindTopo to evaluate multimodal models' reasoning and planning over connectivity, enclosure, order, separation and knotting. Current models perform much better at static-image recognition than interactive planning, with failures mainly in planning rather than perception and overall performance far below humans. Image and video generation helps only when preserving structural relations in a single frame; multi-step operations often change topology or violate constraints.

Aug 11

Tuesday

Jul 31

Friday

Jul 29

Wednesday

Jul 26

Sunday

Jul 15

Wednesday

Jul 10

Friday

Jul 2

Thursday

Jun 27

Saturday

Jun 25

Thursday
  1. Thinking to recall: How reasoning unlocks parametric knowledge in LLMs

    A Google Research paper at COLM 2026 finds reasoning traces unlock factual knowledge LLMs otherwise cannot recall, even for simple single-hop questions. Tests on Gemini-2.5 Flash and Pro and Qwen3-32B identify two mechanisms: extra tokens act as a 'computation buffer', and 'factual priming' produces related facts to semantically prepare the correct answer. Self-generated intermediate facts can also introduce hallucination risks.

Jun 11

Thursday

May 8

Friday
  1. Adaptive Parallel Reasoning: The Next Paradigm in Efficient Inference Scaling

    Berkeley AI Research reviews parallel reasoning, focusing on models deciding when to decompose and parallelise independent subtasks, how many threads to generate and how to coordinate. Existing approaches including Self-consistency, Best-of-N, Tree of Thoughts, MCTS, ParaThinker, GroupThink and Hogwild! Inference mostly impose parallel structures externally rather than teaching adaptive behaviour.

Apr 20

Monday
  1. Gradient-based Planning for World Models at Longer Horizons

    Berkeley AI Research proposed GRASP, a gradient-based planner for learned world models. It lifts trajectories into virtual states for parallel optimisation across time, injects randomness directly into state iterations for exploration and reshapes gradients to give actions clear signals. Avoiding fragile state-input gradients in high-dimensional visual models makes long-horizon planning more practical and robust.

Apr 16

Thursday

Mar 25

Wednesday

Mar 17

Tuesday

Mar 13

Friday
  1. Identifying Interactions at Scale for LLMs

    Berkeley AI Research proposed SPEX and ProxySPEX to identify key interactions driving LLM outputs at scale in feature, data and model-component attribution. SPEX turns interaction search into sparse recovery using sparsity and low order; ProxySPEX exploits hierarchy to match SPEX with roughly ten times fewer ablations.

Nov 1

Saturday
  1. RL without TD learning

    Berkeley AI Research proposed a divide-and-conquer off-policy reinforcement learning algorithm without temporal-difference (TD) learning. It reduces Bellman recursions from linear to logarithmic counts, scaling to long-horizon tasks.

Jul 17

Thursday
  1. Le Chat dives deep.

    Mistral added features to Le Chat including preview Deep Research, Voxtral-powered voice mode, multilingual reasoning with Magistral, Projects for organising conversations and image editing with Black Forest Labs.

Jun 10

Tuesday