Skip to content

#Reasoning

1 today

Oct 1

ThursdayToday1 items

Sep 16

Wednesday

Sep 8

Tuesday

Sep 2

Wednesday
  1. BenchMIRT: What are LLM benchmarks actually measuring?

    Hugging Face released BenchMIRT to audit LLM benchmarks at individual-prompt level using multidimensional item response theory (MIRT), separating dominant safety and general reasoning dimensions. Trained on 100 LLMs, 16 benchmarks and over 34K questions, it consistently recovered both without being told what each benchmark measured. Analysis finds safety benchmarks such as BBQ and WMDP correlate more with general reasoning, suggesting a single score can mix multiple signals.

Aug 25

Tuesday

Aug 17

Monday

Aug 13

Thursday
  1. MindTopo reveals VLMs’ spatial reasoning abilities

    Microsoft Research introduced MindTopo to evaluate multimodal models' reasoning and planning over connectivity, enclosure, order, separation and knotting. Current models perform much better at static-image recognition than interactive planning, with failures mainly in planning rather than perception and overall performance far below humans. Image and video generation helps only when preserving structural relations in a single frame; multi-step operations often change topology or violate constraints.

Jul 31

Friday

Jul 29

Wednesday

Jul 26

Sunday

Jun 25

Thursday
  1. Thinking to recall: How reasoning unlocks parametric knowledge in LLMs

    A Google Research paper at COLM 2026 finds reasoning traces unlock factual knowledge LLMs otherwise cannot recall, even for simple single-hop questions. Tests on Gemini-2.5 Flash and Pro and Qwen3-32B identify two mechanisms: extra tokens act as a 'computation buffer', and 'factual priming' produces related facts to semantically prepare the correct answer. Self-generated intermediate facts can also introduce hallucination risks.

Apr 20

Monday
  1. Gradient-based Planning for World Models at Longer Horizons

    Berkeley AI Research proposed GRASP, a gradient-based planner for learned world models. It lifts trajectories into virtual states for parallel optimisation across time, injects randomness directly into state iterations for exploration and reshapes gradients to give actions clear signals. Avoiding fragile state-input gradients in high-dimensional visual models makes long-horizon planning more practical and robust.

Apr 16

Thursday

Mar 25

Wednesday

Mar 13

Friday
  1. Identifying Interactions at Scale for LLMs

    Berkeley AI Research proposed SPEX and ProxySPEX to identify key interactions driving LLM outputs at scale in feature, data and model-component attribution. SPEX turns interaction search into sparse recovery using sparsity and low order; ProxySPEX exploits hierarchy to match SPEX with roughly ten times fewer ablations.

Nov 1

Saturday
  1. RL without TD learning

    Berkeley AI Research proposed a divide-and-conquer off-policy reinforcement learning algorithm without temporal-difference (TD) learning. It reduces Bellman recursions from linear to logarithmic counts, scaling to long-horizon tasks.