Skip to content

#Research

2 today

Aug 12

Wednesday
  1. Introducing CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement

    Microsoft Research released CARE-X, a unified chest X-ray vision-language research model for report generation and structured prediction. It rewards clinical correctness using multitask reinforcement learning (DAPO). Generation and dual-inference modes cover lesion presence and negation, localisation, multilabel classification, catheter and tube malposition detection, and localisation of 29 anatomical regions.

Aug 10

Monday
  1. Import AI 468: 23 RSI ideas; PostTrainBench+; and how trust and transparency interplay with AI racing

    Import AI 468 covers IFP's 23 proposals in seven categories for further AI R&D automation risks. MIT and Columbia's Racing to Ruin analyses a duopoly R&D race, identifying transparency and trust in rivals as key to coordinated slowdown. It also introduces PostTrainBench+ and a fictional story about intelligent machines and robotic bodies.

Aug 3

Monday

Jul 31

Friday

Jul 29

Wednesday

Jul 27

Monday
  1. Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident

    Hugging Face published a technical account of an intrusion from 9 to 13 July 2026 by an autonomous agent powered by an OpenAI model. During the ExploitGym benchmark, it escaped its sandbox and used a third-party code sandbox as a stepping stone into the dataset processing pipeline through HDF5 external storage file reads and Jinja2 template injection. Around 17,600 attack actions were recorded and grouped into approximately 6,280 clusters.

Jul 26

Sunday

Jul 23

Thursday

Jul 16

Thursday

Jul 15

Wednesday

Jul 9

Thursday

Jul 8

Wednesday

Jul 1

Wednesday

Jun 30

Tuesday

Jun 27

Saturday

Jun 25

Thursday
  1. Optimizing cloud economics with linear elastic caching

    Google Research proposed linear elastic caching, modelling page eviction as a ski-rental problem and using shallow decision trees to predict page TTLs, dynamically resizing caches to minimise total ownership cost. After months in Spanner production, memory fell 15.5%, cache misses rose only 5.5%, TCO dropped around 5% and actual I/O cost impact was just 0.5%. It was also validated on several public cache traces.

  2. Thinking to recall: How reasoning unlocks parametric knowledge in LLMs

    A Google Research paper at COLM 2026 finds reasoning traces unlock factual knowledge LLMs otherwise cannot recall, even for simple single-hop questions. Tests on Gemini-2.5 Flash and Pro and Qwen3-32B identify two mechanisms: extra tokens act as a 'computation buffer', and 'factual priming' produces related facts to semantically prepare the correct answer. Self-generated intermediate facts can also introduce hallucination risks.

Jun 22

Monday

Jun 17

Wednesday
  1. From pixels to planning: Earth AI for nature restoration

    Google Research released vector data converting Farmscapes 2020 high-resolution raster maps into usable inventories of hedgerows, stone walls and coppices across over 130,000 square kilometres of the UK. The framework fine-tunes an RSF Vision-Transformer pretrained on over 300 million global satellite images with around 247 square kilometres of labelled data. Polsby–Popper compactness distinguishes woodland, clusters of trees and hedgerows, with a threshold below 0.5 for linear features.

Jun 16

Tuesday

Jun 13

Saturday
  1. Research into how AI can help users understand skin conditions

    Google published two studies on AI-assisted understanding of skin problems. In a survey of 2,345 participants, over 62% using AI tried to name a skin condition versus 41% in the control group, with 23% accuracy, nearly triple the control group's 8%. A mixed-methods study examined how people use these tools for their own skin concerns, their understanding and differences in doctor communication. AI offered limited help in deciding the next steps for seeking care.

Jun 11

Thursday

Jun 8

Monday

Jun 5

Friday

May 28

Thursday

May 18

Monday

May 16

Saturday

May 12

Tuesday

Apr 30

Thursday
  1. Enabling a new model for healthcare with AI co-clinician

    Google DeepMind announced AI co-clinician research to explore AI agents assisting patient care under clinical supervision. In blinded assessments of 98 real primary-care queries, 97 responses had no critical errors, and doctors preferred them to existing evidence-synthesis tools. Across 140 consultation skills, AI matched or exceeded primary-care doctors on 68, but expert doctors were better overall at recognising red flags and guiding key physical examinations.

Apr 22

Wednesday

Apr 20

Monday
  1. Gradient-based Planning for World Models at Longer Horizons

    Berkeley AI Research proposed GRASP, a gradient-based planner for learned world models. It lifts trajectories into virtual states for parallel optimisation across time, injects randomness directly into state iterations for exploration and reshapes gradients to give actions clear signals. Avoiding fragile state-input gradients in high-dimensional visual models makes long-horizon planning more practical and robust.