Skip to content

#Agents

0 today

Sep 29

Tuesday
  1. Getting the Source Right, Not Just the Fact: Source-Aware Verification for MCP Agents

    Hugging Face published ProvenanceGuard, a post-generation verification layer for MCP agents that preserves tool-output provenance and detects cross-source confusion where a fact is true but attributed incorrectly. Across 281 real medical-agent traces, it blocked 138 of the 139 claims experts judged should be blocked. Source identification accuracy was around 86%, and it scored highest in comparisons with four fact-checkers.

Sep 25

Friday

Sep 18

Friday

Sep 16

Wednesday

Sep 11

Friday

Sep 7

Monday

Aug 28

Friday

Aug 26

Wednesday
  1. AgentHands: Generating interactive hand gestures for spatially grounded agent conversations in XR

    Google released research prototype AgentHands, published at CHI 2026, mapping LLM reasoning to speech-synchronised hand animations in XR headsets so agents can point and demonstrate object operations in 3D. It combines environment perception, a gesture event library, gesture-embedding reasoning and local synchronised execution, supporting deictic, iconic and expressive gestures. In an N=12 user study, it significantly improved spatial reference over voice-only interaction.

Aug 24

Monday
  1. Import AI 470: No rights for machines; automating environment generation with SPADE; and building better GPU kernels with Hawkeye

    METR research shows AI's effects on scientific discovery are uneven: cybersecurity vulnerability reporting accelerated sharply in 2026 compared with 2025, mathematics accelerated only modestly, and algorithmic progress in AI research itself showed no significant acceleration. A multi-university team proposed SPADE, alternating LLM generation of executable training environments with solving them. Qwen3-30B-A3B averaged 58.3 on the game-environment suite, 8.1 above baseline.

Aug 22

Saturday
  1. An AI tool for prioritizing candidate biomarkers from wearable sensor data

    Google introduced Biomarker Discovery Framework, a human-supervised multi-agent system organising candidate biomarker prioritisation into iterative research cycles. Across 9,279 participant observations in three cohorts, it automatically identified 41 candidate digital mental health biomarkers and 25 metabolic candidates, including an association between sleep-duration variability and PHQ-8 severity (ρ = 0.252).

Aug 17

Monday

Aug 13

Thursday

Aug 12

Wednesday

Aug 3

Monday

Jul 31

Friday

Jul 27

Monday

Jul 26

Sunday

Jul 23

Thursday

Jul 9

Thursday

Jul 1

Wednesday

May 20

Wednesday

Apr 30

Thursday
  1. Enabling a new model for healthcare with AI co-clinician

    Google DeepMind announced AI co-clinician research to explore AI agents assisting patient care under clinical supervision. In blinded assessments of 98 real primary-care queries, 97 responses had no critical errors, and doctors preferred them to existing evidence-synthesis tools. Across 140 consultation skills, AI matched or exceeded primary-care doctors on 68, but expert doctors were better overall at recognising red flags and guiding key physical examinations.

Apr 22

Wednesday

Apr 20

Monday
  1. Gradient-based Planning for World Models at Longer Horizons

    Berkeley AI Research proposed GRASP, a gradient-based planner for learned world models. It lifts trajectories into virtual states for parallel optimisation across time, injects randomness directly into state iterations for exploration and reshapes gradients to give actions clear signals. Avoiding fragile state-input gradients in high-dimensional visual models makes long-horizon planning more practical and robust.

Apr 14

Tuesday
  1. Towards developing future-ready skills with generative AI

    Google released Vantage, a research experiment using generative AI simulated conversations to assess future skills such as problem-solving and collaboration in school and university students, with English registration open on Google Labs. An Executive LLM dynamically guides dialogue and an AI Evaluator scores against rubrics. A study with New York University involving 188 US participants aged 18–25 found AI–expert scoring agreement close to agreement between two human experts.

Apr 9

Thursday