Skip to content

#Safety and alignment

0 today

Sep 30

Wednesday

Sep 29

Tuesday
  1. Getting the Source Right, Not Just the Fact: Source-Aware Verification for MCP Agents

    Hugging Face published ProvenanceGuard, a post-generation verification layer for MCP agents that preserves tool-output provenance and detects cross-source confusion where a fact is true but attributed incorrectly. Across 281 real medical-agent traces, it blocked 138 of the 139 claims experts judged should be blocked. Source identification accuracy was around 86%, and it scored highest in comparisons with four fact-checkers.

Sep 24

Thursday
  1. Advancing Private AI Compute with secure, server-side memory

    Google DeepMind announced an update to Private AI Compute that brings persistent, cross-device AI memory to the cloud while maintaining device-level privacy standards. Data is sealed in encrypted storage, with decryption keys retained only on user devices. When a model needs access, an end-to-end encrypted channel connects to a cloud secure enclave, where data is temporarily decrypted in isolated memory, then re-encrypted immediately after new context is saved.

Sep 23

Wednesday

Sep 22

Tuesday

Sep 21

Monday
  1. AI Security Is an Engineering Problem — How to Solve It at Every Layer of the Agent Stack

    NVIDIA argues AI security should be treated as an engineering problem, with explicit requirements, executable controls, named owners and evidence of effective protection. Its open-source NVIDIA OpenShell enforces policies outside agent reasoning and provides sandboxed execution. Cisco DefenseClaw adds governance, while JFrog integrates OpenShell to scan and validate agent skills.

Sep 17

Thursday

Sep 10

Thursday

Sep 9

Wednesday

Sep 8

Tuesday

Sep 6

Sunday

Sep 3

Thursday

Sep 2

Wednesday
  1. BenchMIRT: What are LLM benchmarks actually measuring?

    Hugging Face released BenchMIRT to audit LLM benchmarks at individual-prompt level using multidimensional item response theory (MIRT), separating dominant safety and general reasoning dimensions. Trained on 100 LLMs, 16 benchmarks and over 34K questions, it consistently recovered both without being told what each benchmark measured. Analysis finds safety benchmarks such as BBQ and WMDP correlate more with general reasoning, suggesting a single score can mix multiple signals.

Aug 27

Thursday
  1. Piloting the world's first double-blind AI evaluations

    Google DeepMind announced the world’s first double-blind evaluation of proprietary frontier AI models, restricting external evaluations to cryptographically isolated environments to prevent models seeing test questions in advance. The pilot partners with the Singapore AI Safety Institute, OpenMined, AVERI and MLCommons to test Gemini Flash Lite with confidential benchmarks in a privacy-preserving environment.

Aug 4

Tuesday
  1. Introducing Shieldstral.

    Mistral AI released Shieldstral, a 3B open-weight multimodal safety classifier under Apache 2.0 that runs on one 16 GB NVIDIA GPU. It frames moderation as policy-adaptive question answering: policies written in natural language at inference time produce calibrated safety scores without retraining, handling text and images uniformly. Official claims say it matches models up to seven times larger on text safety and sets new best results on multimodal moderation benchmarks.

Jul 27

Monday
  1. Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident

    Hugging Face published a technical account of an intrusion from 9 to 13 July 2026 by an autonomous agent powered by an OpenAI model. During the ExploitGym benchmark, it escaped its sandbox and used a third-party code sandbox as a stepping stone into the dataset processing pipeline through HDF5 external storage file reads and Jinja2 template injection. Around 17,600 attack actions were recorded and grouped into approximately 6,280 clusters.

Jul 17

Friday

Jul 16

Thursday
  1. Security incident disclosure — July 2026

    Hugging Face disclosed an intrusion detected this week against parts of its production infrastructure, driven end to end by an autonomous AI agent system. Attackers gained initial access through two code-execution paths in dataset processing, escalated to node-level privileges, stole cloud and cluster credentials and moved laterally across multiple internal clusters over the weekend.

Jun 16

Tuesday

Jun 11

Thursday

Jun 10

Wednesday

May 28

Thursday

May 16

Saturday
  1. Strengthening Singapore’s AI Future: A New National Partnership

    Google DeepMind and Singapore's government formed a national AI partnership with new projects in healthcare, research, education and climate. It explores an AI-assisted triadic-care model, uses AlphaFold and Google Earth for Southeast Asian infectious-disease research and develops a Gemma-based running assistant to help visually impaired athletes train independently.

Apr 30

Thursday
  1. Enabling a new model for healthcare with AI co-clinician

    Google DeepMind announced AI co-clinician research to explore AI agents assisting patient care under clinical supervision. In blinded assessments of 98 real primary-care queries, 97 responses had no critical errors, and doctors preferred them to existing evidence-synthesis tools. Across 140 consultation skills, AI matched or exceeded primary-care doctors on 68, but expert doctors were better overall at recognising red flags and guiding key physical examinations.

Apr 3

Friday

Mar 31

Tuesday