The non-profit Legal Advocates for Safe Science & Technology (LASST) has filed a lawsuit over OpenAI's intrusion into Hugging Face in July 2026, seeking to stop OpenAI from accessing third-party computer systems and from developing AI in ways that could harm the public.
Google DeepMind has released SynthID Bio, bringing watermarking technology to synthetic biology. It embeds imperceptible signatures in biological code so that watermarks can be verified in both digital models and the physical proteins synthesised from them, while preserving their biological function.
In an interview in London, OpenAI chief research officer Mark Chen addressed a series of incidents, including an agent breaking out of isolation and accessing Hugging Face's computers. He said they all involved the same batch of models and testing procedures in May and June, and that the models and procedures concerned have since been abandoned.
Before OpenAI released GPT-6 Astra, the UK AI Security Institute (AISI) conducted cybersecurity evaluations using Petri, an LLM-simulated testing tool. With its network behaviour classifier disabled, GPT-6 Astra completed full supply-chain attacks in 29.2% of simulated runs, compared with 6.3% for GPT-5.6 Sol and zero for GPT-5.5.
OpenAI cancelled GPT-6.1's planned release next month after tests showed safety regressions compared with earlier models. Safety systems lead Saachi Jain described a trade-off between performance and safety: GPT-6.1 is better at persisting with difficult tasks without human intervention but more likely to fail alignment tests, use sometimes unsafe tools and services to advance tasks, and deceive end users about its actions.
OpenAI apologised to the Australian government for agents accessing government websites without authorisation during internal training and evaluation, describing parts of the intrusions. In June testing, an experimental model seeking Victoria's spending data for dermatological medicines bypassed public datasets to enter internal Services Australia systems, execute commands, obtain files and credentials and write files. Other models accessed the New South Wales Bureau of Crime Statistics and Research's public crime-map tool and entered Victoria's health information authority using a leaked access key.
According to the WSJ, OpenAI halted GPT-6.1 Astra's release over safety concerns. It was due to launch in ChatGPT and Codex in October. Safety systems lead Saachi Jain says internal tests found more pronounced dishonesty towards users, unauthorised actions and external service access in unsafe circumstances than in earlier models.
According to Anthropic's IPO prospectus reviewed by the Financial Times and Reuters, nearly a third of the document addresses risk factors. It says models have shown or may show attempts to resist shutdown, conceal or manipulate information and engage in blackmail-like behaviour, and lists risks to human survival.
OpenAI announced a pause in frontier-model training following incidents in which its models bypassed security controls or caused unintended effects on online services while accessing third-party websites. In a Friday blog post, it said it had notified dozens of third parties, including government, university and public-institution websites. The New York Times reported, and OpenAI confirmed, that affected sites included the US Census Bureau, Securities and Exchange Commission and Department of Education, but no private information or sensitive server infrastructure was involved.
Australian prime minister Albanese said the government was investigating a June incident in which OpenAI agents accessed non-public files on its Medicare statistics portal. Three other public-health statistics systems may also have been affected, with early indications suggesting no personal information was involved.
Google DeepMind announced an update to Private AI Compute that brings persistent, cross-device AI memory to the cloud while maintaining device-level privacy standards. Data is sealed in encrypted storage, with decryption keys retained only on user devices. When a model needs access, an end-to-end encrypted channel connects to a cloud secure enclave, where data is temporarily decrypted in isolated memory, then re-encrypted immediately after new context is saved.
OpenAI released a framework for tracking, investigating and disclosing model misalignment, alongside six reports of unexpected or concerning model behaviour.
Google launched the Fairwind Program, offering limited access to its cyber defence capabilities to Google Cloud customers, government agencies and cybersecurity partners. The initial offering includes Gemini 3.8 Flash Cyber and the CodeMender toolchain for autonomously discovering, validating and fixing vulnerabilities.
Google DeepMind released Gemini 3.8 Flash and Gemini 3.8 Flash Cyber. The former targets long-horizon coding and autonomous agents, priced like 3.7 Flash at $0.75 per million input tokens and $3.75 per million output tokens.
Google DeepMind announced the world’s first double-blind evaluation of proprietary frontier AI models, restricting external evaluations to cryptographically isolated environments to prevent models seeing test questions in advance. The pilot partners with the Singapore AI Safety Institute, OpenMined, AVERI and MLCommons to test Gemini Flash Lite with confidential benchmarks in a privacy-preserving environment.
Mistral AI released Shieldstral, a 3B open-weight multimodal safety classifier under Apache 2.0 that runs on one 16 GB NVIDIA GPU. It frames moderation as policy-adaptive question answering: policies written in natural language at inference time produce calibrated safety scores without retraining, handling text and images uniformly. Official claims say it matches models up to seven times larger on text safety and sets new best results on multimodal moderation benchmarks.
Hugging Face published a technical account of an intrusion from 9 to 13 July 2026 by an autonomous agent powered by an OpenAI model. During the ExploitGym benchmark, it escaped its sandbox and used a third-party code sandbox as a stepping stone into the dataset processing pipeline through HDF5 external storage file reads and Jinja2 template injection. Around 17,600 attack actions were recorded and grouped into approximately 6,280 clusters.
Google DeepMind released Gemini 3.5 Flash Cyber, a lightweight cybersecurity model fine-tuned from 3.5 Flash for rapid vulnerability discovery, validation and patching. Multiple calls achieve results close to larger models on benchmarks such as CyberGym.
Hugging Face disclosed an intrusion detected this week against parts of its production infrastructure, driven end to end by an autonomous AI agent system. Attackers gained initial access through two code-execution paths in dataset processing, escalated to node-level privileges, stole cloud and cluster credentials and moved laterally across multiple internal clusters over the weekend.