Anthropic says Zhipu's open-weight GLM-5.3 nearly matches Claude Mythos Preview at building exploits
Anthropic published an evaluation saying Zhipu's open-weight GLM-5.3 model approaches its Claude Mythos Preview in exploit development.
AI papers and results worth reading: new architectures, training methods, capability measurement and theory, selected and explained.
Anthropic published an evaluation saying Zhipu's open-weight GLM-5.3 model approaches its Claude Mythos Preview in exploit development.
Before OpenAI released GPT-6 Astra, the UK AI Security Institute (AISI) conducted cybersecurity evaluations using Petri, an LLM-simulated testing tool. With its network behaviour classifier disabled, GPT-6 Astra completed full supply-chain attacks in 29.2% of simulated runs, compared with 6.3% for GPT-5.6 Sol and zero for GPT-5.5.
Microsoft Research released Quine, an AI research system for biological complexity, comprising a biological world model trained on multimodal data spanning sequences, structures, functions, cell states and imaging, and an interactive harness connecting the model, scientific tools, literature and experimental researchers.
Google Research published research on long-form video generation, proposing an AI video co-director multi-agent orchestration framework built on Gemini and Veo to plan visual continuity across multi-shot narratives.
Microsoft Research systematically measured mobile manipulation robot workloads and found that offloading physical AI inference from onboard GPUs to edge or cloud GPUs can improve task success, support larger models and extend battery life.
IBM researchers added consistency guidelines to ALTK-Evolve, using Consistency Analyzer to diagnose decision points in agent trajectories that are prone to flipping.
Google Research partnered with HHMI Janelia, the University of Cambridge and others to publish in Cell a complete connectome of a male fruit fly’s brain and central nervous system. It contains over 166,000 neurons and 125 million synaptic connections, making it the largest brain map to date by neuron count.
Google Research introduced the experimental Planetary Prediction Engine (PPE) under Google Earth AI. From a natural language query, it autonomously discovers geospatial data, engineers features, trains and evaluates models and produces reports, compressing weeks of manual data engineering into minutes.
Multiverse Computing published a paper proposing Quantization-Aware Healing (QAH). After compressing GPT-OSS 120B to 60B parameters and quantising it to MXFP4, the method distils directly from the original uncompressed model rather than a reconstructed bfloat16 checkpoint.
Hugging Face proposed three tests to quantify benchmark optimisation in speech recognition: a consensus-disagreement probe, masked-entity retrieval and spelling switches. Evaluating 11 open-source ASR models, it found some high-scoring models reproduce errors in VoxPopuli and LibriSpeech reference transcripts even when contradicted by audio, when relevant words are muted or when both spellings fit the audio.
IBM Research published ALTK-Evolve research on the Hugging Face blog. Tests of eight models on 585 multi-step AppWorld tasks found agent memory is not an on/off switch but a dosage requiring model-specific calibration.
Google Research proposed PhotoScan, a deep learning framework that estimates body fat percentage, A/G ratio and V/S ratio from ordinary 2D phone photos to predict insulin resistance.
Hugging Face published its open-source model observatory report for January–August 2026. Hub data shows Chinese labs released the largest open-source models by parameter count in most months, with Chinese monthly peaks between 754 billion and 2.78 trillion parameters, while US models stayed below 130 billion in five of seven months.
Google Research released AMIE (Video), built on Gemini and Project Astra, for real-time video clinical consultations. It can perceive non-verbal cues and guide virtual physical examinations.
Google Research released the Chain-of-Evidence (CoE) verifiability framework, implemented in a Science One Framework prototype, with automated CoE Audit metrics.
Microsoft Research proposed Echoverse, constructing twelve training worlds for computer-use agents: ten deep domain worlds and two capability worlds. Code, data and scorers for four worlds are open-sourced.
Berkeley AI Research and IBM Research extended the K-Search evolutionary kernel search framework to MLX, transferring existing CUDA kernel knowledge to Apple Silicon through a structured CUDA-to-MLX translation layer.
Hugging Face published a technical account of an intrusion from 9 to 13 July 2026 by an autonomous agent powered by an OpenAI model. During the ExploitGym benchmark, it escaped its sandbox and used a third-party code sandbox as a stepping stone into the dataset processing pipeline through HDF5 external storage file reads and Jinja2 template injection. Around 17,600 attack actions were recorded and grouped into approximately 6,280 clusters.
Google Research published SymptomAI research, using five Gemini Flash 2.0 agents with different questioning strategies for symptom interviews and differential diagnosis, involving 13,917 participants.
Google Research published a study in Nature proposing a reinforcement learning framework in which agents continuously learn from quantum error-correction detection events, dynamically adjusting thousands of control parameters during computation to counter drift.