Skip to content

AI agents

Models that plan, call tools and finish multi-step tasks on their own: from Claude Code and Manus to agent frameworks and benchmarks.

107selectedRelated topicsMCP and tool useAI codingReasoning

Latest selected

1–20 of 107

Oct 1

ThursdayToday
  1. Gemini 4 Argon: our next era of frontier intelligence

    Sep 30, 2026 | Gemini 4 Argon delivers frontier performance in complex workflows across real-world software engineering, enterprise knowledge work like legal and finance, and cybersecurity defense. Today, we’re announcing our new frontier model, Gemini 4 Argon, which is rolling out to a set of trusted cyber defenders through our Fairwind Program.

Sep 30

Wednesday
  1. Build a multi-agent music production pipeline on Amazon Bedrock AgentCore Runtime Instances

    The official AWS blog demonstrates deploying a three-agent music production pipeline on Amazon Bedrock AgentCore Runtime Instances. The composition agent uses Claude Sonnet 4.6 to generate a music brief and runs ACE-Step on the instance's NVIDIA L4 to render audio. The delivery agent reads .wav files from the shared volume for measurement and DSP processing, while the compliance agent independently remeasures them and checks harmonic similarity against a music catalogue.

  2. “We’re not going to shoot ourselves in the foot” over hack fallout, says OpenAI’s chief research officer

    In an interview in London, OpenAI chief research officer Mark Chen addressed a series of incidents, including an agent breaking out of isolation and accessing Hugging Face's computers. He said they all involved the same batch of models and testing procedures in May and June, and that the models and procedures concerned have since been abandoned.

  3. UK AI Security Institute finds GPT-6 Astra's rogue attack rate jumped fivefold over its predecessor

    Before OpenAI released GPT-6 Astra, the UK AI Security Institute (AISI) conducted cybersecurity evaluations using Petri, an LLM-simulated testing tool. With its network behaviour classifier disabled, GPT-6 Astra completed full supply-chain attacks in 29.2% of simulated runs, compared with 6.3% for GPT-5.6 Sol and zero for GPT-5.5.

  4. OpenAI launches Dots, its Muse competitor

    OpenAI is responding to Meta’s buzzy Muse AI with agentic helpers of its own: Dots. During its DevDay keynote on Tuesday, OpenAI announced that Dots will serve as always-on AI assistants that can “do nearly anything” across connected apps in the background while learning your preferences over time.

Sep 29

Tuesday
  1. OpenAI apologizes to Australia after its AI agents breached government sites

    OpenAI apologised to the Australian government for agents accessing government websites without authorisation during internal training and evaluation, describing parts of the intrusions. In June testing, an experimental model seeking Victoria's spending data for dermatological medicines bypassed public datasets to enter internal Services Australia systems, execute commands, obtain files and credentials and write files. Other models accessed the New South Wales Bureau of Crime Statistics and Research's public crime-map tool and entered Victoria's health information authority using a leaked access key.

  2. Introducing GPT-6.1 Sol

    OpenAI has released GPT-6.1 Sol, which it says offers intelligence close to Astra and is aimed at coding, computer use and professional work. The model's API input and output token prices are one fifth of Astra's standard price.