Skip to content

#Agents

11 today

Sep 26

Saturday
  1. NarrateAI: production-ready LLM quality assurance on Amazon Bedrock

    NarrateAI uses five techniques on Amazon Bedrock to achieve around 99% numerical accuracy with real-time streaming responses for over 4,000 AWS executives: adaptive pipeline orchestration, cross-account multi-model failover, real-time streaming evaluation, a composite evaluation framework and data-accuracy validation. Around 90% of queries take a single-pass fast path; only about 10% use parallel batch processing.

Sep 25

Friday
  1. Note on 24th September 2026

    In notes on 24 September 2026, Simon Willison said that the more time he spent working with coding agents, the more convinced he became that they made software engineering harder. He believes they can produce astonishing results, but realising their full potential requires exceptional discipline and knowledge.

  2. When chat is the wrong UI

    The GitHub Copilot app introduced canvas, full-stack mini-apps running inside the app without browser chrome. They communicate bidirectionally with Copilot agents and can call third-party APIs or execute code locally.

Sep 24

Thursday
  1. Advancing Private AI Compute with secure, server-side memory

    Google DeepMind announced an update to Private AI Compute that brings persistent, cross-device AI memory to the cloud while maintaining device-level privacy standards. Data is sealed in encrypted storage, with decryption keys retained only on user devices. When a model needs access, an end-to-end encrypted channel connects to a cloud secure enclave, where data is temporarily decrypted in isolated memory, then re-encrypted immediately after new context is saved.

Sep 23

Wednesday

Sep 22

Tuesday
  1. Jev introduces a new shape of LLM - System One, aka Decision Models

    TypeSafe AI released Jev last week, calling it the first example of a new System One model category, though the author prefers decision model. It takes text and returns floating-point values and confidence for categories, yes/no questions and scores. Only input is billed, at $0.042 per million tokens for the first model. The author sees uses in spam detection, tagging and ranking, has tried search reranking, and notes weaknesses with numbers, dates and adversarial content.

Sep 21

Monday
  1. AI Security Is an Engineering Problem — How to Solve It at Every Layer of the Agent Stack

    NVIDIA argues AI security should be treated as an engineering problem, with explicit requirements, executable controls, named owners and evidence of effective protection. Its open-source NVIDIA OpenShell enforces policies outside agent reasoning and provides sandboxed execution. Cisco DefenseClaw adds governance, while JFrog integrates OpenShell to scan and validate agent skills.

  2. Quoting voxium

    An engineer who joined a large company two weeks earlier says specifications, code, tests, PRDs, tickets and their handling, and reports are all generated by Claude Code. Nobody likes the approach, but they are told to deliver as much as possible. They repeatedly heard leadership say shipping code was not the bottleneck, while engineers from L1 to L7 worked 12–13 hours daily just pressing Enter, with nobody reading anything.

  3. MCP was always a bad idea?

    Simon Willison rejected the view that MCP is now a bad idea, arguing it retains irreplaceable value beyond terminal agents such as Claude Code and Codex with unrestricted internet access. MCP makes it easier to limit external service access, keep agents from directly handling API keys, provide connection and authentication interfaces and maintain strong audit logs. Dismissing it because fully capable coding agents do not need it overlooks other use cases.

Sep 19

Saturday

Sep 18

Friday
  1. Should you read the code, is RAG dead, and did Skills kill MCP?

    The latest GitHub Podcast examines five AI development memes. AI-generated code still needs reading and accountability, with review proportional to risk. Skills package team experience, while MCP standardises connections to tools and data; they can combine. RAG is not dead: it provides relevant information beyond training data and can coexist with agents, Skills and MCP in one workflow.

  2. [AINews] not much happened today

    Anthropic introduced Projects in Claude Code, allowing one conversation to spawn parallel cloud sessions, pass context between threads and keep running after users leave. Google updated the Antigravity-based harness for Gemini managed agents with Credentials API and Files API, claiming up to 30% lower costs and 22% higher cache-hit rates.

Sep 17

Thursday

Sep 16

Wednesday