Skip to content

Deployment

The engineering of running models: inference optimisation, memory and cost, serving architecture and infrastructure choices.

Latest selected

1–20 of 42

Sep 30

Wednesday
  1. Build a multi-agent music production pipeline on Amazon Bedrock AgentCore Runtime Instances

    The official AWS blog demonstrates deploying a three-agent music production pipeline on Amazon Bedrock AgentCore Runtime Instances. The composition agent uses Claude Sonnet 4.6 to generate a music brief and runs ACE-Step on the instance's NVIDIA L4 to render audio. The delivery agent reads .wav files from the shared volume for measurement and DSP processing, while the compliance agent independently remeasures them and checks harmonic similarity against a music catalogue.

Sep 29

Tuesday

Sep 24

Thursday

Sep 22

Tuesday

Sep 17

Thursday

Sep 16

Wednesday
  1. From Megawatts to Tokens: How NVIDIA Maximizes AI Factory Production

    NVIDIA introduced DSX at AI Infra Summit to improve AI-factory token throughput within a fixed power budget. Lambda validated DSX MaxLPS on a five-rack, 19-node HGX B200 cluster, running 19 nodes within the same budget as 16 at full power. Cluster token throughput rose 24%, from around four million to five million tokens per second, with performance per watt up 23%.

Sep 5

Saturday

Sep 1

Tuesday

Aug 21

Friday

Aug 14

Friday

Aug 11

Tuesday

Jul 8

Wednesday

Jul 7

Tuesday