BLAST RADIUS

Economics and capital · Inference economics

IBM Research demonstrates llm-d serving throughput

Sep 8, 2026

IBM Research reported llm-d serving more than 6.6 million output tokens per minute on H100 GPUs.

Read the original at llm-d.aiOpens the publisher's site in a new tab

More in Economics and capital