Economics and capital · Inference economics
IBM Research demonstrates llm-d serving throughput
IBM Research reported llm-d serving more than 6.6 million output tokens per minute on H100 GPUs.
Read the original at llm-d.aiOpens the publisher's site in a new tab