It may be measured in requests, examples, or tokens and usually improves with batching and parallel hardware. Maximizing throughput can increase individual request latency, so serving systems choose an operating tradeoff.
Throughput is the amount of inference or training work completed per unit of time.
It may be measured in requests, examples, or tokens and usually improves with batching and parallel hardware. Maximizing throughput can increase individual request latency, so serving systems choose an operating tradeoff.