The rate at which a model generates output tokens during the decode phase — the primary throughput metric users feel after the first token arrives.
The rate at which a model generates output tokens during the decode phase.
The rate at which a model generates output tokens during the decode phase — the primary throughput metric users feel after the first token arrives.