For generated text, systems often separate time to first token from time per subsequent token. Queueing, input length, model compute, network transfer, and decoding all contribute.
Latency is the elapsed time between starting a request or operation and receiving its result.
For generated text, systems often separate time to first token from time per subsequent token. Queueing, input length, model compute, network transfer, and decoding all contribute.