The maximum number of tokens a model can process in a single request, including both input and output — exceeding it requires truncation, summarization, or RAG.
The maximum number of tokens a model can process in a single request.
The maximum number of tokens a model can process in a single request, including both input and output — exceeding it requires truncation, summarization, or RAG.