The limit is set by the model architecture, training, and serving configuration. Applications budget space among instructions, history, retrieved evidence, tool results, and the expected response.
Context length is the maximum number of tokens a model interface can process across its input and generated output.
The limit is set by the model architecture, training, and serving configuration. Applications budget space among instructions, history, retrieved evidence, tool results, and the expected response.