Together, Fireworks, Groq, and the hyperscalers sell tokens on someone else's (or their own) weights. Price and latency depend on batching, quantization, and GPU supply, not just the model name. Switching providers can change tokenizer behavior, tool-calling quirks, and rate limits even on 'the same' model.