R1 Distill Llama 70B vs Granite 4.0 Micro: pricing comparison
Granite 4.0 Micro is the cheaper option at $0.02/$0.11 per 1M input/output tokens — R1 Distill Llama 70B ($0.80/$0.80) costs about 47x more per input token. Full spec-by-spec breakdown below.
| R1 Distill Llama 70B | Granite 4.0 Micro | |
|---|---|---|
| Input /1M tokens | $0.80 | $0.02 |
| Output /1M tokens | $0.80 | $0.11 |
| Cache read /1M | — | — |
| Context window | 8K | 131K |
| Provider | DeepSeek | Ibm-granite |
| Vision input | No | No |
| Released | Jan 2025 | Oct 2025 |
Real workload costs: R1 Distill Llama 70B vs Granite 4.0 Micro
Cost per single request at common token profiles — the cheapest model for each workload is highlighted.
| Workload | Tokens (in / out) | R1 Distill Llama 70B | Granite 4.0 Micro |
|---|---|---|---|
| Chatbot message | 500 / 300 | $0.0006 | <$0.0001 |
| RAG query with context | 4,000 / 500 | $0.0036 | $0.0001 |
| Document summarization | 20,000 / 1,000 | $0.02 | $0.0005 |
| Agent coding session | 100,000 / 5,000 | $0.08 | $0.0023 |
Frequently asked questions
Which is cheaper: R1 Distill Llama 70B vs Granite 4.0 Micro?
Granite 4.0 Micro is the cheapest on input tokens at $0.02 per 1M, and Granite 4.0 Micro is cheapest on output at $0.11 per 1M. R1 Distill Llama 70B costs about 47x more per input token than Granite 4.0 Micro.
How much does R1 Distill Llama 70B cost per 1M tokens?
R1 Distill Llama 70B costs $0.80 per 1M input tokens and $0.80 per 1M output tokens on the DeepSeek API, with a 8K-token context window.
How much does Granite 4.0 Micro cost per 1M tokens?
Granite 4.0 Micro costs $0.02 per 1M input tokens and $0.11 per 1M output tokens on the Ibm-granite API, with a 131K-token context window.
What does a typical request cost on R1 Distill Llama 70B vs Granite 4.0 Micro?
For a typical request with 2,000 input and 500 output tokens: R1 Distill Llama 70B costs $0.0020, Granite 4.0 Micro costs <$0.0001. At 1,000 requests/day for a month that is $60.00 for R1 Distill Llama 70B vs $2.70 for Granite 4.0 Micro.
Estimate your own workload with the per-model calculators (R1 Distill Llama 70B, Granite 4.0 Micro), browse all current prices on the live model pricing dashboard, or run these models head-to-head on a real prompt (free account required). Prices may differ from provider list prices for batch or tiered usage.
Pricing data via the OpenRouter models API.