# LLM API cost calculator: compare OpenAI, Claude and Gemini token prices

> LLM APIs charge per million tokens, with output tokens costing 5–8 times more than input. Monthly cost = requests × (input tokens × input price + output tokens × output price) ÷ 1,000,000. For example, 1,000 requests a day with 1,500 input and 400 output tokens costs about $210 a month on GPT-6.1 Sol or Claude Sonnet 5.5, and about $11 on GPT-6 Luna.

- Interactive version: https://infikeytechnologies.com/tools/llm-token-cost-calculator
- Type: free instant calculator
- Category: AI
- Last reviewed: 2026-10-05
- Publisher: Infikey Technologies (https://infikeytechnologies.com)

## Example result (default inputs)

| Input | Value |
| --- | --- |
| Model | OpenAI — GPT-6.1 Sol |
| Requests per day | 1,000 |
| Input tokens per request | 1,500 |
| Output tokens per request | 400 |
| Input served from prompt cache | 0% |
| Days per month | 30 |

**Estimated monthly API cost: $210.** 30,000 requests a month on GPT-6.1 Sol cost about $210 ($2,520 a year).

Cheapest for this workload: GPT-6 Luna at $10.50 a month — 95% less, if its quality is enough for the task.

| Metric | Value |
| --- | --- |
| Cost per request | $0.00700 |
| Cost per day | $7.00 |
| Cost per year | $2,520 |
| Input tokens per month | 45M ($90.00) |
| Output tokens per month | 12M ($120) |

### Same workload on every model (monthly)

| Item | Value | Detail |
| --- | --- | --- |
| GPT-6 Luna | $10.50 | OpenAI · $0.10 in / $0.50 out per 1M |
| Gemini 3.5 Flash-Lite | $43.50 | Google · $0.30 in / $2.50 out per 1M |
| Gemini 3.8 Flash | $78.75 | Google · $0.75 in / $3.75 out per 1M |
| Claude Haiku 4.5 | $105 | Anthropic · $1.00 in / $5.00 out per 1M |
| GPT-6.1 Sol (selected) | $210 | OpenAI · $2.00 in / $10.00 out per 1M |
| Claude Sonnet 5.5 | $210 | Anthropic · $2.00 in / $10.00 out per 1M |
| Gemini 3.1 Pro (preview) | $234 | Google · $2.00 in / $12.00 out per 1M |
| Claude Opus 5.5 | $420 | Anthropic · $4.00 in / $20.00 out per 1M |
| GPT-6 Astra | $1,050 | OpenAI · $10.00 in / $50.00 out per 1M |
| Claude Fable 5.1 | $1,050 | Anthropic · $10.00 in / $50.00 out per 1M |

- Standard tier, short-context prices in USD. Batch processing is usually about 50% cheaper; cache-write charges and taxes are not included.

Open this result on the website: https://infikeytechnologies.com/tools/llm-token-cost-calculator?model=gpt-6.1-sol&requests=1000&input_tokens=1500&output_tokens=400&cached=0&days=30

## Use from a link or API

Add these query parameters to https://infikeytechnologies.com/tools/llm-token-cost-calculator (pre-filled page), https://infikeytechnologies.com/tools/llm-token-cost-calculator.md (this plain-text page) or https://infikeytechnologies.com/api/tools/llm-token-cost-calculator (JSON).

| Parameter | Meaning | Accepted values |
| --- | --- | --- |
| `model` | Model | one of gpt-6-astra, gpt-6.1-sol, gpt-6-luna, claude-fable-5.1, claude-opus-5.5, claude-sonnet-5.5, claude-haiku-4.5, gemini-3.1-pro, gemini-3.8-flash, gemini-3.5-flash-lite |
| `requests` | Requests per day | number from 1 to 100000000, default 1000 |
| `input_tokens` | Input tokens per request | number from 0 to 200000, default 1500 |
| `output_tokens` | Output tokens per request | number from 0 to 128000, default 400 |
| `cached` | Input served from prompt cache | number from 0 to 100 (%), default 0 |
| `days` | Days per month | number from 1 to 31, default 30 |

- Support chatbot with RAG: https://infikeytechnologies.com/tools/llm-token-cost-calculator?model=gpt-6.1-sol&requests=2000&input_tokens=3000&output_tokens=300&cached=60&days=30
- Document summaries: https://infikeytechnologies.com/tools/llm-token-cost-calculator?model=gpt-6.1-sol&requests=500&input_tokens=12000&output_tokens=800&cached=0&days=30
- Classification at scale: https://infikeytechnologies.com/tools/llm-token-cost-calculator?model=gpt-6.1-sol&requests=50000&input_tokens=400&output_tokens=20&cached=50&days=30
- Coding assistant: https://infikeytechnologies.com/tools/llm-token-cost-calculator?model=gpt-6.1-sol&requests=3000&input_tokens=20000&output_tokens=1500&cached=70&days=30

## How it is calculated

- **Monthly requests**: requests per day × days per month
- **Input cost**: input tokens × (1 − cached %) × input price ÷ 1M + input tokens × cached % × cached price ÷ 1M
- **Output cost**: output tokens × output price ÷ 1M
- **Monthly cost**: monthly requests × (input cost + output cost per request)

## Current API prices per million tokens

Standard (pay-as-you-go) tier for prompts in the short-context range, checked on 5 October 2026 against each provider’s official pricing page. Prices change often; always confirm before committing to a budget.

| Provider | Model | Input | Cached input | Output |
| --- | --- | --- | --- | --- |
| OpenAI | GPT-6 Astra | $10.00 | $1.00 | $50.00 |
| OpenAI | GPT-6.1 Sol | $2.00 | $0.10 | $10.00 |
| OpenAI | GPT-6 Luna | $0.10 | $0.01 | $0.50 |
| Anthropic | Claude Fable 5.1 | $10.00 | $0.25 | $50.00 |
| Anthropic | Claude Opus 5.5 | $4.00 | $0.20 | $20.00 |
| Anthropic | Claude Sonnet 5.5 | $2.00 | $0.20 | $10.00 |
| Anthropic | Claude Haiku 4.5 | $1.00 | $0.10 | $5.00 |
| Google | Gemini 3.1 Pro (preview) | $2.00 | $0.20 | $12.00 |
| Google | Gemini 3.8 Flash | $0.75 | $0.075 | $3.75 |
| Google | Gemini 3.5 Flash-Lite | $0.30 | $0.03 | $2.50 |

## What drives LLM API cost

- Output tokens: they cost several times more than input, so long answers and visible reasoning add up fastest.
- Context size: retrieval (RAG) and chat history are resent on every request. Trimming context is usually the biggest saving.
- Model choice: small models such as GPT-6 Luna, Gemini Flash-Lite and Claude Haiku cost 10–100 times less than flagship models and handle classification, extraction and routing well.
- Prompt caching: repeated prefixes such as a long system prompt are billed at a fraction of the input price.
- Batch processing: work that can wait a few hours is typically billed at half price.

## How to estimate tokens for your use case

1. Count the words in a typical system prompt, retrieved context and user message, then divide by 0.75 to get input tokens.
2. Do the same for a typical answer to get output tokens; add reasoning tokens if you use a thinking model.
3. Multiply by expected daily requests, then run the numbers above for two or three candidate models.
4. After launch, log real token counts from the API response and update the estimate.

## FAQ

### How are LLM API costs calculated?

Providers charge separately for input tokens (your prompt and context) and output tokens (the reply), priced per million tokens. Multiply each token count by its price, add them, and multiply by the number of requests.

### How many words are 1,000 tokens?

Roughly 750 English words. Code, non-English text and newer tokenizers can produce more tokens for the same text.

### Which is cheaper: OpenAI, Claude or Gemini?

It depends on the model tier, not the provider. GPT-6.1 Sol costs $2 input / $10 output per million tokens, Claude Sonnet 5.5 $2 / $10, and Gemini 3.8 Flash $0.75 / $3.75. Small models from each provider are far cheaper than their flagships.

### What is prompt caching and how much does it save?

When the start of a prompt repeats across requests, providers bill those tokens at a reduced cached rate — typically 2.5–10% of the normal input price for current models. It helps most with long system prompts and shared documents.

### How much does it cost to run an AI chatbot per month?

Model usage is often the smaller part: 1,000 conversations a day typically costs from a few dollars to a few hundred dollars a month depending on the model and context size. Hosting, vector search, monitoring and ongoing improvements can cost as much as the tokens or more.

### How often are these prices updated?

The table shows the date it was last checked against the official OpenAI, Anthropic and Google pricing pages. Scheduled changes, such as Gemini 3.8 Flash doubling on 1 January 2027, are applied automatically.

## Sources

- [OpenAI API pricing](https://platform.openai.com/docs/pricing)
- [Anthropic Claude pricing](https://docs.claude.com/en/docs/about-claude/pricing)
- [Google Gemini API pricing](https://ai.google.dev/gemini-api/docs/pricing)

## Get expert help

Send these results to an Infikey Technologies specialist from the form on https://infikeytechnologies.com/tools/llm-token-cost-calculator#estimate or via https://infikeytechnologies.com/contact.

More free calculators: https://infikeytechnologies.com/tools.md

## Disclaimer

This calculator is provided free for general information and planning only. Results are estimates based on the inputs you enter and the assumptions described on this page, reference data such as published prices may change, and actual costs and outcomes will differ. Nothing on this page is financial, legal, tax, investment or other professional advice. Infikey Technologies Private Limited, Infikey Technologies LLC and their directors, employees and affiliates make no warranty, express or implied, about the accuracy, completeness or suitability of this tool or its results, and accept no liability for any loss or damage, direct or indirect, arising from its use or from reliance on its results. Verify all figures independently and seek professional advice before making any decision. Use of this tool is at your own risk.
