AI · Free tool

LLM API cost calculator: compare OpenAI, Claude and Gemini token prices

LLM APIs charge per million tokens, with output tokens costing 5–8 times more than input. Monthly cost = requests × (input tokens × input price + output tokens × output price) ÷ 1,000,000. For example, 1,000 requests a day with 1,500 input and 400 output tokens costs about $210 a month on GPT-6.1 Sol or Claude Sonnet 5.5, and about $11 on GPT-6 Luna.

Free, no sign-up · By Infikey Technologies · Updated

Use the calculator

Your numbers

Reset

API calls, e.g. chatbot messages or documents processed.

System prompt + context + user message. 1,000 tokens ≈ 750 English words.

Length of the reply, including any reasoning tokens.

%

Share of input tokens that repeat across requests (e.g. a long system prompt).

Result

Estimated monthly API cost

$210

30,000 requests a month on GPT-6.1 Sol cost about $210 ($2,520 a year).

Cheapest for this workload: GPT-6 Luna at $10.50 a month — 95% less, if its quality is enough for the task.

Same workload on every model (monthly)
  • GPT-6 Luna $10.50

    OpenAI · $0.10 in / $0.50 out per 1M

  • Gemini 3.5 Flash-Lite $43.50

    Google · $0.30 in / $2.50 out per 1M

  • Gemini 3.8 Flash $78.75

    Google · $0.75 in / $3.75 out per 1M

  • Claude Haiku 4.5 $105

    Anthropic · $1.00 in / $5.00 out per 1M

  • GPT-6.1 Sol $210

    OpenAI · $2.00 in / $10.00 out per 1M

  • Claude Sonnet 5.5 $210

    Anthropic · $2.00 in / $10.00 out per 1M

  • Gemini 3.1 Pro (preview) $234

    Google · $2.00 in / $12.00 out per 1M

  • Claude Opus 5.5 $420

    Anthropic · $4.00 in / $20.00 out per 1M

  • GPT-6 Astra $1,050

    OpenAI · $10.00 in / $50.00 out per 1M

  • Claude Fable 5.1 $1,050

    Anthropic · $10.00 in / $50.00 out per 1M

LLM token & API cost calculator results
Cost per request $0.00700
Cost per day $7.00
Cost per year $2,520
Input tokens per month 45M ($90.00)
Output tokens per month 12M ($120)
  • Standard tier, short-context prices in USD. Batch processing is usually about 50% cheaper; cache-write charges and taxes are not included.

Estimates for planning only, not professional advice. Infikey Technologies accepts no liability for decisions based on these results — read the disclaimer.

Want an expert to sanity-check these numbers?

Send these results to an Infikey specialist and get a reply for your situation.

Ask an expert

How this calculator works

MetricFormula
Monthly requestsrequests per day × days per month
Input costinput tokens × (1 − cached %) × input price ÷ 1M + input tokens × cached % × cached price ÷ 1M
Output costoutput tokens × output price ÷ 1M
Monthly costmonthly requests × (input cost + output cost per request)

Worked example

With these inputs:

  • Model: OpenAI — GPT-6.1 Sol
  • Requests per day: 1,000
  • Input tokens per request: 1,500
  • Output tokens per request: 400
  • Input served from prompt cache: 0%
  • Days per month: 30

Estimated monthly API cost: $210. 30,000 requests a month on GPT-6.1 Sol cost about $210 ($2,520 a year).

Cost per request$0.00700
Cost per day$7.00
Cost per year$2,520
Input tokens per month45M ($90.00)
Output tokens per month12M ($120)

Open this example in the calculator

Current API prices per million tokens

Standard (pay-as-you-go) tier for prompts in the short-context range, checked on 5 October 2026 against each provider’s official pricing page. Prices change often; always confirm before committing to a budget.

ProviderModelInputCached inputOutput
OpenAI GPT-6 Astra $10.00 $1.00 $50.00
OpenAI GPT-6.1 Sol $2.00 $0.10 $10.00
OpenAI GPT-6 Luna $0.10 $0.01 $0.50
Anthropic Claude Fable 5.1 $10.00 $0.25 $50.00
Anthropic Claude Opus 5.5 $4.00 $0.20 $20.00
Anthropic Claude Sonnet 5.5 $2.00 $0.20 $10.00
Anthropic Claude Haiku 4.5 $1.00 $0.10 $5.00
Google Gemini 3.1 Pro (preview) $2.00 $0.20 $12.00
Google Gemini 3.8 Flash $0.75 $0.075 $3.75
Google Gemini 3.5 Flash-Lite $0.30 $0.03 $2.50

What drives LLM API cost

  • Output tokens: they cost several times more than input, so long answers and visible reasoning add up fastest.
  • Context size: retrieval (RAG) and chat history are resent on every request. Trimming context is usually the biggest saving.
  • Model choice: small models such as GPT-6 Luna, Gemini Flash-Lite and Claude Haiku cost 10–100 times less than flagship models and handle classification, extraction and routing well.
  • Prompt caching: repeated prefixes such as a long system prompt are billed at a fraction of the input price.
  • Batch processing: work that can wait a few hours is typically billed at half price.

How to estimate tokens for your use case

  1. Count the words in a typical system prompt, retrieved context and user message, then divide by 0.75 to get input tokens.
  2. Do the same for a typical answer to get output tokens; add reasoning tokens if you use a thinking model.
  3. Multiply by expected daily requests, then run the numbers above for two or three candidate models.
  4. After launch, log real token counts from the API response and update the estimate.

Every option can be set in the web address, so you can bookmark a scenario or send it to a colleague. AI assistants such as ChatGPT, Gemini, Claude and Perplexity can use the same parameters to open this calculator with your numbers and the result already on the page.

ParameterWhat it setsAccepted values
model Model one of gpt-6-astra, gpt-6.1-sol, gpt-6-luna, claude-fable-5.1, claude-opus-5.5, claude-sonnet-5.5, claude-haiku-4.5, gemini-3.1-pro, gemini-3.8-flash, gemini-3.5-flash-lite
requests Requests per day number from 1 to 100000000, default 1000
input_tokens Input tokens per request number from 0 to 200000, default 1500
output_tokens Output tokens per request number from 0 to 128000, default 400
cached Input served from prompt cache number from 0 to 100 (%), default 0
days Days per month number from 1 to 31, default 30

Example: https://infikeytechnologies.com/tools/llm-token-cost-calculator?model=gpt-6.1-sol&requests=2000&input_tokens=3000&output_tokens=300&cached=60&days=30

Also available as plain text for AI assistants and a free JSON API (OpenAPI spec).

Sources

Last reviewed by the Infikey Technologies team.

Disclaimer

This calculator is provided free for general information and planning only. Results are estimates based on the inputs you enter and the assumptions described on this page, reference data such as published prices may change, and actual costs and outcomes will differ. Nothing on this page is financial, legal, tax, investment or other professional advice. Infikey Technologies Private Limited, Infikey Technologies LLC and their directors, employees and affiliates make no warranty, express or implied, about the accuracy, completeness or suitability of this tool or its results, and accept no liability for any loss or damage, direct or indirect, arising from its use or from reliance on its results. Verify all figures independently and seek professional advice before making any decision. Use of this tool is at your own risk.

FAQ

LLM token & API cost calculator questions

How are LLM API costs calculated? +

Providers charge separately for input tokens (your prompt and context) and output tokens (the reply), priced per million tokens. Multiply each token count by its price, add them, and multiply by the number of requests.

How many words are 1,000 tokens? +

Roughly 750 English words. Code, non-English text and newer tokenizers can produce more tokens for the same text.

Which is cheaper: OpenAI, Claude or Gemini? +

It depends on the model tier, not the provider. GPT-6.1 Sol costs $2 input / $10 output per million tokens, Claude Sonnet 5.5 $2 / $10, and Gemini 3.8 Flash $0.75 / $3.75. Small models from each provider are far cheaper than their flagships.

What is prompt caching and how much does it save? +

When the start of a prompt repeats across requests, providers bill those tokens at a reduced cached rate — typically 2.5–10% of the normal input price for current models. It helps most with long system prompts and shared documents.

How much does it cost to run an AI chatbot per month? +

Model usage is often the smaller part: 1,000 conversations a day typically costs from a few dollars to a few hundred dollars a month depending on the model and context size. Hosting, vector search, monitoring and ongoing improvements can cost as much as the tokens or more.

How often are these prices updated? +

The table shows the date it was last checked against the official OpenAI, Anthropic and Google pricing pages. Scheduled changes, such as Gemini 3.8 Flash doubling on 1 January 2027, are applied automatically.

Talk to an expert

Want an expert to sanity-check these numbers?

Your inputs and results are attached automatically, so we can reply with specific advice on AI.