← Back to tools

AI pricing calculator

Compare LLM API costs across major providers. Model monthly spend from request volume, context size, and generated output.

Model APIs / cost model

Select Use Case

Configuration

1,000 tokens

500 tokens

1,000 / day (30,000 / month)

Estimated Monthly Cost

Based on 30,000 requests/month with 1,000 input + 500 output tokens each

Gemini 1.5 Flash

Best Value

Google • Fast and efficient

Input: $2.25 | Output: $4.50 | Per request: <$0.01

$6.75

/month

GPT-4o mini

OpenAI • Affordable and intelligent

Input: $4.50 | Output: $9.00 | Per request: <$0.01

$13.50

/month

Mistral Small

Mistral AI • Cost-efficient for simple tasks

Input: $6.00 | Output: $9.00 | Per request: <$0.01

$15.00

/month

Claude 3 Haiku

Anthropic • Fastest and most cost-effective

Input: $7.50 | Output: $18.75 | Per request: <$0.01

$26.25

/month

Gemini 1.5 Pro

Google • Long context up to 2M tokens

Input: $37.50 | Output: $75.00 | Per request: <$0.01

$112.50

/month

Mistral Large

Mistral AI • Top-tier reasoning

Input: $60.00 | Output: $90.00 | Per request: <$0.01

$150.00

/month

GPT-4o

OpenAI • Flagship multimodal model

Input: $75.00 | Output: $150.00 | Per request: <$0.01

$225.00

/month

Claude 3.5 Sonnet

Anthropic • Best balance of speed and intelligence

Input: $90.00 | Output: $225.00 | Per request: $0.0105

$315.00

/month

GPT-4 Turbo

OpenAI • Previous flagship model

Input: $300.00 | Output: $450.00 | Per request: $0.0250

$750.00

/month

Claude 3 Opus

Anthropic • Most powerful for complex tasks

Input: $450.00 | Output: $1125.00 | Per request: $0.0525

$1575.00

/month

* Prices as of January 2025. Actual costs may vary. Does not include fine-tuning, image generation, or other specialized features.

Understanding AI API pricing

AI model pricing is typically based on tokens, which are pieces of text—roughly four characters or 0.75 words in English. Pricing is usually quoted per million tokens, with separate rates for input and output.

Key pricing factors

  • Input tokens—system prompts, user messages, and context sent to the model
  • Output tokens—the text the model generates in response
  • Model tier—more capable models generally cost more
  • Context window—larger windows admit more input and can increase costs

Cost optimization tips

  • Choose the right model—use smaller models for routine tasks and reserve more capable models for complex reasoning
  • Optimize prompts—concise, well-structured prompts reduce input token costs
  • Limit output length—set output limits when the task has a predictable response shape
  • Use caching—reuse stable prompt prefixes and common responses where providers support it
  • Batch requests—some providers offer discounted asynchronous processing

Provider comparison

Provider Strengths Best for
Anthropic Long context, coding, safety Complex analysis and code generation
OpenAI Ecosystem and multimodal models General-purpose and vision tasks
Google Large context and multimodal inputs Document processing and long content
Mistral AI Open weights and efficient models Cost-sensitive applications