Model pricing

756 models across 40 providers, per million tokens, synced from the public price feeds on 2026-08-07.

Every price table has input and output. 645 of these also carry the rates that decide the bill of anyone reusing context: the cost of WRITING to the prompt cache, the dearer one-hour cache entry, the separate reasoning rate, and the long-context threshold where a call reprices entirely. Those are the same numbers KostLens bills from.

openrouter

azure

bedrock

google

deepinfra

fireworks

xai

anthropic

novita

openai

gmi

moonshot

vercel

deepseek

hyperbolic

alibaba

ai21

minimax

tensormesh

zhipu

replicate

baseten

cloudflare

mistral

nebius

ovhcloud

snowflake

cerebras

nscale

sambanova

together

anyscale

cohere

darkbloom

groq

inception

lambda

meta

perplexity

scaleway