Global model marketplace

Every TokensChain model,
one searchable view.

Live metadata for 300+ TokensChain models — price, context, modality, all visible.

400 models
inclusionAI: Ling 3.0 Tiny (free)
inclusionai/ling-3.0-tiny:free

Ling 3.0 Tiny is a mixture-of-experts model from InclusionAI, with 1.3B active parameters out of 7.9B total. It is designed for responsive agents, instruction following, and multi-turn conversations, with switchable...

Ctx
262K
In
Free
Out
Free
Meta: Muse Spark 1.2
meta/muse-spark-1.2

Muse Spark 1.2 is a reasoning model from Meta, designed for complex agentic tasks. It accepts text, images, video, audio, and PDF documents, returns text, and offers a 1M-token context...

Ctx
1049K
In
$1.25 / 1M
Out
$4.25 / 1M
Qwen: Qwen3.8 Max
qwen/qwen3.8-max

Qwen3.8 Max is the flagship model in Alibaba's Qwen3.8 series, the general-availability successor to the Qwen3.8 Max Preview. It is a multimodal reasoning model intended for complex reasoning, visual understanding,...

Ctx
1000K
In
$2.00 / 1M
Out
$6.00 / 1M
DeepSeek V4 Flash Latest
~deepseek/deepseek-v4-flash-latest

This model always redirects to the latest model in the DeepSeek V4 Flash family.

Ctx
1049K
In
$0.090 / 1M
Out
$0.179 / 1M
DeepSeek: DeepSeek V4 Flash 0731
deepseek/deepseek-v4-flash-0731

DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows.

Ctx
1049K
In
$0.090 / 1M
Out
$0.180 / 1M
Thinking Machines: Inkling Small
thinkingmachines/inkling-small

Inkling Small is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 12B active parameters out of 276B total. It is positioned as the smaller, more efficient member of...

Ctx
524K
In
$0.450 / 1M
Out
$1.20 / 1M
Qwen: Qwen3.7 Flash
qwen/qwen3.7-flash

Qwen3.7 Flash is a vision-language reasoning model from Alibaba. It is suited for multimodal agents, visual coding, search, and computer interaction, with strengths in object recognition, spatial understanding, and real-world...

Ctx
1000K
In
$0.030 / 1M
Out
$0.130 / 1M
Claude Opus 5 (Fast)
anthropic/claude-opus-5-fast

Fast-mode variant of [Opus 5](/anthropic/claude-opus-5) - identical capabilities with higher output speed at 2x pricing relative to regular Opus 5. Learn more in Anthropic's docs: https://platform.claude.com/docs/en/build-with-claude/fast-mode

Ctx
1000K
In
$10.00 / 1M
Out
$50.00 / 1M
Claude Opus 5
anthropic/claude-opus-5

Claude Opus 5 is Anthropic’s flagship model for demanding reasoning, coding, and long-horizon agentic work. It is particularly strong at end-to-end software tasks, code review and bug finding, visual analysis...

Ctx
1000K
In
$5.00 / 1M
Out
$25.00 / 1M
Claude Opus 5 (batch)
anthropic/claude-opus-5:batch

Claude Opus 5 is Anthropic’s flagship model for demanding reasoning, coding, and long-horizon agentic work. It is particularly strong at end-to-end software tasks, code review and bug finding, visual analysis...

Ctx
1000K
In
$2.50 / 1M
Out
$12.50 / 1M
Ling-3.0-flash
inclusionai/ling-3.0-flash

*Ling-3.0-flash* is a *124B-parameter Mixture-of-Experts (MoE) model*, with approximately *5.1B parameters activated per token*. The model is designed with *token efficiency and production-scale agentic inference* as key priorities, enabling developers...

Ctx
262K
In
$0.021 / 1M
Out
$0.063 / 1M
Poolside: Laguna S 2.1
poolside/laguna-s-2.1

Laguna S 2.1 is the latest coding agent model from [Poolside](<https://poolside.ai/>). Laguna S 2.1 is a 118B total parameter model with 8B active parameters, scoring 70.2% on Terminal-Bench 2.1 and...

Ctx
1049K
In
$0.090 / 1M
Out
$0.180 / 1M
Poolside: Laguna S 2.1 (free)
poolside/laguna-s-2.1:free

Laguna S 2.1 is the latest coding agent model from [Poolside](<https://poolside.ai/>). Laguna S 2.1 is a 118B total parameter model with 8B active parameters, scoring 70.2% on Terminal-Bench 2.1 and...

Ctx
262K
In
Free
Out
Free
Google: Gemini 3.6 Flash
google/gemini-3.6-flash

Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to produce polished outputs with fewer unnecessary edits and...

Ctx
1049K
In
$1.50 / 1M
Out
$7.50 / 1M
Google: Gemini 3.6 Flash (batch)
google/gemini-3.6-flash:batch

Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to produce polished outputs with fewer unnecessary edits and...

Ctx
1049K
In
$0.750 / 1M
Out
$3.75 / 1M
Google: Gemini 3.5 Flash Lite
google/gemini-3.5-flash-lite

Gemini 3.5 Flash Lite is a high-efficiency model from Google with upgraded agentic capabilities. It is suited for subagents that execute focused tasks within complex, multi-agent workflows.

Ctx
1049K
In
$0.300 / 1M
Out
$2.50 / 1M
Google: Gemini 3.5 Flash Lite (batch)
google/gemini-3.5-flash-lite:batch

Gemini 3.5 Flash Lite is a high-efficiency model from Google with upgraded agentic capabilities. It is suited for subagents that execute focused tasks within complex, multi-agent workflows.

Ctx
1049K
In
$0.150 / 1M
Out
$1.25 / 1M
Meituan: LongCat 2.0
meituan/longcat-2.0

LongCat 2.0 is a sparse mixture-of-experts language model from Meituan, with 48B active parameters out of 1.6T total. It is suited for coding, repository-level changes, long-horizon problem solving, and agentic...

Ctx
1049K
In
$0.300 / 1M
Out
$1.20 / 1M
Thinking Machines: Inkling
thinkingmachines/inkling

Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning, coding, agentic and tool-use systems,...

Ctx
1049K
In
$0.950 / 1M
Out
$4.05 / 1M
Thinking Machines: Inkling (batch)
thinkingmachines/inkling:batch

Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning, coding, agentic and tool-use systems,...

Ctx
524K
In
$0.500 / 1M
Out
$2.02 / 1M
Auto Router (Beta)
openrouter/auto-beta

Auto Router (Beta) is a task-aware router from OpenRouter. It classifies each request, then routes it the [most popular model](/rankings#task-spend) for that task based on aggregate spend, filtered by your...

Ctx
2000K
In
$-1000000.000 / 1M
Out
$-1000000.000 / 1M
MoonshotAI: Kimi K3
moonshotai/kimi-k3

Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic workflows, and is particularly strong at...

Ctx
1049K
In
$3.00 / 1M
Out
$15.00 / 1M
Meta: Muse Spark 1.1
meta/muse-spark-1.1

Muse Spark 1.1 is a multimodal reasoning model from Meta, built for agentic tasks. It accepts text, images, video, audio, and PDF documents and returns text, with a 1M-token context...

Ctx
1049K
In
$1.25 / 1M
Out
$4.25 / 1M
Kwaipilot: KAT-Coder-Air V2.5
kwaipilot/kat-coder-air-v2.5

KAT-Coder-Air V2.5 is a flagship-level Agentic Coding model that can directly hand over an entire issue or an entire business workflow to it, allowing it to autonomously locate and make...

Ctx
256K
In
$0.150 / 1M
Out
$0.600 / 1M
Kwaipilot: KAT-Coder-Pro V2.5
kwaipilot/kat-coder-pro-v2.5

KAT-Coder-Pro V2.5 is a flagship-level Agentic Coding model that can directly hand over an entire issue or an entire business workflow to it, allowing it to autonomously locate and make...

Ctx
256K
In
$0.740 / 1M
Out
$2.96 / 1M
OpenAI: GPT-5.6 Luna Pro
openai/gpt-5.6-luna-pro

GPT-5.6 Luna Pro is the same underlying model as [GPT-5.6 Luna](https://openrouter.ai/openai/gpt-5.6-luna), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode

Ctx
1050K
In
$0.100 / 1M
Out
$0.600 / 1M
OpenAI: GPT-5.6 Luna Pro (batch)
openai/gpt-5.6-luna-pro:batch

GPT-5.6 Luna Pro is the same underlying model as [GPT-5.6 Luna](https://openrouter.ai/openai/gpt-5.6-luna), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode

Ctx
1050K
In
$0.100 / 1M
Out
$0.600 / 1M
OpenAI: GPT-5.6 Luna
openai/gpt-5.6-luna

GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series. It is suited for high-volume, latency-sensitive tasks such as chat, classification, and lightweight agentic workflows, providing capable reasoning for...

Ctx
1050K
In
$0.100 / 1M
Out
$0.600 / 1M
OpenAI: GPT-5.6 Luna (batch)
openai/gpt-5.6-luna:batch

GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series. It is suited for high-volume, latency-sensitive tasks such as chat, classification, and lightweight agentic workflows, providing capable reasoning for...

Ctx
1050K
In
$0.100 / 1M
Out
$0.600 / 1M
OpenAI: GPT-5.6 Terra Pro
openai/gpt-5.6-terra-pro

GPT-5.6 Terra Pro is the same underlying model as [GPT-5.6 Terra](https://openrouter.ai/openai/gpt-5.6-terra), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode

Ctx
1050K
In
$1.00 / 1M
Out
$6.00 / 1M
OpenAI: GPT-5.6 Terra Pro (batch)
openai/gpt-5.6-terra-pro:batch

GPT-5.6 Terra Pro is the same underlying model as [GPT-5.6 Terra](https://openrouter.ai/openai/gpt-5.6-terra), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode

Ctx
1050K
In
$1.00 / 1M
Out
$6.00 / 1M
OpenAI: GPT-5.6 Terra
openai/gpt-5.6-terra

GPT-5.6 Terra is a balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier. It is suited for everyday coding, reasoning, and agentic...

Ctx
1050K
In
$1.00 / 1M
Out
$6.00 / 1M
OpenAI: GPT-5.6 Terra (batch)
openai/gpt-5.6-terra:batch

GPT-5.6 Terra is a balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier. It is suited for everyday coding, reasoning, and agentic...

Ctx
1050K
In
$1.00 / 1M
Out
$6.00 / 1M
OpenAI: GPT-5.6 Sol Pro
openai/gpt-5.6-sol-pro

GPT-5.6 Sol Pro is the same underlying model as [GPT-5.6 Sol](https://openrouter.ai/openai/gpt-5.6-sol), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode

Ctx
1050K
In
$5.00 / 1M
Out
$30.00 / 1M
OpenAI: GPT-5.6 Sol Pro (batch)
openai/gpt-5.6-sol-pro:batch

GPT-5.6 Sol Pro is the same underlying model as [GPT-5.6 Sol](https://openrouter.ai/openai/gpt-5.6-sol), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode

Ctx
1050K
In
$2.50 / 1M
Out
$15.00 / 1M
OpenAI: GPT-5.6 Sol
openai/gpt-5.6-sol

GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series. It is suited for complex reasoning, coding, and agentic workflows, and is particularly strong at command-line and multi-step coding tasks...

Ctx
1050K
In
$5.00 / 1M
Out
$30.00 / 1M
OpenAI: GPT-5.6 Sol (batch)
openai/gpt-5.6-sol:batch

GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series. It is suited for complex reasoning, coding, and agentic workflows, and is particularly strong at command-line and multi-step coding tasks...

Ctx
1050K
In
$2.50 / 1M
Out
$15.00 / 1M
SpaceXAI: Grok 4.5
x-ai/grok-4.5

Grok 4.5 is SpaceXAI's smartest model with frontier performance on coding, knowledge work, and STEM.

Ctx
500K
In
$2.00 / 1M
Out
$6.00 / 1M
xAI: Grok Latest
~x-ai/grok-latest

This model always redirects to the latest Grok model from xAI.

Ctx
500K
In
$2.00 / 1M
Out
$6.00 / 1M
AionLabs: Aion-3.0-Mini
aion-labs/aion-3.0-mini

Aion-3.0 Mini is a multi-model roleplaying and storytelling system from AionLabs, built on the DeepSeek family of models. It uses a collaborative generation process in which multiple specialized models each...

Ctx
131K
In
$0.700 / 1M
Out
$1.40 / 1M
AionLabs: Aion-3.0
aion-labs/aion-3.0

Aion-3.0 is a multi-model roleplaying and storytelling system from AionLabs, built on the GLM family of models. It uses a collaborative generation process in which multiple specialized models each contribute...

Ctx
131K
In
$3.00 / 1M
Out
$6.00 / 1M
Tencent: Hy3
tencent/hy3

Hy3 is a 295B-parameter Mixture-of-Experts model from Tencent (21B active, 192 experts with top-8 routing) built for reasoning, agentic workflows, and real-world production use. It supports a configurable reasoning effort:...

Ctx
262K
In
$0.132 / 1M
Out
$0.528 / 1M
Poolside: Laguna XS 2.1
poolside/laguna-xs-2.1

Laguna XS 2.1 is the latest coding agent model in the 33B-A3B category from [Poolside](https://poolside.ai/) and a step forward from their Laguna XS.2 model (released in April 2026). It combines...

Ctx
262K
In
$0.060 / 1M
Out
$0.120 / 1M
Poolside: Laguna XS 2.1 (free)
poolside/laguna-xs-2.1:free

Laguna XS 2.1 is the latest coding agent model in the 33B-A3B category from [Poolside](https://poolside.ai/) and a step forward from their Laguna XS.2 model (released in April 2026). It combines...

Ctx
262K
In
Free
Out
Free
Anthropic: Claude Sonnet 5
anthropic/claude-sonnet-5

Sonnet 5 is Anthropic's most capable Sonnet-class model, with frontier performance across coding, agents, and professional work. It supports adaptive thinking with selectable reasoning effort levels (low, medium, high, max,...

Ctx
1000K
In
$2.00 / 1M
Out
$10.00 / 1M
Anthropic: Claude Sonnet 5 (batch)
anthropic/claude-sonnet-5:batch

Sonnet 5 is Anthropic's most capable Sonnet-class model, with frontier performance across coding, agents, and professional work. It supports adaptive thinking with selectable reasoning effort levels (low, medium, high, max,...

Ctx
1000K
In
$1.00 / 1M
Out
$5.00 / 1M
Google: Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image)
google/gemini-3.1-flash-lite-image

Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image) is Google's fastest, most cost-efficient Gemini image model, built for high-velocity developer pipelines and rapid-fire visual exploration. It delivers text-to-image generation...

Ctx
66K
In
$0.250 / 1M
Out
$1.50 / 1M
Nex AGI: Nex-N2-Mini
nex-agi/nex-n2-mini

Nex-N2-Mini is an open-source agentic mixture-of-experts model from Nex AGI, the smaller sibling in the Nex-N2 series. It accepts text and image input and is built for coding, tool use,...

Ctx
262K
In
$0.025 / 1M
Out
$0.100 / 1M
Sakana: Fugu Ultra
sakana/fugu-ultra

Fugu Ultra is the higher-performance model in Sakana AI's Fugu family. Rather than a single monolithic model, Fugu is a learned multi-agent orchestration system: a language model trained to route...

Ctx
1000K
In
$5.00 / 1M
Out
$30.00 / 1M
Google: Nano Banana 2 (Gemini 3.1 Flash Image)
google/gemini-3.1-flash-image

Gemini 3.1 Flash Image, a.k.a. "Nano Banana 2," is Google’s latest state of the art image generation and editing model, delivering Pro-level visual quality at Flash speed. It combines advanced...

Ctx
131K
In
$0.500 / 1M
Out
$3.00 / 1M
Google: Nano Banana Pro (Gemini 3 Pro Image)
google/gemini-3-pro-image

Nano Banana Pro is Google’s most advanced image-generation and editing model, built on Gemini 3 Pro. It extends the original Nano Banana with significantly improved multimodal reasoning, real-world grounding, and...

Ctx
131K
In
$2.00 / 1M
Out
$12.00 / 1M
Cohere: North Mini Code (free)
cohere/north-mini-code:free

North Mini Code is Cohere's first agentic coding model and the debut of its North family. A sparse mixture-of-experts model with 30B total parameters and 3B active, it is optimized...

Ctx
256K
In
Free
Out
Free
Z.ai: GLM 5.2
z-ai/glm-5.2

GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent workflows, project-level software engineering,...

Ctx
1049K
In
$0.182 / 1M
Out
$0.572 / 1M
Z.ai: GLM 5.2 (batch)
z-ai/glm-5.2:batch

GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent workflows, project-level software engineering,...

Ctx
512K
In
$0.700 / 1M
Out
$2.20 / 1M
OpenRouter: Fusion
openrouter/fusion

Fusion turns your prompt into a small multi-model deliberation. A panel of expert models (see below) analyzes your prompt in parallel with web search and web fetch enabled, then a...

Ctx
1000K
In
$-1000000.000 / 1M
Out
$-1000000.000 / 1M
MoonshotAI: Kimi K2.7 Code
moonshotai/kimi-k2.7-code

MoonshotAI: Kimi K2.7 Code is a coding-focused model in Moonshot AI's Kimi K2 family, built to complete end-to-end programming tasks reliably over long contexts. It uses a native multimodal mixture-of-experts...

Ctx
262K
In
$0.700 / 1M
Out
$3.50 / 1M
MoonshotAI: Kimi K2.7 Code (batch)
moonshotai/kimi-k2.7-code:batch

MoonshotAI: Kimi K2.7 Code is a coding-focused model in Moonshot AI's Kimi K2 family, built to complete end-to-end programming tasks reliably over long contexts. It uses a native multimodal mixture-of-experts...

Ctx
262K
In
$0.475 / 1M
Out
$2.00 / 1M
Anthropic: Claude Fable Latest
~anthropic/claude-fable-latest

This model always redirects to the latest model in the Claude Fable family.

Ctx
1000K
In
$10.00 / 1M
Out
$50.00 / 1M
Anthropic: Claude Fable 5
anthropic/claude-fable-5

Claude Fable 5 is a Mythos-class model from Anthropic, built for autonomous knowledge work and coding. It supports text, image, and file inputs with text output, with reasoning support and...

Ctx
1000K
In
$10.00 / 1M
Out
$50.00 / 1M
Anthropic: Claude Fable 5 (batch)
anthropic/claude-fable-5:batch

Claude Fable 5 is a Mythos-class model from Anthropic, built for autonomous knowledge work and coding. It supports text, image, and file inputs with text output, with reasoning support and...

Ctx
1000K
In
$5.00 / 1M
Out
$25.00 / 1M
Nex AGI: Nex-N2-Pro
nex-agi/nex-n2-pro

Nex-N2-Pro is an agentic mixture-of-experts model from Nex AGI, with 17B active parameters out of 397B total. Built on the Qwen3.5 architecture, it accepts text and image input and produces...

Ctx
262K
In
$0.250 / 1M
Out
$1.00 / 1M
NVIDIA: Nemotron 3.5 Content Safety (free)
nvidia/nemotron-3.5-content-safety:free

NVIDIA Nemotron 3.5 Content Safety is a compact 4B-parameter multimodal guardrail model from NVIDIA, fine-tuned from Google Gemma-3-4B. It moderates both inputs to and responses from LLMs and VLMs, accepting...

Ctx
128K
In
Free
Out
Free
NVIDIA: Nemotron 3 Ultra
nvidia/nemotron-3-ultra-550b-a55b

NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...

Ctx
512K
In
$0.600 / 1M
Out
$3.60 / 1M
NVIDIA: Nemotron 3 Ultra (batch)
nvidia/nemotron-3-ultra-550b-a55b:batch

NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...

Ctx
512K
In
$0.300 / 1M
Out
$1.80 / 1M
NVIDIA: Nemotron 3 Ultra (free)
nvidia/nemotron-3-ultra-550b-a55b:free

NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...

Ctx
1000K
In
Free
Out
Free
Qwen: Qwen3.7 Plus
qwen/qwen3.7-plus

Qwen3.7-Plus is a cost-effective model in Alibaba's Qwen3.7 series. It supports text and image input with text output, building on the series' text capabilities with a comprehensive upgrade to its...

Ctx
1000K
In
$0.320 / 1M
Out
$1.28 / 1M
MiniMax: MiniMax M3
minimax/minimax-m3

MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,...

Ctx
1049K
In
$0.300 / 1M
Out
$1.20 / 1M
MiniMax: MiniMax M3 (batch)
minimax/minimax-m3:batch

MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,...

Ctx
524K
In
$0.150 / 1M
Out
$0.600 / 1M
StepFun: Step 3.7 Flash
stepfun/step-3.7-flash

Step 3.7 Flash is StepFun's latest high-efficiency multimodal Mixture-of-Experts model. It pairs a 196B-parameter language backbone with a vision encoder for native image and video understanding, activating roughly 11B parameters...

Ctx
262K
In
$0.200 / 1M
Out
$1.15 / 1M
Anthropic: Claude Opus 4.8 (Fast)
anthropic/claude-opus-4.8-fast

Fast-mode variant of [Opus 4.8](/anthropic/claude-opus-4.8) - identical capabilities with higher output speed at 2x pricing relative to regular Opus 4.8. Learn more in Anthropic's docs: https://platform.claude.com/docs/en/build-with-claude/fast-mode

Ctx
1000K
In
$10.00 / 1M
Out
$50.00 / 1M
Anthropic: Claude Opus 4.8
anthropic/claude-opus-4.8

Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family. It supports text, image, and file inputs with text output, with reasoning support and a 1M-token...

Ctx
1000K
In
$5.00 / 1M
Out
$25.00 / 1M
Anthropic: Claude Opus 4.8 (batch)
anthropic/claude-opus-4.8:batch

Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family. It supports text, image, and file inputs with text output, with reasoning support and a 1M-token...

Ctx
1000K
In
$2.50 / 1M
Out
$12.50 / 1M
Qwen: Qwen3.7 Max
qwen/qwen3.7-max

Qwen3.7-Max is the flagship model in Alibaba's Qwen3.7 series. It supports text input and output and is designed for agent-centric workloads, with particular strengths in coding, office and productivity tasks,...

Ctx
1000K
In
$1.48 / 1M
Out
$4.42 / 1M
SpaceXAI: Grok Build 0.1
x-ai/grok-build-0.1

Grok Build 0.1 is SpaceXAI’s fast coding model trained specifically for agentic software engineering workflows. It supports text and image inputs with text output, and is optimized for interactive coding...

Ctx
256K
In
$1.00 / 1M
Out
$2.00 / 1M
Google: Gemini 3.5 Flash
google/gemini-3.5-flash

Gemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed. It is highly optimized for coding proficiency and parallel agentic execution...

Ctx
1049K
In
$1.50 / 1M
Out
$9.00 / 1M
Google: Gemini 3.5 Flash (batch)
google/gemini-3.5-flash:batch

Gemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed. It is highly optimized for coding proficiency and parallel agentic execution...

Ctx
1049K
In
$0.750 / 1M
Out
$4.50 / 1M
Anthropic: Claude Opus 4.7 (Fast)
anthropic/claude-opus-4.7-fast

Fast-mode variant of [Opus 4.7](/anthropic/claude-opus-4.7) - identical capabilities with higher output speed at premium 6x pricing. Learn more in Anthropic's docs: https://platform.claude.com/docs/en/build-with-claude/fast-mode

Ctx
1000K
In
$30.00 / 1M
Out
$150.00 / 1M
Perceptron: Perceptron Mk1
perceptron/perceptron-mk1

Perceptron Mk1 (Mark One) is Perceptron's highest-quality vision-language model for video and embodied reasoning.** It accepts image and video inputs paired with natural language queries, and produces detailed visual understanding...

Ctx
33K
In
$0.150 / 1M
Out
$1.50 / 1M
inclusionAI: Ring-2.6-1T
inclusionai/ring-2.6-1t

Ring-2.6-1T is a 1T-parameter-scale thinking model with 63B active parameters, built for real-world agent workflows that require both strong capability and operational efficiency. It is optimized for coding agents, tool...

Ctx
262K
In
$0.075 / 1M
Out
$0.625 / 1M
Google: Gemini 3.1 Flash Lite
google/gemini-3.1-flash-lite

Gemini 3.1 Flash Lite is Google’s GA high-efficiency multimodal model optimized for low-latency, high-volume workloads. It supports text, image, video, audio, and PDF inputs, and is designed for lightweight agentic...

Ctx
1049K
In
$0.250 / 1M
Out
$1.50 / 1M
Google: Gemini 3.1 Flash Lite (batch)
google/gemini-3.1-flash-lite:batch

Gemini 3.1 Flash Lite is Google’s GA high-efficiency multimodal model optimized for low-latency, high-volume workloads. It supports text, image, video, audio, and PDF inputs, and is designed for lightweight agentic...

Ctx
1049K
In
$0.125 / 1M
Out
$0.750 / 1M
OpenAI: GPT Chat Latest
openai/gpt-chat-latest

GPT Chat Latest points to OpenAI's stable API alias `chat-latest` that always resolves to the latest Instant chat model used in ChatGPT. As OpenAI rolls out new Instant model updates...

Ctx
400K
In
$5.00 / 1M
Out
$30.00 / 1M
SpaceXAI: Grok 4.3
x-ai/grok-4.3

Grok 4.3 is a reasoning model from SpaceXAI. It accepts text and image inputs with text output, and is suited for agentic workflows, instruction-following tasks, and applications requiring high factual...

Ctx
1000K
In
$1.25 / 1M
Out
$2.50 / 1M
IBM: Granite 4.1 8B
ibm-granite/granite-4.1-8b

Granite 4.1 8B is a dense, decoder-only 8-billion-parameter language model from IBM, part of the Granite 4.1 family. It supports a 131K-token context window and is designed for enterprise tasks...

Ctx
131K
In
$0.050 / 1M
Out
$0.100 / 1M
Mistral: Mistral Medium 3.5
mistralai/mistral-medium-3-5

Mistral Medium 3.5 is a dense 128B instruction-following model from Mistral AI. It supports text and image inputs with text output, and is designed for agentic workflows, coding, and complex...

Ctx
262K
In
$1.50 / 1M
Out
$7.50 / 1M
NVIDIA: Nemotron 3 Nano Omni (free)
nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free

NVIDIA Nemotron™ 3 Nano Omni is a 30B-A3B open multimodal model designed to function as a perception and context sub-agent in enterprise agent systems. It accepts text, image, video, and...

Ctx
256K
In
Free
Out
Free
Anthropic Claude Haiku Latest
~anthropic/claude-haiku-latest

This model always redirects to the latest model in the Anthropic Claude Haiku family.

Ctx
200K
In
$1.00 / 1M
Out
$5.00 / 1M
OpenAI GPT Mini Latest
~openai/gpt-mini-latest

This model always redirects to the latest model in the OpenAI GPT Mini family.

Ctx
400K
In
$0.750 / 1M
Out
$4.50 / 1M
Google Gemini Pro Latest
~google/gemini-pro-latest

This model always redirects to the latest model in the Google Gemini Pro family.

Ctx
1049K
In
$2.00 / 1M
Out
$12.00 / 1M
MoonshotAI Kimi Latest
~moonshotai/kimi-latest

This model always redirects to the latest model in the MoonshotAI Kimi family.

Ctx
1049K
In
$2.80 / 1M
Out
$14.00 / 1M
Google Gemini Flash Latest
~google/gemini-flash-latest

This model always redirects to the latest model in the Google Gemini Flash family.

Ctx
1049K
In
$1.50 / 1M
Out
$7.50 / 1M
Anthropic Claude Sonnet Latest
~anthropic/claude-sonnet-latest

This model always redirects to the latest model in the Anthropic Claude Sonnet family.

Ctx
1000K
In
$2.00 / 1M
Out
$10.00 / 1M
OpenAI GPT Latest
~openai/gpt-latest

This model always redirects to the latest model in the OpenAI GPT family.

Ctx
1050K
In
$5.00 / 1M
Out
$30.00 / 1M
Qwen: Qwen3.5 Plus 2026-04-20
qwen/qwen3.5-plus-20260420

Qwen3.5 Plus (April 2026) is a large-scale multimodal language model from Alibaba. It accepts text, image, and video input and produces text output, with a 1M token context window. This...

Ctx
1000K
In
$0.300 / 1M
Out
$1.80 / 1M
Qwen: Qwen3.6 Flash
qwen/qwen3.6-flash

Qwen3.6 Flash is a fast, efficient language model from Alibaba's Qwen 3.6 series. It supports text, image, and video input with a 1M token context window. Tiered pricing kicks in...

Ctx
1000K
In
$0.188 / 1M
Out
$1.13 / 1M
Qwen: Qwen3.6 35B A3B
qwen/qwen3.6-35b-a3b

Qwen3.6-35B-A3B is an open-weight multimodal model from Alibaba Cloud with 35 billion total parameters and 3 billion active parameters per token. It uses a hybrid sparse mixture-of-experts architecture combining Gated...

Ctx
262K
In
$0.150 / 1M
Out
$1.00 / 1M
Qwen: Qwen3.6 Max Preview
qwen/qwen3.6-max-preview

Qwen3.6-Max-Preview is a proprietary frontier model from Alibaba Cloud built on a sparse mixture-of-experts architecture with approximately 1 trillion total parameters. It is optimized for agentic coding, tool use, and...

Ctx
262K
In
$1.03 / 1M
Out
$6.16 / 1M
Qwen: Qwen3.6 27B
qwen/qwen3.6-27b

Qwen3.6 27B is a dense 27-billion-parameter language model from the Qwen Team at Alibaba, released in April 2026. It features hybrid multimodal capabilities — accepting text, image, and video inputs...

Ctx
262K
In
$0.600 / 1M
Out
$3.60 / 1M
OpenAI: GPT-5.5 Pro
openai/gpt-5.5-pro

GPT-5.5 Pro is OpenAI’s high-capability model optimized for deep reasoning and accuracy on complex, high-stakes workloads. It features a 1M+ token context window (922K input, 128K output) with support for...

Ctx
1050K
In
$30.00 / 1M
Out
$180.00 / 1M
OpenAI: GPT-5.5 Pro (batch)
openai/gpt-5.5-pro:batch

GPT-5.5 Pro is OpenAI’s high-capability model optimized for deep reasoning and accuracy on complex, high-stakes workloads. It features a 1M+ token context window (922K input, 128K output) with support for...

Ctx
1050K
In
$15.00 / 1M
Out
$90.00 / 1M
OpenAI: GPT-5.5
openai/gpt-5.5

GPT-5.5 is OpenAI’s frontier model designed for complex professional workloads, building on GPT-5.4 with stronger reasoning, higher reliability, and improved token efficiency on hard tasks. It features a 1M+ token...

Ctx
1050K
In
$5.00 / 1M
Out
$30.00 / 1M
OpenAI: GPT-5.5 (batch)
openai/gpt-5.5:batch

GPT-5.5 is OpenAI’s frontier model designed for complex professional workloads, building on GPT-5.4 with stronger reasoning, higher reliability, and improved token efficiency on hard tasks. It features a 1M+ token...

Ctx
1050K
In
$2.50 / 1M
Out
$15.00 / 1M
DeepSeek: DeepSeek V4 Pro
deepseek/deepseek-v4-pro

DeepSeek V4 Pro is a large-scale Mixture-of-Experts model from DeepSeek with 1.6T total parameters and 49B activated parameters, supporting a 1M-token context window. It is designed for advanced reasoning, coding,...

Ctx
1049K
In
$0.435 / 1M
Out
$0.870 / 1M
DeepSeek: DeepSeek V4 Flash 0423
deepseek/deepseek-v4-flash

DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters, supporting a 1M-token context window. It is designed for fast inference and...

Ctx
1049K
In
$0.140 / 1M
Out
$0.280 / 1M
inclusionAI: Ling-2.6-1T
inclusionai/ling-2.6-1t

Ling-2.6-1T is an instant (instruct) model from inclusionAI and the company’s trillion-parameter flagship, designed for real-world agents that require fast execution and high efficiency at scale. It uses a “fast...

Ctx
262K
In
$0.075 / 1M
Out
$0.625 / 1M
Tencent: Hy3 preview
tencent/hy3-preview

Hy3 preview is a high-efficiency Mixture-of-Experts model from Tencent designed for agentic workflows and production use. It supports configurable reasoning levels across disabled, low, and high modes, allowing it to...

Ctx
262K
In
$0.063 / 1M
Out
$0.210 / 1M
Xiaomi: MiMo-V2.5-Pro
xiaomi/mimo-v2.5-pro

MiMo-V2.5-Pro is Xiaomi’s flagship model, delivering strong performance in general agentic capabilities, complex software engineering, and long-horizon tasks, with top rankings on benchmarks such as ClawEval, GDPVal, and SWE-bench Pro....

Ctx
1050K
In
$0.435 / 1M
Out
$0.870 / 1M
Xiaomi: MiMo-V2.5
xiaomi/mimo-v2.5

MiMo-V2.5 is a native omnimodal model by Xiaomi. It delivers Pro-level agentic performance at roughly half the inference cost, while surpassing MiMo-V2-Omni in multimodal perception across image and video understanding...

Ctx
1050K
In
$0.140 / 1M
Out
$0.280 / 1M
OpenAI: GPT-5.4 Image 2
openai/gpt-5.4-image-2

[GPT-5.4](https://openrouter.ai/openai/gpt-5.4) Image 2 combines OpenAI's GPT-5.4 model with state-of-the-art image generation capabilities from GPT Image 2. It enables rich multimodal workflows, allowing users to seamlessly move between reasoning, coding, and...

Ctx
272K
In
$8.00 / 1M
Out
$15.00 / 1M
inclusionAI: Ling-2.6-flash
inclusionai/ling-2.6-flash

Ling-2.6-flash is an instant (instruct) model from inclusionAI with 104B total parameters and 7.4B active parameters, designed for real-world agents that require fast responses, strong execution, and high token efficiency....

Ctx
262K
In
$0.010 / 1M
Out
$0.030 / 1M
Anthropic: Claude Opus Latest
~anthropic/claude-opus-latest

This model always redirects to the latest model in the Claude Opus family.

Ctx
1000K
In
$5.00 / 1M
Out
$25.00 / 1M
Pareto Code Router
openrouter/pareto-code

The Pareto Router maintains a tiered shortlist of strong coding models, ranked by [Artificial Analysis](https://artificialanalysis.ai/) coding percentiles. Set min_coding_score between 0 and 1 on the [pareto-router plugin](https://openrouter.ai/docs/guides/routing/routers/pareto-router#the-min_coding_score-parameter) to control how...

Ctx
2000K
In
$-1000000.000 / 1M
Out
$-1000000.000 / 1M
MoonshotAI: Kimi K2.6
moonshotai/kimi-k2.6

Kimi K2.6 is Moonshot AI's next-generation multimodal model, designed for long-horizon coding, coding-driven UI/UX generation, and multi-agent orchestration. It handles complex end-to-end coding tasks across Python, Rust, and Go, and...

Ctx
262K
In
$0.580 / 1M
Out
$2.44 / 1M
Anthropic: Claude Opus 4.7
anthropic/claude-opus-4.7

Opus 4.7 is the next generation of Anthropic's Opus family, built for long-running, asynchronous agents. Building on the coding and agentic strengths of Opus 4.6, it delivers stronger performance on...

Ctx
1000K
In
$5.00 / 1M
Out
$25.00 / 1M
Anthropic: Claude Opus 4.7 (batch)
anthropic/claude-opus-4.7:batch

Opus 4.7 is the next generation of Anthropic's Opus family, built for long-running, asynchronous agents. Building on the coding and agentic strengths of Opus 4.6, it delivers stronger performance on...

Ctx
1000K
In
$2.50 / 1M
Out
$12.50 / 1M
Z.ai: GLM 5.1
z-ai/glm-5.1

GLM-5.1 delivers a major leap in coding capability, with particularly significant gains in handling long-horizon tasks. Unlike previous models built around minute-level interactions, GLM-5.1 can work independently and continuously on...

Ctx
205K
In
$0.952 / 1M
Out
$2.99 / 1M
Google: Gemma 4 26B A4B
google/gemma-4-26b-a4b-it

Gemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind. Despite 25.2B total parameters, only 3.8B activate per token during inference — delivering near-31B quality at...

Ctx
262K
In
$0.070 / 1M
Out
$0.340 / 1M
Google: Gemma 4 26B A4B (free)
google/gemma-4-26b-a4b-it:free

Gemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind. Despite 25.2B total parameters, only 3.8B activate per token during inference — delivering near-31B quality at...

Ctx
262K
In
Free
Out
Free
Google: Gemma 4 31B
google/gemma-4-31b-it

Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function...

Ctx
262K
In
$0.100 / 1M
Out
$0.340 / 1M
Google: Gemma 4 31B (free)
google/gemma-4-31b-it:free

Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function...

Ctx
262K
In
Free
Out
Free
Showing first 120 — refine search to narrow results.

Unified access layer

One key. Every model on the network.

Stop juggling per-vendor SDKs and quotas — TokensChain offers an OpenAI-compatible front door with smart routing baked in.
tokenschain · preview