Global model marketplace

Every TokensChain model,
one searchable view.

Live metadata for 300+ TokensChain models — price, context, modality, all visible.

459 models
Z.ai: GLM 5.3 Prime
z-ai/glm-5.3-prime

GLM-5.3-Prime is the high-speed variant of Z.ai's GLM-5.3, inheriting its full capabilities while delivering 1.5–2× the output throughput through inference acceleration. It supports text input and output with a 1M-token...

Ctx
1000K
In
$2.80 / 1M
Out
$8.80 / 1M
Qwen: Qwen3.8 Max Prime
qwen/qwen3.8-max-prime

Qwen3.8 Max Prime is a higher-throughput variant of Qwen3.8 Max from Alibaba's Qwen team, served as a separate SKU at a higher price point. It accepts text, image, and video...

Ctx
1000K
In
$4.00 / 1M
Out
$12.00 / 1M
Space Bunny Alpha
stealth/space-bunny-alpha

Space Bunny Alpha is an anonymous large model with blazing-fast inference, strong coding capabilities and native multimodal input support. It delivers adjustable reasoning effort, and a 1M-token context window. Space...

Ctx
1000K
In
Free
Out
Free
AionLabs: Aion 3.5 Mini
aion-labs/aion-3.5-mini

Aion 3.5 Mini is a multi-model roleplaying and storytelling system from AionLabs, built on the GLM family of models. It is the smaller, lower-cost sibling of Aion 3.5 and uses...

Ctx
262K
In
$0.700 / 1M
Out
$1.40 / 1M
AionLabs: Aion 3.5
aion-labs/aion-3.5

Aion 3.5 is a multi-model roleplaying and storytelling system from AionLabs, built on the GLM family of models. It uses a collaborative generation process in which multiple specialized models each...

Ctx
262K
In
$3.00 / 1M
Out
$6.00 / 1M
Upstage: Solar Mini 4
upstage/solar-mini4

Solar Mini 4 is Upstage's compact, cost-efficient language model, a 35B-parameter mixture-of-experts with 3B active parameters and a 524K context window. It is built for agentic use cases where response...

Ctx
524K
In
$0.050 / 1M
Out
$0.200 / 1M
Cohere: Command A+
cohere/command-a-plus

Command A+ is Cohere's flagship model for enterprise agentic workflows. It accepts text and image inputs with a 192K context window, supports native tool calling with strict tool schemas, structured...

Ctx
192K
In
$0.300 / 1M
Out
$1.50 / 1M
OpenAI: GPT-6 Luna Pro
openai/gpt-6-luna-pro

GPT-6 Luna Pro is the same underlying model as [GPT-6 Luna](https://openrouter.ai/openai/gpt-6-luna), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode

Ctx
1050K
In
$0.100 / 1M
Out
$0.500 / 1M
OpenAI: GPT-6 Luna Pro (batch)
openai/gpt-6-luna-pro:batch

GPT-6 Luna Pro is the same underlying model as [GPT-6 Luna](https://openrouter.ai/openai/gpt-6-luna), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode

Ctx
1050K
In
$0.050 / 1M
Out
$0.250 / 1M
OpenAI: GPT-6 Luna
openai/gpt-6-luna

GPT-6 Luna is the fast, cost-efficient model in OpenAI's GPT-6 series, positioned below GPT-6 Sol. It is suited for high-volume and latency-sensitive workloads such as chat, classification, and lightweight agentic...

Ctx
1050K
In
$0.100 / 1M
Out
$0.500 / 1M
OpenAI: GPT-6 Luna (batch)
openai/gpt-6-luna:batch

GPT-6 Luna is the fast, cost-efficient model in OpenAI's GPT-6 series, positioned below GPT-6 Sol. It is suited for high-volume and latency-sensitive workloads such as chat, classification, and lightweight agentic...

Ctx
1050K
In
$0.050 / 1M
Out
$0.250 / 1M
OpenAI: GPT-6 Sol Pro
openai/gpt-6-sol-pro

GPT-6 Sol Pro is the same underlying model as [GPT-6 Sol](https://openrouter.ai/openai/gpt-6-sol), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode

Ctx
1050K
In
$2.00 / 1M
Out
$10.00 / 1M
OpenAI: GPT-6 Sol Pro (batch)
openai/gpt-6-sol-pro:batch

GPT-6 Sol Pro is the same underlying model as [GPT-6 Sol](https://openrouter.ai/openai/gpt-6-sol), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode

Ctx
1050K
In
$1.00 / 1M
Out
$5.00 / 1M
OpenAI: GPT-6 Sol
openai/gpt-6-sol

GPT-6 Sol is the cost-efficient high-end model in OpenAI's GPT-6 series, positioned below the flagship GPT-6 Astra and above the fast GPT-6 Luna tier. It is suited for demanding professional...

Ctx
1050K
In
$2.00 / 1M
Out
$10.00 / 1M
OpenAI: GPT-6 Sol (batch)
openai/gpt-6-sol:batch

GPT-6 Sol is the cost-efficient high-end model in OpenAI's GPT-6 series, positioned below the flagship GPT-6 Astra and above the fast GPT-6 Luna tier. It is suited for demanding professional...

Ctx
1050K
In
$1.00 / 1M
Out
$5.00 / 1M
Anthropic: Claude Opus 5.5
anthropic/claude-opus-5.5

Claude Opus 5.5 is Anthropic's flagship model for demanding reasoning, coding, and long-horizon agentic work, succeeding Claude Opus 5. It is particularly strong at multi-step changes in large codebases, code...

Ctx
1000K
In
$4.00 / 1M
Out
$20.00 / 1M
Anthropic: Claude Opus 5.5 (batch)
anthropic/claude-opus-5.5:batch

Claude Opus 5.5 is Anthropic's flagship model for demanding reasoning, coding, and long-horizon agentic work, succeeding Claude Opus 5. It is particularly strong at multi-step changes in large codebases, code...

Ctx
1000K
In
$2.00 / 1M
Out
$10.00 / 1M
Xiaomi: MiMo-V2.6-Pro-UltraSpeed
xiaomi/mimo-v2.6-pro-ultraspeed

MiMo-V2.6-Pro-UltraSpeed is the fast speed edition of Xiaomi's flagship foundation model, MiMo-V2.6-Pro. Built from the same 1T MiMo-V2.6-Pro checkpoint, it matches the original model in quality while delivering roughly 10x...

Ctx
1049K
In
$4.35 / 1M
Out
$8.70 / 1M
Xiaomi: MiMo-V2.6-Flash
xiaomi/mimo-v2.6-flash

MiMo-V2.6-Flash is an open-source foundation model developed by Xiaomi. Built on a Mixture-of-Experts architecture with 309B total parameters and 15B activated per token, it employs a hybrid attention mechanism for...

Ctx
1049K
In
$0.140 / 1M
Out
$0.280 / 1M
Xiaomi: MiMo-V2.6-Pro
xiaomi/mimo-v2.6-pro

MiMo-V2.6-Pro is the flagship foundation model developed by Xiaomi. Built at a scale of over 1T parameters, it is designed to push the ceiling of capability for the most demanding...

Ctx
1049K
In
$0.435 / 1M
Out
$0.870 / 1M
SpaceXAI: Grok 4.7
x-ai/grok-4.7

Grok 4.7 is SpaceXAI's flagship model for coding, agentic tasks, and knowledge work, succeeding Grok 4.6. It is particularly strong at long-running software engineering tasks, verifying its own work, and...

Ctx
500K
In
$1.60 / 1M
Out
$4.80 / 1M
Qwen: Qwen3.8 Omni Flash
qwen/qwen3.8-omni-flash

Qwen3.8 Omni Flash is an omni-modal reasoning model from Alibaba, the first Qwen model built around agentic capabilities with native audio-video understanding. It is suited for audio-video analysis and summarization,...

Ctx
1000K
In
$0.150 / 1M
Out
$0.470 / 1M
PrismML: Ternary Bonsai 2 27B
prism-ml/ternary-bonsai-2-27b

Bonsai 2 27B is a 27B-parameter reasoning model from PrismML derived from Qwen3.8-27B. It supports coding, mathematics, tool calling, and image understanding with a 262K-token context window. Ternary compression shrinks...

Ctx
262K
In
$0.075 / 1M
Out
$0.500 / 1M
Z.ai: GLM 5.3 FlashX
z-ai/glm-5.3-flashx

GLM-5.3-FlashX is the high-speed variant of Z.ai's GLM-5.3-Flash, a native multimodal model delivering inference speeds of up to 200 tokens/s. Built on the same hybrid sparse and linear attention architecture...

Ctx
1049K
In
$0.370 / 1M
Out
$1.25 / 1M
Pareto
unbiased/pareto

Pareto is a multimodal composite model built for research, coding, and agentic workflows, while delivering frontier-level performance across a broad range of general-purpose tasks.

Ctx
262K
In
$2.50 / 1M
Out
$7.50 / 1M
DeepSeek: DeepSeek Pro Latest
~deepseek/deepseek-pro-latest

This model always redirects to the latest model in the DeepSeek Pro family.

Ctx
1049K
In
$0.387 / 1M
Out
$2.90 / 1M
DeepSeek: DeepSeek Flash Latest
~deepseek/deepseek-flash-latest

This model always redirects to the latest model in the DeepSeek Flash family.

Ctx
1049K
In
$0.099 / 1M
Out
$0.600 / 1M
Inference.net: Schematron V2 Turbo
inference-net/schematron-v2-turbo

Schematron V2 Turbo is a 3B-parameter HTML-to-JSON extraction model from Inference.net. It prioritizes throughput for high-volume extraction workloads. Extraction instructions must be supplied through a JSON schema in response_format rather...

Ctx
128K
In
$0.030 / 1M
Out
$0.150 / 1M
Inference.net: Schematron V2 Small
inference-net/schematron-v2-small

Schematron V2 Small is a 3B-parameter HTML-to-JSON extraction model from Inference.net. It prioritizes extraction quality for complex schemas and long pages. Extraction instructions must be supplied through a JSON schema...

Ctx
128K
In
$0.050 / 1M
Out
$0.230 / 1M
OpenAI: GPT Astra Latest
~openai/gpt-astra-latest

This model always redirects to the latest model in the GPT Astra family.

Ctx
1050K
In
$10.00 / 1M
Out
$50.00 / 1M
OpenAI: GPT Sol Latest
~openai/gpt-sol-latest

This model always redirects to the latest model in the GPT Sol family.

Ctx
1050K
In
$2.00 / 1M
Out
$10.00 / 1M
OpenAI: GPT Terra Latest
~openai/gpt-terra-latest

This model always redirects to the latest model in the GPT Terra family.

Ctx
1050K
In
$2.00 / 1M
Out
$12.00 / 1M
OpenAI: GPT Luna Latest
~openai/gpt-luna-latest

This model always redirects to the latest model in the GPT Luna family.

Ctx
1050K
In
$0.100 / 1M
Out
$0.500 / 1M
Sakana: Fugu Ultra v2
sakana/fugu-ultra-v2

Fugu Ultra v2 is the higher-performance model in Sakana AI's Fugu family. Rather than a single monolithic model, Fugu is a learned multi-agent orchestration system: a language model trained to...

Ctx
1000K
In
$5.00 / 1M
Out
$30.00 / 1M
Sakana: Fugu Max
sakana/fugu-max

Fugu Max is the cost-performance model in Sakana AI's Fugu family. Rather than a single monolithic model, Fugu is a learned multi-agent orchestration system: a language model trained to route...

Ctx
1000K
In
$2.00 / 1M
Out
$6.00 / 1M
inclusionAI: Ling 3.0 Flash VL
inclusionai/ling-3.0-flash-vl

Ling 3.0 Flash VL builds on Ling 3.0 Flash (124B total / 5.5B active MoE from InclusionAI), further strengthening its language capabilities while adding native visual perception and advanced visual...

Ctx
262K
In
$0.060 / 1M
Out
$0.180 / 1M
DeepSeek: DeepSeek V4.1 Flash
deepseek/deepseek-v4.1-flash

DeepSeek V4.1 Flash is a sparse mixture-of-experts model from DeepSeek, and the first built on the company's Causal Encoder-Decoder (CED) architecture. It activates 8B parameters on input and 16B on...

Ctx
1049K
In
$0.140 / 1M
Out
$0.420 / 1M
DeepSeek: DeepSeek V4.1 Flash (batch)
deepseek/deepseek-v4.1-flash:batch

DeepSeek V4.1 Flash is a sparse mixture-of-experts model from DeepSeek, and the first built on the company's Causal Encoder-Decoder (CED) architecture. It activates 8B parameters on input and 16B on...

Ctx
1049K
In
$0.112 / 1M
Out
$0.336 / 1M
Inception: Mercury 2.5
inception/mercury-2.5

Mercury 2.5 is the fastest reasoning LLM, and the latest diffusion LLM (dLLM) from Inception. Instead of generating tokens sequentially, Mercury 2.5 produces and refines multiple tokens in parallel, achieving...

Ctx
260K
In
$0.040 / 1M
Out
$0.150 / 1M
Nex AGI: Nex-N2.5-Mini
nex-agi/nex-n2.5-mini

Nex-N2.5 is an agentic model built to turn goals into working, verified outcomes. Its core strength is agentic coding within a visual feedback loop: it can explore codebases, implement multi-file...

Ctx
262K
In
$0.025 / 1M
Out
$0.100 / 1M
Nex AGI: Nex-N2.5-Mini (free)
nex-agi/nex-n2.5-mini:free

Nex-N2.5 is an agentic model built to turn goals into working, verified outcomes. Its core strength is agentic coding within a visual feedback loop: it can explore codebases, implement multi-file...

Ctx
262K
In
Free
Out
Free
Nex AGI: Nex-N2.5-Pro
nex-agi/nex-n2.5-pro

Nex-N2.5 is an agentic model built to turn goals into working, verified outcomes. Its core strength is agentic coding within a visual feedback loop: it can explore codebases, implement multi-file...

Ctx
262K
In
$0.075 / 1M
Out
$0.250 / 1M
Nex AGI: Nex-N2.5-Pro (free)
nex-agi/nex-n2.5-pro:free

Nex-N2.5 is an agentic model built to turn goals into working, verified outcomes. Its core strength is agentic coding within a visual feedback loop: it can explore codebases, implement multi-file...

Ctx
262K
In
Free
Out
Free
OpenAI: GPT-6 Astra
openai/gpt-6-astra

GPT-6 Astra is OpenAI's flagship model for demanding end-to-end work. It is suited for advanced analysis, software engineering, deep research, scientific work, and document creation, with particular strengths in long-horizon...

Ctx
1050K
In
$10.00 / 1M
Out
$50.00 / 1M
OpenAI: GPT-6 Astra (batch)
openai/gpt-6-astra:batch

GPT-6 Astra is OpenAI's flagship model for demanding end-to-end work. It is suited for advanced analysis, software engineering, deep research, scientific work, and document creation, with particular strengths in long-horizon...

Ctx
1050K
In
$5.00 / 1M
Out
$25.00 / 1M
OpenAI: GPT-6 Astra Pro
openai/gpt-6-astra-pro

GPT-6 Astra Pro is the same underlying model as [GPT-6 Astra](https://openrouter.ai/openai/gpt-6-astra), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode

Ctx
1050K
In
$10.00 / 1M
Out
$50.00 / 1M
OpenAI: GPT-6 Astra Pro (batch)
openai/gpt-6-astra-pro:batch

GPT-6 Astra Pro is the same underlying model as [GPT-6 Astra](https://openrouter.ai/openai/gpt-6-astra), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode

Ctx
1050K
In
$5.00 / 1M
Out
$25.00 / 1M
inclusionAI: Ling 3.0 Flash Sante (free)
inclusionai/ling-3.0-flash-sante:free

Ling 3.0 Flash Sante is a health and medicine-focused mixture-of-experts model from InclusionAI, built on Ling 3.0 Flash with 5.1B active parameters out of 124B total. It is designed for...

Ctx
262K
In
Free
Out
Free
Qwen: Qwen3.8 Max (0902)
qwen/qwen3.8-max-0902

Qwen3.8 Max 0902 is an updated snapshot of Qwen3.8 Max from Alibaba's Qwen team. It is a 2.4-trillion-parameter mixture-of-experts model that accepts text, image, and video input and returns text,...

Ctx
1000K
In
$2.00 / 1M
Out
$6.00 / 1M
Meta: Muse Spark 1.3 Contributor
meta/muse-spark-1.3-contributor

Muse Spark 1.3 Contributor is the cost-efficient contributor tier of Meta’s multimodal reasoning model for experimentation, learning, and early-stage agentic, multi-agent, and coding workflows. It is designed to track information...

Ctx
1049K
In
$0.100 / 1M
Out
$0.200 / 1M
Meta: Muse Spark 1.3
meta/muse-spark-1.3

Muse Spark 1.3 is a multimodal reasoning model from Meta for long-running agentic, multi-agent, and coding workflows. It is designed to keep track of information across extended tasks, work through...

Ctx
1049K
In
$1.25 / 1M
Out
$4.25 / 1M
Google: Gemini 3.8 Flash
google/gemini-3.8-flash

Gemini 3.8 Flash is Google's most intelligent Flash model with significant gains from 3.7 Flash across software engineering, agentic tasks, and multi-step reasoning.

Ctx
1049K
In
$0.750 / 1M
Out
$3.75 / 1M
Google: Gemini 3.8 Flash (batch)
google/gemini-3.8-flash:batch

Gemini 3.8 Flash is Google's most intelligent Flash model with significant gains from 3.7 Flash across software engineering, agentic tasks, and multi-step reasoning.

Ctx
1049K
In
$0.375 / 1M
Out
$1.88 / 1M
Anthropic: Claude Fable 5.1
anthropic/claude-fable-5.1

Claude Fable 5.1 improves on Claude Fable 5 across the board, with the biggest gains in agentic coding, long-running agentic workflows, and knowledge work: long code refactors, front-end and visual...

Ctx
1000K
In
$10.00 / 1M
Out
$50.00 / 1M
Anthropic: Claude Fable 5.1 (batch)
anthropic/claude-fable-5.1:batch

Claude Fable 5.1 improves on Claude Fable 5 across the board, with the biggest gains in agentic coding, long-running agentic workflows, and knowledge work: long code refactors, front-end and visual...

Ctx
1000K
In
$5.00 / 1M
Out
$25.00 / 1M
IBM: Granite 4.2 8B
ibm-granite/granite-4.2-8b

Granite 4.2 8B is a dense reasoning model from IBM. It is suited for mathematics, code generation, multilingual dialogue, and agentic workflows that need multi-step reasoning. It supports full, low-effort,...

Ctx
131K
In
$0.060 / 1M
Out
$0.250 / 1M
Tencent: Hy4 preview
tencent/hy4-preview

Tencent: Hy4 preview is a mixture-of-experts model from Tencent, with 49B active parameters out of 770B total. It is designed for coding agents, complex tool-use workflows, and productivity tasks that...

Ctx
1049K
In
$0.834 / 1M
Out
$2.50 / 1M
inclusionAI: Ling 3.0 Flash Fin
inclusionai/ling-3.0-flash-fin

Ling 3.0 Flash Fin is a finance-focused mixture-of-experts model from InclusionAI, built on Ling 3.0 Flash with 5.1B active parameters out of 124B total. It is designed for real-world investment...

Ctx
262K
In
$0.060 / 1M
Out
$0.180 / 1M
inclusionAI: Ling 3.0 Flash Fin (free)
inclusionai/ling-3.0-flash-fin:free

Ling 3.0 Flash Fin is a finance-focused mixture-of-experts model from InclusionAI, built on Ling 3.0 Flash with 5.1B active parameters out of 124B total. It is designed for real-world investment...

Ctx
262K
In
Free
Out
Free
Z.ai: GLM Flash Latest
~z-ai/glm-flash-latest

This model always redirects to the latest model in the GLM Flash family.

Ctx
1311K
In
$0.075 / 1M
Out
$0.250 / 1M
Qwen: Qwen3.8 Flash
qwen/qwen3.8-flash

Qwen3.8 Flash is a multimodal reasoning model from Alibaba. It is suited for coding assistance, agentic workflows, visual understanding, document and codebase analysis, desktop interaction, chart analysis, and long-video analysis.

Ctx
1000K
In
$0.150 / 1M
Out
$0.470 / 1M
Z.ai: GLM 5.3 Flash
z-ai/glm-5.3-flash

GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while...

Ctx
1311K
In
$0.150 / 1M
Out
$0.500 / 1M
Z.ai: GLM 5.3 Flash (batch)
z-ai/glm-5.3-flash:batch

GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while...

Ctx
1049K
In
$0.060 / 1M
Out
$0.200 / 1M
Meta: Muse Spark 1.2 Contributor
meta/muse-spark-1.2-contributor

Muse Spark 1.2 contributor tier is a reasoning model from Meta designed for developers who want to start building at an even lower cost. It’s meaningfully cheaper than Muse Spark...

Ctx
1049K
In
$0.100 / 1M
Out
$0.200 / 1M
DeepSeek: DeepSeek V4 Flash Vision Exp
deepseek/deepseek-v4-flash-vision-exp

DeepSeek V4 Flash Vision Exp is an experimental vision-enabled version of [DeepSeek V4 Flash 0731](https://openrouter.ai/deepseek/deepseek-v4-flash-0731) from DeepSeek, adding image understanding while matching the base model on text capabilities including agents,...

Ctx
1049K
In
$0.220 / 1M
Out
$0.660 / 1M
Tencent: Hy-MT2-1.8B
tencent/hy-mt2-1.8b

Hy-MT2-1.8B is a compact 1.8B-parameter translation model from Tencent. It supports 33 language pairs and five Chinese dialect and minority-language pairs, with workflows for structured, delimiter-based, contextual, glossary-based, and style-guided...

Ctx
8K
In
$0.044 / 1M
Out
$0.177 / 1M
Tencent: Hy-MT2-30B-A3B
tencent/hy-mt2-30b-a3b

Hy-MT2-30B-A3B is Tencent's flagship translation model in the Hy-MT2 family. It supports 33 language pairs and five Chinese dialect and minority-language pairs, with workflows for structured, delimiter-based, contextual, glossary-based, and...

Ctx
8K
In
$0.074 / 1M
Out
$0.295 / 1M
Z.ai: GLM Latest
~z-ai/glm-latest

This model always redirects to the latest GLM model from Z.ai.

Ctx
1311K
In
$0.563 / 1M
Out
$2.50 / 1M
Tencent: Hy-MT2-7B
tencent/hy-mt2-7b

Hy-MT2-7B is a 7B-parameter translation model from Tencent. It supports 33 language pairs and five Chinese dialect and minority-language pairs, with workflows for structured, delimiter-based, contextual, glossary-based, and style-guided translation.

Ctx
8K
In
$0.074 / 1M
Out
$0.295 / 1M
Z.ai: GLM 5.3
z-ai/glm-5.3

GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks. It supports text input and output with a 1M-token context window, and improves...

Ctx
1311K
In
$0.840 / 1M
Out
$2.64 / 1M
Z.ai: GLM 5.3 (batch)
z-ai/glm-5.3:batch

GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks. It supports text input and output with a 1M-token context window, and improves...

Ctx
1049K
In
$0.450 / 1M
Out
$2.00 / 1M
Qwen: Qwen3.8 27B
qwen/qwen3.8-27b

Qwen3.8 27B is an open-weight dense vision-language model from Qwen. It is suited for coding, professional workflows, research, multimodal interaction, and long-running agent tasks, with flexible thinking that can be...

Ctx
1000K
In
$0.420 / 1M
Out
$3.00 / 1M
Qwen: Qwen3.8 27B (free)
qwen/qwen3.8-27b:free

Qwen3.8 27B is an open-weight dense vision-language model from Qwen. It is suited for coding, professional workflows, research, multimodal interaction, and long-running agent tasks, with flexible thinking that can be...

Ctx
262K
In
Free
Out
Free
Dots Studio: Dots3-Note Preview (free)
dots-studio/dots-3-note-preview:free

Dots3-Note Preview is an open-weight mixture-of-experts model from Dots Studio, with 16B active parameters out of 280B total. It is the lightest model in the Dots 3 family and is...

Ctx
512K
In
Free
Out
Free
Google: Gemini 3.7 Flash
google/gemini-3.7-flash

Gemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed for tasks that require responsive performance and reliable multi-step...

Ctx
1049K
In
$0.750 / 1M
Out
$3.75 / 1M
Google: Gemini 3.7 Flash (batch)
google/gemini-3.7-flash:batch

Gemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed for tasks that require responsive performance and reliable multi-step...

Ctx
1049K
In
$0.375 / 1M
Out
$1.88 / 1M
ByteDance Seed: Seed 2.1 Turbo
bytedance-seed/seed-2-1-turbo

Seed 2.1 Turbo is a multimodal model from ByteDance Seed for coding and long-horizon agent workflows. It is suited for end-to-end software delivery, multi-step task execution, and understanding visual and...

Ctx
262K
In
$0.500 / 1M
Out
$2.50 / 1M
Qwen: Qwen3.8 2.4T A95B
qwen/qwen3.8-2.4t-a95b

Qwen3.8 2.4T A95B is an open-weight sparse mixture-of-experts model from Qwen and the open-weight variant of [Qwen3.8 Max](/qwen/qwen3.8-max), with 95 billion active parameters out of 2.4 trillion total. It is...

Ctx
1049K
In
$2.00 / 1M
Out
$6.00 / 1M
ByteDance Seed: Seed-2.0-Code
bytedance-seed/seed-2.0-code

Seed 2.0 Code is a model from ByteDance Seed optimized for agentic coding. It is suited for frontend development, multilingual programming tasks, and coding-agent workflows in tools such as Claude...

Ctx
262K
In
$0.500 / 1M
Out
$3.00 / 1M
DeepSeek: DeepSeek V4 Pro 0813
deepseek/deepseek-v4-pro-0813

DeepSeek V4 Pro 0813 is a large-scale mixture-of-experts model from DeepSeek. This is the GA release of DeepSeek V4 Pro.

Ctx
1049K
In
$0.462 / 1M
Out
$1.39 / 1M
SpaceXAI: Grok 4.6
x-ai/grok-4.6

Grok 4.6 is a model from SpaceXAI with frontier performance on coding, knowledge work, and STEM. It is succeeded by [Grok 4.7](/x-ai/grok-4.7).

Ctx
500K
In
$2.00 / 1M
Out
$6.00 / 1M
LiquidAI: LFM2.5-2.6B (free)
liquid/lfm-2.5-2.6b:free

LFM2.5-2.6B is a compact reasoning model from Liquid AI. It is suited for agent workflows, data extraction, RAG, and long-context processing. Liquid advises against using it for agentic coding or...

Ctx
66K
In
Free
Out
Free
NVIDIA: Nemotron 3.5 Lightning
nvidia/nemotron-3.5-lightning

NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total. It is suited for high-throughput agentic workloads and specialized tasks that...

Ctx
262K
In
$0.080 / 1M
Out
$0.200 / 1M
NVIDIA: Nemotron 3.5 Lightning (free)
nvidia/nemotron-3.5-lightning:free

NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total. It is suited for high-throughput agentic workloads and specialized tasks that...

Ctx
1000K
In
Free
Out
Free
Sakana: Sakana Namazu
sakana/sakana-namazu

Sakana Namazu is a Japanese-specialized reasoning model from Sakana AI, based on Kimi K2.6 with additional training for Japanese language and business contexts. It is suited for Japanese instruction following,...

Ctx
262K
In
$0.950 / 1M
Out
$4.00 / 1M
Upstage: Solar Pro 4
upstage/solar-pro4

Solar Pro 4 is Upstage's cost-efficient large language model, featuring a 524K context window. It is built for long-horizon tasks and agentic workflows, with strong capabilities in office productivity, document-intensive...

Ctx
524K
In
$0.090 / 1M
Out
$0.360 / 1M
Meta: Muse Glimmer 30B
meta/muse-glimmer-30b

Muse Glimmer 30B is a dense, open-weight multimodal model from Meta Superintelligence Labs, distilled from Muse Spark and optimized for autonomous agents on consumer hardware. It is suited for long-horizon...

Ctx
131K
In
$0.300 / 1M
Out
$1.20 / 1M
Meta: Muse Spark 1.2
meta/muse-spark-1.2

Muse Spark 1.2 is a reasoning model from Meta, designed for complex agentic tasks. It accepts text, images, video, audio, and PDF documents, returns text, and offers a 1M-token context...

Ctx
1049K
In
$1.25 / 1M
Out
$4.25 / 1M
DeepSeek: DeepSeek V4 Flash Latest
~deepseek/deepseek-v4-flash-latest

This model always redirects to the latest model in the DeepSeek V4 Flash family.

Ctx
1311K
In
$0.038 / 1M
Out
$0.550 / 1M
DeepSeek: DeepSeek V4 Flash 0731
deepseek/deepseek-v4-flash-0731

DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows....

Ctx
1311K
In
$0.040 / 1M
Out
$0.640 / 1M
Thinking Machines: Inkling Small
thinkingmachines/inkling-small

Inkling Small is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 12B active parameters out of 276B total. It is positioned as the smaller, more efficient member of...

Ctx
1049K
In
$0.450 / 1M
Out
$1.20 / 1M
Thinking Machines: Inkling Small (free)
thinkingmachines/inkling-small:free

Inkling Small is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 12B active parameters out of 276B total. It is positioned as the smaller, more efficient member of...

Ctx
1049K
In
Free
Out
Free
Qwen: Qwen3.7 Flash
qwen/qwen3.7-flash

Qwen3.7 Flash is a vision-language reasoning model from Alibaba. It is suited for multimodal agents, visual coding, search, and computer interaction, with strengths in object recognition, spatial understanding, and real-world...

Ctx
1000K
In
$0.030 / 1M
Out
$0.130 / 1M
Anthropic: Claude Opus 5
anthropic/claude-opus-5

Claude Opus 5 is Anthropic’s flagship model for demanding reasoning, coding, and long-horizon agentic work. It is particularly strong at end-to-end software tasks, code review and bug finding, visual analysis...

Ctx
1000K
In
$5.00 / 1M
Out
$25.00 / 1M
Anthropic: Claude Opus 5 (batch)
anthropic/claude-opus-5:batch

Claude Opus 5 is Anthropic’s flagship model for demanding reasoning, coding, and long-horizon agentic work. It is particularly strong at end-to-end software tasks, code review and bug finding, visual analysis...

Ctx
1000K
In
$2.50 / 1M
Out
$12.50 / 1M
inclusionAI: Ling 3.0 Flash
inclusionai/ling-3.0-flash

*Ling-3.0-flash* is a *124B-parameter Mixture-of-Experts (MoE) model*, with approximately *5.1B parameters activated per token*. The model is designed with *token efficiency and production-scale agentic inference* as key priorities, enabling developers...

Ctx
262K
In
$0.021 / 1M
Out
$0.063 / 1M
Poolside: Laguna S 2.1
poolside/laguna-s-2.1

Laguna S 2.1 is the latest coding agent model from [Poolside](<https://poolside.ai/>). Laguna S 2.1 is a 118B total parameter model with 8B active parameters, scoring 70.2% on Terminal-Bench 2.1 and...

Ctx
1049K
In
$0.090 / 1M
Out
$0.180 / 1M
Poolside: Laguna S 2.1 (free)
poolside/laguna-s-2.1:free

Laguna S 2.1 is the latest coding agent model from [Poolside](<https://poolside.ai/>). Laguna S 2.1 is a 118B total parameter model with 8B active parameters, scoring 70.2% on Terminal-Bench 2.1 and...

Ctx
262K
In
Free
Out
Free
Google: Gemini 3.6 Flash
google/gemini-3.6-flash

Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to produce polished outputs with fewer unnecessary edits and...

Ctx
1049K
In
$0.750 / 1M
Out
$3.75 / 1M
Google: Gemini 3.6 Flash (batch)
google/gemini-3.6-flash:batch

Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to produce polished outputs with fewer unnecessary edits and...

Ctx
1049K
In
$0.375 / 1M
Out
$1.88 / 1M
Google: Gemini 3.5 Flash Lite
google/gemini-3.5-flash-lite

Gemini 3.5 Flash Lite is a high-efficiency model from Google with upgraded agentic capabilities. It is suited for subagents that execute focused tasks within complex, multi-agent workflows.

Ctx
1049K
In
$0.300 / 1M
Out
$2.50 / 1M
Google: Gemini 3.5 Flash Lite (batch)
google/gemini-3.5-flash-lite:batch

Gemini 3.5 Flash Lite is a high-efficiency model from Google with upgraded agentic capabilities. It is suited for subagents that execute focused tasks within complex, multi-agent workflows.

Ctx
1049K
In
$0.150 / 1M
Out
$1.25 / 1M
Meituan: LongCat 2.0
meituan/longcat-2.0

LongCat 2.0 is a sparse mixture-of-experts language model from Meituan, with 48B active parameters out of 1.6T total. It is suited for coding, repository-level changes, long-horizon problem solving, and agentic...

Ctx
1049K
In
$0.300 / 1M
Out
$1.20 / 1M
Thinking Machines: Inkling
thinkingmachines/inkling

Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning, coding, agentic and tool-use systems,...

Ctx
1049K
In
$1.00 / 1M
Out
$4.05 / 1M
Thinking Machines: Inkling (free)
thinkingmachines/inkling:free

Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning, coding, agentic and tool-use systems,...

Ctx
1049K
In
Free
Out
Free
Auto Router (Beta)
openrouter/auto-beta

The experimental version of our Auto Router where we test new improvements. Use it to get the latest and greatest version of our general purpose auto router, but expect beta...

Ctx
2000K
In
$-1000000.000 / 1M
Out
$-1000000.000 / 1M
MoonshotAI: Kimi K3
moonshotai/kimi-k3

Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic workflows, and is particularly strong at...

Ctx
1049K
In
$3.00 / 1M
Out
$15.00 / 1M
MoonshotAI: Kimi K3 (batch)
moonshotai/kimi-k3:batch

Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic workflows, and is particularly strong at...

Ctx
1049K
In
$2.28 / 1M
Out
$11.40 / 1M
Meta: Muse Spark 1.1
meta/muse-spark-1.1

Muse Spark 1.1 is a multimodal reasoning model from Meta, built for agentic tasks. It accepts text, images, video, audio, and PDF documents and returns text, with a 1M-token context...

Ctx
1049K
In
$1.25 / 1M
Out
$4.25 / 1M
Kwaipilot: KAT-Coder-Pro V2.5
kwaipilot/kat-coder-pro-v2.5

KAT-Coder-Pro V2.5 is a flagship-level Agentic Coding model that can directly hand over an entire issue or an entire business workflow to it, allowing it to autonomously locate and make...

Ctx
262K
In
$0.740 / 1M
Out
$2.96 / 1M
OpenAI: GPT-5.6 Luna Pro
openai/gpt-5.6-luna-pro

GPT-5.6 Luna Pro is the same underlying model as [GPT-5.6 Luna](https://openrouter.ai/openai/gpt-5.6-luna), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode

Ctx
1050K
In
$0.200 / 1M
Out
$1.20 / 1M
OpenAI: GPT-5.6 Luna Pro (batch)
openai/gpt-5.6-luna-pro:batch

GPT-5.6 Luna Pro is the same underlying model as [GPT-5.6 Luna](https://openrouter.ai/openai/gpt-5.6-luna), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode

Ctx
1050K
In
$0.100 / 1M
Out
$0.600 / 1M
OpenAI: GPT-5.6 Luna
openai/gpt-5.6-luna

GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series. It is suited for high-volume, latency-sensitive tasks such as chat, classification, and lightweight agentic workflows, providing capable reasoning for...

Ctx
1050K
In
$0.200 / 1M
Out
$1.20 / 1M
OpenAI: GPT-5.6 Luna (batch)
openai/gpt-5.6-luna:batch

GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series. It is suited for high-volume, latency-sensitive tasks such as chat, classification, and lightweight agentic workflows, providing capable reasoning for...

Ctx
1050K
In
$0.100 / 1M
Out
$0.600 / 1M
OpenAI: GPT-5.6 Terra Pro
openai/gpt-5.6-terra-pro

GPT-5.6 Terra Pro is the same underlying model as [GPT-5.6 Terra](https://openrouter.ai/openai/gpt-5.6-terra), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode

Ctx
1050K
In
$2.00 / 1M
Out
$12.00 / 1M
OpenAI: GPT-5.6 Terra Pro (batch)
openai/gpt-5.6-terra-pro:batch

GPT-5.6 Terra Pro is the same underlying model as [GPT-5.6 Terra](https://openrouter.ai/openai/gpt-5.6-terra), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode

Ctx
1050K
In
$1.00 / 1M
Out
$6.00 / 1M
OpenAI: GPT-5.6 Terra
openai/gpt-5.6-terra

GPT-5.6 Terra is a balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier. It is suited for everyday coding, reasoning, and agentic...

Ctx
1050K
In
$2.00 / 1M
Out
$12.00 / 1M
OpenAI: GPT-5.6 Terra (batch)
openai/gpt-5.6-terra:batch

GPT-5.6 Terra is a balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier. It is suited for everyday coding, reasoning, and agentic...

Ctx
1050K
In
$1.00 / 1M
Out
$6.00 / 1M
OpenAI: GPT-5.6 Sol Pro
openai/gpt-5.6-sol-pro

GPT-5.6 Sol Pro is the same underlying model as [GPT-5.6 Sol](https://openrouter.ai/openai/gpt-5.6-sol), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode

Ctx
1050K
In
$2.00 / 1M
Out
$10.00 / 1M
OpenAI: GPT-5.6 Sol Pro (batch)
openai/gpt-5.6-sol-pro:batch

GPT-5.6 Sol Pro is the same underlying model as [GPT-5.6 Sol](https://openrouter.ai/openai/gpt-5.6-sol), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode

Ctx
1050K
In
$1.00 / 1M
Out
$5.00 / 1M
Showing first 120 — refine search to narrow results.

Unified access layer

One key. Every model on the network.

Stop juggling per-vendor SDKs and quotas — TokensChain offers an OpenAI-compatible front door with smart routing baked in.
tokenschain · preview