🌐 2025/2026 Model Catalog

Supported Models & Smart Routing

Access frontier reasoning models, ultra-fast coding LLMs, multimodal vision, and embedding endpoints through a unified interface with automatic provider failover.

1. Flagship Models

Model SlugProviderContext WindowInput ($/M)Output ($/M)
mistral/ministral-14b
Mistral: Ministral 14B
mistral262k ctx$0.20$0.20
mistral/ministral-8b
Mistral: Ministral 8B
mistral262k ctx$0.15$0.15
mistral/ministral-3b
Mistral: Ministral 3B
mistral131k ctx$0.10$0.10
mistral/mistral-embed
Mistral: Mistral Embed
mistral8k ctx$0.10TBD
mistral/codestral-embed
Mistral: Codestral Embed
mistral8k ctx$0.15TBD
openai/text-embedding-3-small
OpenAI: text-embedding-3-small
openai8k ctx$0.02TBD
openai/text-embedding-3-large
OpenAI: text-embedding-3-large
openai8k ctx$0.13TBD
voyage/voyage-4
Voyage: voyage-4
voyage32k ctx$0.06TBD
voyage/voyage-4-large
Voyage: voyage-4-large
voyage32k ctx$0.12TBD
voyage/voyage-4-lite
Voyage: voyage-4-lite
voyage32k ctx$0.02TBD
voyage/voyage-code-4
Voyage: voyage-code-4
voyage32k ctx$0.12TBD
voyage/voyage-finance-2
Voyage: voyage-finance-2
voyage32k ctx$0.12TBD
voyage/voyage-law-2
Voyage: voyage-law-2
voyage16k ctx$0.12TBD
anthropic/claude-fable-5
Anthropic: Claude Fable 5
anthropic1000k ctx$10.00$50.00
anthropic/claude-sonnet-5
Anthropic: Claude Sonnet 5
anthropic1000k ctx$2.00$10.00
anthropic/claude-opus-5
Anthropic: Claude Opus 5
anthropic1000k ctx$5.00$25.00
anthropic/claude-haiku-4-5
Anthropic: Claude Haiku 4.5
anthropic200k ctx$1.00$5.00
anthropic/claude-opus-4-8
Anthropic: Claude Opus 4.8
anthropic1000k ctx$5.00$25.00
anthropic/claude-opus-4-7
Anthropic: Claude Opus 4.7
anthropic1000k ctx$5.00$25.00
anthropic/claude-opus-4-6
Anthropic: Claude Opus 4.6
anthropic1000k ctx$5.00$25.00
anthropic/claude-sonnet-4-6
Anthropic: Claude Sonnet 4.6
anthropic1000k ctx$3.00$15.00
anthropic/claude-opus-4-5
Anthropic: Claude Opus 4.5
anthropic200k ctx$5.00$25.00
anthropic/claude-sonnet-4-5
Anthropic: Claude Sonnet 4.5
anthropic200k ctx$3.00$15.00
openai/gpt-5.6-sol
OpenAI: GPT-5.6 Sol
openai272k ctx$5.00$30.00
openai/gpt-5.6-terra
OpenAI: GPT-5.6 Terra
openai272k ctx$2.00$12.00
openai/gpt-5.6-luna
OpenAI: GPT-5.6 Luna
openai272k ctx$0.20$1.20
openai/gpt-5.5
OpenAI: GPT-5.5
openai272k ctx$5.00$30.00
openai/gpt-5.4
OpenAI: GPT-5.4
openai272k ctx$2.50$15.00
openai/gpt-5.4-mini
OpenAI: GPT-5.4 mini
openai400k ctx$0.75$4.50
openai/gpt-5.4-nano
OpenAI: GPT-5.4 nano
openai400k ctx$0.20$1.25
openai/gpt-5.2
OpenAI: GPT-5.2
openai400k ctx$1.75$14.00
openai/gpt-5.1
OpenAI: GPT-5.1
openai400k ctx$1.25$10.00
openai/gpt-5
OpenAI: GPT-5
openai400k ctx$1.25$10.00
openai/gpt-5-mini
OpenAI: GPT-5 mini
openai400k ctx$0.25$2.00
openai/gpt-5-nano
OpenAI: GPT-5 nano
openai400k ctx$0.05$0.40
openai/gpt-4.1
OpenAI: GPT-4.1
openai1048k ctx$2.00$8.00
openai/gpt-4.1-mini
OpenAI: GPT-4.1 mini
openai1048k ctx$0.40$1.60
openai/gpt-4.1-nano
OpenAI: GPT-4.1 nano
openai1048k ctx$0.10$0.40
openai/gpt-4o
OpenAI: GPT-4o
openai128k ctx$2.50$10.00
openai/gpt-4o-mini
OpenAI: GPT-4o mini
openai128k ctx$0.15$0.60
openai/o1
OpenAI: o1
openai200k ctx$15.00$60.00
openai/o3-mini
OpenAI: o3-mini
openai200k ctx$1.10$4.40
openai/o3
OpenAI: o3
openai200k ctx$2.00$8.00
openai/o4-mini
OpenAI: o4-mini
openai200k ctx$1.10$4.40
deepseek/deepseek-v4-flash
DeepSeek: DeepSeek V4 Flash
deepseek1000k ctx$0.44$1.32
deepseek/deepseek-v4-pro
DeepSeek: DeepSeek V4 Pro
deepseek1000k ctx$1.32$3.96
moonshot/kimi-k3
Moonshot: Kimi K3
moonshot1049k ctx$3.00$15.00
moonshot/kimi-k2.7-code
Moonshot: Kimi K2.7 Code
moonshot262k ctx$0.95$4.00
moonshot/kimi-k2.7-code-highspeed
Moonshot: Kimi K2.7 Code (high speed)
moonshot262k ctx$1.90$8.00
moonshot/kimi-k2.6
Moonshot: Kimi K2.6
moonshot262k ctx$0.95$4.00
mistral/mistral-large-3
Mistral: Mistral Large 3
mistral262k ctx$0.50$1.50
mistral/mistral-medium-3.5
Mistral: Mistral Medium 3.5
mistral262k ctx$1.50$7.50
mistral/mistral-small-4
Mistral: Mistral Small 4
mistral262k ctx$0.15$0.60
mistral/codestral
Mistral: Codestral
mistral256k ctx$0.30$0.90
mistral/glm-5.2
Mistral: GLM 5.2
mistral1049k ctx$1.40$4.40

2. Virtual Smart Routing Slugs

Instead of hardcoding a provider or specific model, route dynamically by use case:

nicrron/auto

Dynamically evaluates prompt token length and complexity to route to either heavy reasoning or lightweight fast models.

nicrron/reasoning

Prioritizes DeepSeek V4 Pro, OpenAI o3-mini, and Claude Opus 5 for multi-step reasoning, mathematical derivations, and architecture design.

nicrron/coding

Prioritizes Claude Opus 5, Claude Sonnet 5, and Kimi K2.7 Code for syntax precision, refactoring, and code generation.

nicrron/fastest

Prioritizes Claude Haiku 4.5, GPT-5.6 Luna, and DeepSeek V4 Flash for the lowest TTFT and highest tokens/second.

3. Automatic Prompt Caching Pass-Through

Nicrron records upstream prompt cache hits directly in response telemetry and billing settlements. You automatically receive the full provider discount without extra configuration:

  • Anthropic Cache Read: 90% discount on prompt tokens (0.1x base prompt rate).
  • OpenAI Cache Read: 50% discount on prompt tokens (0.5x base prompt rate).
  • Telemetred Usage: Inspect usage.cache_read_tokens in your API response or dashboard logs.

4. Provider Data Handling & Privacy Disclosures

Your prompts and completions are handled by whichever upstream provider serves the model you called, under that provider's policy rather than ours. Two things differ between them and are worth reading separately: training (whether your content is used to improve their models) and retention (how long they keep it). On our own side, we store the token counts needed to bill you and never the text of your prompts. Completions are cached for up to 24 hours so a repeated request does not pay twice, keyed by a hash of the full request and scoped to your account β€” a cached completion is only ever replayed to the account that generated it. Caching applies only when you send temperature: 0 or opt in with x-nicrron-cache: true. Send x-nicrron-cache: false to disable it for a request. Cache hits are free and appear in your logs as cached.

  • OpenAI, Anthropic: Not used for training by default. OpenAI: β€œdata sent to the OpenAI API is not used to train or improve OpenAI models (unless you explicitly opt in).” Anthropic: β€œBy default, we will not use your inputs or outputs from our commercial products … to train our models.” Both still retain content for a period for abuse monitoring β€” OpenAI states up to 30 days; Anthropic publishes no single figure. Zero retention is a separately negotiated arrangement that Nicrron does not hold, so these endpoints are no-training but not zero-retention.
  • Mistral: Nicrron has opted out of training on its Mistral account, so content you send to Mistral models is not used to improve them. One carve-out we cannot opt out of: their commercial terms Β§4.3 place Labs and Preview models outside any opt-out. We do not currently route to any Labs or Preview model, and will note it here if that changes.
  • Moonshot (Kimi): Their terms state Moonshot β€œmay use Content to provide, maintain, develop, support, and improve the Services.” Assume prompts and completions sent to Kimi models may be used for model improvement. Opting out requires an enterprise agreement we do not hold.
  • DeepSeek: DeepSeek publishes no no-training commitment covering API traffic, and its policy lists foundation model training among the purposes personal data may be processed for. Data is processed in the People's Republic of China. Treat DeepSeek as both trained on and retained.
  • Voyage (embeddings): Voyage's policy takes a licence to train on Customer Content unless the account holder opts out. Nicrron has opted out, which under their terms also means zero-day retention β€” text you embed through Voyage is neither trained on nor stored by them.
  • xAI (Grok): We have not found a published xAI statement covering training or retention for their API, as opposed to the consumer Grok apps. Treat it as unresolved rather than as either commitment.

Summarised from each provider's published terms. They change without notice, and the provider's own policy governs β€” this page is a convenience, not a warranty. If your workload cannot tolerate training or retention, pin your requests to a specific model: automatic failover may route you to a different provider than the one you chose.