Supported Models & Smart Routing
Access frontier reasoning models, ultra-fast coding LLMs, multimodal vision, and embedding endpoints through a unified interface with automatic provider failover.
1. Flagship Models
| Model Slug | Provider | Context Window | Input ($/M) | Output ($/M) |
|---|---|---|---|---|
mistral/ministral-14b Mistral: Ministral 14B | mistral | 262k ctx | $0.20 | $0.20 |
mistral/ministral-8b Mistral: Ministral 8B | mistral | 262k ctx | $0.15 | $0.15 |
mistral/ministral-3b Mistral: Ministral 3B | mistral | 131k ctx | $0.10 | $0.10 |
mistral/mistral-embed Mistral: Mistral Embed | mistral | 8k ctx | $0.10 | TBD |
mistral/codestral-embed Mistral: Codestral Embed | mistral | 8k ctx | $0.15 | TBD |
openai/text-embedding-3-small OpenAI: text-embedding-3-small | openai | 8k ctx | $0.02 | TBD |
openai/text-embedding-3-large OpenAI: text-embedding-3-large | openai | 8k ctx | $0.13 | TBD |
voyage/voyage-4 Voyage: voyage-4 | voyage | 32k ctx | $0.06 | TBD |
voyage/voyage-4-large Voyage: voyage-4-large | voyage | 32k ctx | $0.12 | TBD |
voyage/voyage-4-lite Voyage: voyage-4-lite | voyage | 32k ctx | $0.02 | TBD |
voyage/voyage-code-4 Voyage: voyage-code-4 | voyage | 32k ctx | $0.12 | TBD |
voyage/voyage-finance-2 Voyage: voyage-finance-2 | voyage | 32k ctx | $0.12 | TBD |
voyage/voyage-law-2 Voyage: voyage-law-2 | voyage | 16k ctx | $0.12 | TBD |
anthropic/claude-fable-5 Anthropic: Claude Fable 5 | anthropic | 1000k ctx | $10.00 | $50.00 |
anthropic/claude-sonnet-5 Anthropic: Claude Sonnet 5 | anthropic | 1000k ctx | $2.00 | $10.00 |
anthropic/claude-opus-5 Anthropic: Claude Opus 5 | anthropic | 1000k ctx | $5.00 | $25.00 |
anthropic/claude-haiku-4-5 Anthropic: Claude Haiku 4.5 | anthropic | 200k ctx | $1.00 | $5.00 |
anthropic/claude-opus-4-8 Anthropic: Claude Opus 4.8 | anthropic | 1000k ctx | $5.00 | $25.00 |
anthropic/claude-opus-4-7 Anthropic: Claude Opus 4.7 | anthropic | 1000k ctx | $5.00 | $25.00 |
anthropic/claude-opus-4-6 Anthropic: Claude Opus 4.6 | anthropic | 1000k ctx | $5.00 | $25.00 |
anthropic/claude-sonnet-4-6 Anthropic: Claude Sonnet 4.6 | anthropic | 1000k ctx | $3.00 | $15.00 |
anthropic/claude-opus-4-5 Anthropic: Claude Opus 4.5 | anthropic | 200k ctx | $5.00 | $25.00 |
anthropic/claude-sonnet-4-5 Anthropic: Claude Sonnet 4.5 | anthropic | 200k ctx | $3.00 | $15.00 |
openai/gpt-5.6-sol OpenAI: GPT-5.6 Sol | openai | 272k ctx | $5.00 | $30.00 |
openai/gpt-5.6-terra OpenAI: GPT-5.6 Terra | openai | 272k ctx | $2.00 | $12.00 |
openai/gpt-5.6-luna OpenAI: GPT-5.6 Luna | openai | 272k ctx | $0.20 | $1.20 |
openai/gpt-5.5 OpenAI: GPT-5.5 | openai | 272k ctx | $5.00 | $30.00 |
openai/gpt-5.4 OpenAI: GPT-5.4 | openai | 272k ctx | $2.50 | $15.00 |
openai/gpt-5.4-mini OpenAI: GPT-5.4 mini | openai | 400k ctx | $0.75 | $4.50 |
openai/gpt-5.4-nano OpenAI: GPT-5.4 nano | openai | 400k ctx | $0.20 | $1.25 |
openai/gpt-5.2 OpenAI: GPT-5.2 | openai | 400k ctx | $1.75 | $14.00 |
openai/gpt-5.1 OpenAI: GPT-5.1 | openai | 400k ctx | $1.25 | $10.00 |
openai/gpt-5 OpenAI: GPT-5 | openai | 400k ctx | $1.25 | $10.00 |
openai/gpt-5-mini OpenAI: GPT-5 mini | openai | 400k ctx | $0.25 | $2.00 |
openai/gpt-5-nano OpenAI: GPT-5 nano | openai | 400k ctx | $0.05 | $0.40 |
openai/gpt-4.1 OpenAI: GPT-4.1 | openai | 1048k ctx | $2.00 | $8.00 |
openai/gpt-4.1-mini OpenAI: GPT-4.1 mini | openai | 1048k ctx | $0.40 | $1.60 |
openai/gpt-4.1-nano OpenAI: GPT-4.1 nano | openai | 1048k ctx | $0.10 | $0.40 |
openai/gpt-4o OpenAI: GPT-4o | openai | 128k ctx | $2.50 | $10.00 |
openai/gpt-4o-mini OpenAI: GPT-4o mini | openai | 128k ctx | $0.15 | $0.60 |
openai/o1 OpenAI: o1 | openai | 200k ctx | $15.00 | $60.00 |
openai/o3-mini OpenAI: o3-mini | openai | 200k ctx | $1.10 | $4.40 |
openai/o3 OpenAI: o3 | openai | 200k ctx | $2.00 | $8.00 |
openai/o4-mini OpenAI: o4-mini | openai | 200k ctx | $1.10 | $4.40 |
deepseek/deepseek-v4-flash DeepSeek: DeepSeek V4 Flash | deepseek | 1000k ctx | $0.44 | $1.32 |
deepseek/deepseek-v4-pro DeepSeek: DeepSeek V4 Pro | deepseek | 1000k ctx | $1.32 | $3.96 |
moonshot/kimi-k3 Moonshot: Kimi K3 | moonshot | 1049k ctx | $3.00 | $15.00 |
moonshot/kimi-k2.7-code Moonshot: Kimi K2.7 Code | moonshot | 262k ctx | $0.95 | $4.00 |
moonshot/kimi-k2.7-code-highspeed Moonshot: Kimi K2.7 Code (high speed) | moonshot | 262k ctx | $1.90 | $8.00 |
moonshot/kimi-k2.6 Moonshot: Kimi K2.6 | moonshot | 262k ctx | $0.95 | $4.00 |
mistral/mistral-large-3 Mistral: Mistral Large 3 | mistral | 262k ctx | $0.50 | $1.50 |
mistral/mistral-medium-3.5 Mistral: Mistral Medium 3.5 | mistral | 262k ctx | $1.50 | $7.50 |
mistral/mistral-small-4 Mistral: Mistral Small 4 | mistral | 262k ctx | $0.15 | $0.60 |
mistral/codestral Mistral: Codestral | mistral | 256k ctx | $0.30 | $0.90 |
mistral/glm-5.2 Mistral: GLM 5.2 | mistral | 1049k ctx | $1.40 | $4.40 |
2. Virtual Smart Routing Slugs
Instead of hardcoding a provider or specific model, route dynamically by use case:
Dynamically evaluates prompt token length and complexity to route to either heavy reasoning or lightweight fast models.
Prioritizes DeepSeek V4 Pro, OpenAI o3-mini, and Claude Opus 5 for multi-step reasoning, mathematical derivations, and architecture design.
Prioritizes Claude Opus 5, Claude Sonnet 5, and Kimi K2.7 Code for syntax precision, refactoring, and code generation.
Prioritizes Claude Haiku 4.5, GPT-5.6 Luna, and DeepSeek V4 Flash for the lowest TTFT and highest tokens/second.
3. Automatic Prompt Caching Pass-Through
Nicrron records upstream prompt cache hits directly in response telemetry and billing settlements. You automatically receive the full provider discount without extra configuration:
- Anthropic Cache Read: 90% discount on prompt tokens (
0.1xbase prompt rate). - OpenAI Cache Read: 50% discount on prompt tokens (
0.5xbase prompt rate). - Telemetred Usage: Inspect
usage.cache_read_tokensin your API response or dashboard logs.
4. Provider Data Handling & Privacy Disclosures
Your prompts and completions are handled by whichever upstream provider serves the model you called, under that provider's policy rather than ours. Two things differ between them and are worth reading separately: training (whether your content is used to improve their models) and retention (how long they keep it). On our own side, we store the token counts needed to bill you and never the text of your prompts. Completions are cached for up to 24 hours so a repeated request does not pay twice, keyed by a hash of the full request and scoped to your account β a cached completion is only ever replayed to the account that generated it. Caching applies only when you send temperature: 0 or opt in with x-nicrron-cache: true. Send x-nicrron-cache: false to disable it for a request. Cache hits are free and appear in your logs as cached.
- OpenAI, Anthropic: Not used for training by default. OpenAI: βdata sent to the OpenAI API is not used to train or improve OpenAI models (unless you explicitly opt in).β Anthropic: βBy default, we will not use your inputs or outputs from our commercial products β¦ to train our models.β Both still retain content for a period for abuse monitoring β OpenAI states up to 30 days; Anthropic publishes no single figure. Zero retention is a separately negotiated arrangement that Nicrron does not hold, so these endpoints are no-training but not zero-retention.
- Mistral: Nicrron has opted out of training on its Mistral account, so content you send to Mistral models is not used to improve them. One carve-out we cannot opt out of: their commercial terms Β§4.3 place Labs and Preview models outside any opt-out. We do not currently route to any Labs or Preview model, and will note it here if that changes.
- Moonshot (Kimi): Their terms state Moonshot βmay use Content to provide, maintain, develop, support, and improve the Services.β Assume prompts and completions sent to Kimi models may be used for model improvement. Opting out requires an enterprise agreement we do not hold.
- DeepSeek: DeepSeek publishes no no-training commitment covering API traffic, and its policy lists foundation model training among the purposes personal data may be processed for. Data is processed in the People's Republic of China. Treat DeepSeek as both trained on and retained.
- Voyage (embeddings): Voyage's policy takes a licence to train on Customer Content unless the account holder opts out. Nicrron has opted out, which under their terms also means zero-day retention β text you embed through Voyage is neither trained on nor stored by them.
- xAI (Grok): We have not found a published xAI statement covering training or retention for their API, as opposed to the consumer Grok apps. Treat it as unresolved rather than as either commitment.
Summarised from each provider's published terms. They change without notice, and the provider's own policy governs β this page is a convenience, not a warranty. If your workload cannot tolerate training or retention, pin your requests to a specific model: automatic failover may route you to a different provider than the one you chose.