realtime
Realtime
p95 < 300ms TTFT
At direct rate
Direct passthrough to the upstream provider, full precision. The only tier with a hard delivery guarantee on closed-weight models. Use when latency is non-negotiable.
We never charge more than the provider would charge you direct — at least 5% off from your first call, climbing as Atlas ramps, and you keep the savings. Don’t like a call? Thumbs-down it for a full refund, no questions asked. Start on free models with no card at all.
export OPENAI_BASE_URL=https://api.newmen.ai/v1 (or ANTHROPIC_BASE_URL=https://api.newmen.ai) plus a Newmen API key. That’s the migration. Card statement descriptor: NEWMEN.AI*CREDITS.
Pass any model id in the model field. Search the full provider-grouped catalogue and compare rates. Pay as you go is always below going direct — at least 5% off from call one. Reliability Loop adds the platform fee for the verification workflow.
311 models supported
| Model ID | Strictpin exact model | Atlas modeauto-optimize | |
|---|---|---|---|
| Most popular this week · top 12 | |||
DeepSeek: DeepSeek V4 Flash | $0.098 / 1Mvia Baidu | auto-optimized | |
Tencent: Hy3 preview | $0.063 / 1Mvia GMICloud | auto-optimized | |
Xiaomi: MiMo-V2.5 | $0.14 / 1Mvia Xiaomi | auto-optimized | |
Anthropic: Claude Sonnet 4.6 | $3.00 / 1Mvia Amazon Bedrock | auto-optimized | |
Anthropic: Claude Opus 4.7 | $5.00 / 1Mvia Amazon Bedrock | auto-optimized | |
DeepSeek: DeepSeek V4 Pro | $0.435 / 1Mvia DeepSeek | auto-optimized | |
MiniMax: MiniMax M3 | $0.3 / 1Mvia Minimax | auto-optimized | |
Xiaomi: MiMo-V2.5-Pro | $0.435 / 1Mvia Xiaomi | auto-optimized | |
DeepSeek: DeepSeek V3.2 | $0.229 / 1Mvia StreamLake | auto-optimized | |
Google: Gemini 3 Flash Preview | $0.5 / 1Mvia Google | auto-optimized | |
Anthropic: Claude Opus 4.8 | $5.00 / 1Mvia Amazon Bedrock | auto-optimized | |
NVIDIA: Nemotron 3 Super | $0.09 / 1Mvia DekaLLM | auto-optimized | |
| OpenAI · 61 models | |||
OpenAI: GPT Audio | $2.50 / 1Mvia OpenAI | auto-optimized | |
OpenAI: GPT Audio Mini | $0.6 / 1Mvia OpenAI | auto-optimized | |
OpenAI: GPT Chat Latest | $5.00 / 1Mvia OpenAI | auto-optimized | |
OpenAI: GPT-3.5 Turbo | $0.5 / 1Mvia OpenAI | auto-optimized | |
OpenAI: GPT-3.5 Turbo (older v0613) | $1.00 / 1Mvia Azure | auto-optimized | |
OpenAI: GPT-3.5 Turbo 16k | $3.00 / 1Mvia Azure | auto-optimized | |
OpenAI: GPT-3.5 Turbo Instruct | $1.50 / 1Mvia OpenAI | auto-optimized | |
OpenAI: GPT-4 | $30.00 / 1Mvia Azure | auto-optimized | |
OpenAI: GPT-4 Turbo | $10.00 / 1Mvia OpenAI | auto-optimized | |
OpenAI: GPT-4 Turbo (older v1106) | $10.00 / 1Mvia OpenAI | auto-optimized | |
OpenAI: GPT-4 Turbo Preview | $10.00 / 1Mvia OpenAI | auto-optimized | |
OpenAI: GPT-4.1 | $2.00 / 1Mvia Azure | auto-optimized | |
OpenAI: GPT-4.1 Mini | $0.4 / 1Mvia Azure | auto-optimized | |
OpenAI: GPT-4.1 Nano | $0.1 / 1Mvia Azure | auto-optimized | |
OpenAI: GPT-4o | $2.50 / 1Mvia Azure | auto-optimized | |
OpenAI: GPT-4o (2024-05-13) | $5.00 / 1Mvia Azure | auto-optimized | |
OpenAI: GPT-4o (2024-08-06) | $2.50 / 1Mvia Azure | auto-optimized | |
OpenAI: GPT-4o (2024-11-20) | $2.50 / 1Mvia OpenAI | auto-optimized | |
OpenAI: GPT-4o Search Preview | $2.50 / 1Mvia OpenAI | auto-optimized | |
OpenAI: GPT-4o-mini | $0.15 / 1Mvia Azure | auto-optimized | |
OpenAI: GPT-4o-mini (2024-07-18) | $0.15 / 1Mvia OpenAI | auto-optimized | |
OpenAI: GPT-4o-mini Search Preview | $0.15 / 1Mvia OpenAI | auto-optimized | |
OpenAI: GPT-5 | $1.25 / 1Mvia Azure | auto-optimized | |
OpenAI: GPT-5 Chat | $1.25 / 1Mvia OpenAI | auto-optimized | |
OpenAI: GPT-5 Codex | $1.25 / 1Mvia OpenAI | auto-optimized | |
OpenAI: GPT-5 Image | $10.00 / 1Mvia OpenAI | auto-optimized | |
OpenAI: GPT-5 Image Mini | $2.50 / 1Mvia OpenAI | auto-optimized | |
OpenAI: GPT-5 Mini | $0.25 / 1Mvia Azure | auto-optimized | |
OpenAI: GPT-5 Nano | $0.05 / 1Mvia Azure | auto-optimized | |
OpenAI: GPT-5 Pro | $15.00 / 1Mvia OpenAI | auto-optimized | |
OpenAI: GPT-5.1 | $1.25 / 1Mvia Azure | auto-optimized | |
OpenAI: GPT-5.1 Chat | $1.25 / 1Mvia Azure | auto-optimized | |
OpenAI: GPT-5.1-Codex | $1.25 / 1Mvia Azure | auto-optimized | |
OpenAI: GPT-5.1-Codex-Max | $1.25 / 1Mvia Azure | auto-optimized | |
OpenAI: GPT-5.1-Codex-Mini | $0.25 / 1Mvia Azure | auto-optimized | |
OpenAI: GPT-5.2 | $1.75 / 1Mvia Azure | auto-optimized | |
OpenAI: GPT-5.2 Chat | $1.75 / 1Mvia Azure | auto-optimized | |
OpenAI: GPT-5.2 Pro | $21.00 / 1Mvia OpenAI | auto-optimized | |
OpenAI: GPT-5.2-Codex | $1.75 / 1Mvia Azure | auto-optimized | |
OpenAI: GPT-5.3 Chat | $1.75 / 1Mvia Azure | auto-optimized | |
OpenAI: GPT-5.3-Codex | $1.75 / 1Mvia Azure | auto-optimized | |
OpenAI: GPT-5.4 | $2.50 / 1Mvia Azure | auto-optimized | |
OpenAI: GPT-5.4 Image 2 | $8.00 / 1Mvia OpenAI | auto-optimized | |
OpenAI: GPT-5.4 Mini | $0.75 / 1Mvia Azure | auto-optimized | |
OpenAI: GPT-5.4 Nano | $0.2 / 1Mvia Azure | auto-optimized | |
OpenAI: GPT-5.4 Pro | $30.00 / 1Mvia Azure | auto-optimized | |
OpenAI: GPT-5.5 | $5.00 / 1Mvia Azure | auto-optimized | |
OpenAI: GPT-5.5 Pro | $30.00 / 1Mvia OpenAI | auto-optimized | |
OpenAI: gpt-oss-120b | $0.039 / 1Mvia DeepInfra | auto-optimized | |
OpenAI: gpt-oss-20b | $0.029 / 1Mvia DekaLLM | auto-optimized | |
OpenAI: gpt-oss-safeguard-20b | $0.075 / 1Mvia Groq | auto-optimized | |
OpenAI: o1 | $15.00 / 1Mvia OpenAI | auto-optimized | |
OpenAI: o1-pro | $150.00 / 1Mvia OpenAI | auto-optimized | |
OpenAI: o3 | $2.00 / 1Mvia OpenAI | auto-optimized | |
OpenAI: o3 Deep Research | $10.00 / 1Mvia OpenAI | auto-optimized | |
OpenAI: o3 Mini | $1.10 / 1Mvia OpenAI | auto-optimized | |
OpenAI: o3 Mini High | $1.10 / 1Mvia OpenAI | auto-optimized | |
OpenAI: o3 Pro | $20.00 / 1Mvia OpenAI | auto-optimized | |
OpenAI: o4 Mini | $1.10 / 1Mvia OpenAI | auto-optimized | |
OpenAI: o4 Mini Deep Research | $2.00 / 1Mvia OpenAI | auto-optimized | |
OpenAI: o4 Mini High | $1.10 / 1Mvia OpenAI | auto-optimized | |
| Anthropic · 15 models | |||
Anthropic: Claude 3 Haiku | $0.25 / 1Mvia Amazon Bedrock | auto-optimized | |
Anthropic: Claude 3.5 Haiku | $0.8 / 1Mvia Amazon Bedrock | auto-optimized | |
Anthropic: Claude Haiku 4.5 | $1.00 / 1Mvia Amazon Bedrock | auto-optimized | |
Anthropic: Claude Opus 4 | $15.00 / 1Mvia Anthropic | auto-optimized | |
Anthropic: Claude Opus 4.1 | $15.00 / 1Mvia Amazon Bedrock | auto-optimized | |
Anthropic: Claude Opus 4.5 | $5.00 / 1Mvia Amazon Bedrock | auto-optimized | |
Anthropic: Claude Opus 4.6 | $5.00 / 1Mvia Amazon Bedrock | auto-optimized | |
Anthropic: Claude Opus 4.6 (Fast) | $30.00 / 1Mvia Anthropic | auto-optimized | |
Anthropic: Claude Opus 4.7 | $5.00 / 1Mvia Amazon Bedrock | auto-optimized | |
Anthropic: Claude Opus 4.7 (Fast) | $30.00 / 1Mvia Anthropic | auto-optimized | |
Anthropic: Claude Opus 4.8 | $5.00 / 1Mvia Amazon Bedrock | auto-optimized | |
Anthropic: Claude Opus 4.8 (Fast) | $10.00 / 1Mvia Anthropic | auto-optimized | |
Anthropic: Claude Sonnet 4 | $3.00 / 1Mvia Amazon Bedrock | auto-optimized | |
Anthropic: Claude Sonnet 4.5 | $3.00 / 1Mvia Amazon Bedrock | auto-optimized | |
Anthropic: Claude Sonnet 4.6 | $3.00 / 1Mvia Amazon Bedrock | auto-optimized | |
| Google · 22 models | |||
Google: Gemini 2.5 Flash | $0.3 / 1Mvia Google | auto-optimized | |
Google: Gemini 2.5 Flash Lite | $0.1 / 1Mvia Google | auto-optimized | |
Google: Gemini 2.5 Flash Lite Preview 09-2025 | $0.1 / 1Mvia Google | auto-optimized | |
Google: Gemini 2.5 Pro | $1.25 / 1Mvia Google | auto-optimized | |
Google: Gemini 2.5 Pro Preview 05-06 | $1.25 / 1Mvia Google | auto-optimized | |
Google: Gemini 2.5 Pro Preview 06-05 | $1.25 / 1Mvia Google | auto-optimized | |
Google: Gemini 3 Flash Preview | $0.5 / 1Mvia Google | auto-optimized | |
Google: Gemini 3.1 Flash Lite | $0.25 / 1Mvia Google | auto-optimized | |
Google: Gemini 3.1 Flash Lite Preview | $0.25 / 1Mvia Google | auto-optimized | |
Google: Gemini 3.1 Pro Preview | $2.00 / 1Mvia Google | auto-optimized | |
Google: Gemini 3.1 Pro Preview Custom Tools | $2.00 / 1Mvia Google AI Studio | auto-optimized | |
Google: Gemini 3.5 Flash | $1.50 / 1Mvia Google | auto-optimized | |
Google: Gemma 2 27B | $0.65 / 1Mvia NextBit | auto-optimized | |
Google: Gemma 3 12B | $0.04 / 1Mvia DeepInfra | auto-optimized | |
Google: Gemma 3 27B | $0.08 / 1Mvia DeepInfra | auto-optimized | |
Google: Gemma 3 4B | $0.04 / 1Mvia DeepInfra | auto-optimized | |
Google: Gemma 3n 4B | $0.06 / 1Mvia Together | auto-optimized | |
Google: Gemma 4 26B A4B | $0.06 / 1Mvia DekaLLM | auto-optimized | |
Google: Gemma 4 31B | $0.12 / 1Mvia DeepInfra | auto-optimized | |
Google: Nano Banana (Gemini 2.5 Flash Image) | $0.3 / 1Mvia Google | auto-optimized | |
Google: Nano Banana 2 (Gemini 3.1 Flash Image Preview) | $0.5 / 1Mvia Google | auto-optimized | |
Google: Nano Banana Pro (Gemini 3 Pro Image Preview) | $2.00 / 1Mvia Google | auto-optimized | |
| Meta · 12 models | |||
Llama Guard 3 8B | $0.484 / 1Mvia Cloudflare | auto-optimized | |
Meta: Llama 3 70B Instruct | $0.51 / 1Mvia Novita | auto-optimized | |
Meta: Llama 3 8B Instruct | $0.04 / 1Mvia Novita | auto-optimized | |
Meta: Llama 3.1 70B Instruct | $0.4 / 1Mvia DeepInfra | auto-optimized | |
Meta: Llama 3.1 8B Instruct | $0.02 / 1Mvia DeepInfra | auto-optimized | |
Meta: Llama 3.2 11B Vision Instruct | $0.245 / 1Mvia DeepInfra | auto-optimized | |
Meta: Llama 3.2 1B Instruct | $0.027 / 1Mvia Cloudflare | auto-optimized | |
Meta: Llama 3.2 3B Instruct | $0.051 / 1Mvia Cloudflare | auto-optimized | |
Meta: Llama 3.3 70B Instruct | $0.1 / 1Mvia DeepInfra | auto-optimized | |
Meta: Llama 4 Maverick | $0.15 / 1Mvia DeepInfra | auto-optimized | |
Meta: Llama 4 Scout | $0.08 / 1Mvia DeepInfra | auto-optimized | |
Meta: Llama Guard 4 12B | $0.18 / 1Mvia DeepInfra | auto-optimized | |
| Mistral · 19 models | |||
Mistral Large | $2.00 / 1Mvia Mistral | auto-optimized | |
Mistral Large 2407 | $2.00 / 1Mvia Mistral | auto-optimized | |
Mistral: Codestral 2508 | $0.3 / 1Mvia Mistral | auto-optimized | |
Mistral: Devstral 2 2512 | $0.4 / 1Mvia Mistral | auto-optimized | |
Mistral: Ministral 3 14B 2512 | $0.2 / 1Mvia Mistral | auto-optimized | |
Mistral: Ministral 3 3B 2512 | $0.1 / 1Mvia Mistral | auto-optimized | |
Mistral: Ministral 3 8B 2512 | $0.15 / 1Mvia Mistral | auto-optimized | |
Mistral: Mistral Large 3 2512 | $0.5 / 1Mvia Mistral | auto-optimized | |
Mistral: Mistral Medium 3 | $0.4 / 1Mvia Mistral | auto-optimized | |
Mistral: Mistral Medium 3.1 | $0.4 / 1Mvia Mistral | auto-optimized | |
Mistral: Mistral Medium 3.5 | $1.50 / 1Mvia Mistral | auto-optimized | |
Mistral: Mistral Nemo | $0.02 / 1Mvia DeepInfra | auto-optimized | |
Mistral: Mistral Small 3 | $0.05 / 1Mvia DeepInfra | auto-optimized | |
Mistral: Mistral Small 3.1 24B | $0.351 / 1Mvia Cloudflare | auto-optimized | |
Mistral: Mistral Small 3.2 24B | $0.075 / 1Mvia DeepInfra | auto-optimized | |
Mistral: Mistral Small 4 | $0.15 / 1Mvia Mistral | auto-optimized | |
Mistral: Mixtral 8x22B Instruct | $2.00 / 1Mvia Mistral | auto-optimized | |
Mistral: Saba | $0.2 / 1Mvia Mistral | auto-optimized | |
Mistral: Voxtral Small 24B 2507 | $0.1 / 1Mvia Mistral | auto-optimized | |
| xAI · 4 models | |||
xAI: Grok 4.20 | $1.25 / 1Mvia xAI | auto-optimized | |
xAI: Grok 4.20 Multi-Agent | $2.00 / 1Mvia xAI | auto-optimized | |
xAI: Grok 4.3 | $1.25 / 1Mvia xAI | auto-optimized | |
xAI: Grok Build 0.1 | $1.00 / 1Mvia xAI | auto-optimized | |
| DeepSeek · 12 models | |||
DeepSeek: DeepSeek V3 | $0.2 / 1Mvia StreamLake | auto-optimized | |
DeepSeek: DeepSeek V3 0324 | $0.2 / 1Mvia DeepInfra | auto-optimized | |
DeepSeek: DeepSeek V3.1 | $0.21 / 1Mvia DeepInfra | auto-optimized | |
DeepSeek: DeepSeek V3.1 Terminus | $0.27 / 1Mvia DeepInfra | auto-optimized | |
DeepSeek: DeepSeek V3.2 | $0.229 / 1Mvia StreamLake | auto-optimized | |
DeepSeek: DeepSeek V3.2 Exp | $0.27 / 1Mvia AtlasCloud | auto-optimized | |
DeepSeek: DeepSeek V4 Flash | $0.098 / 1Mvia Baidu | auto-optimized | |
DeepSeek: DeepSeek V4 Pro | $0.435 / 1Mvia DeepSeek | auto-optimized | |
DeepSeek: R1 | $0.7 / 1Mvia Novita | auto-optimized | |
DeepSeek: R1 0528 | $0.5 / 1Mvia DeepInfra | auto-optimized | |
DeepSeek: R1 Distill Llama 70B | $0.7 / 1Mvia DeepInfra | auto-optimized | |
DeepSeek: R1 Distill Qwen 32B | $0.29 / 1Mvia NextBit | auto-optimized | |
| Alibaba (Qwen) · 47 models | |||
Qwen: Qwen Plus 0728 | $0.26 / 1Mvia Alibaba | auto-optimized | |
Qwen: Qwen Plus 0728 (thinking) | $0.26 / 1Mvia Alibaba | auto-optimized | |
Qwen: Qwen-Plus | $0.26 / 1Mvia Alibaba | auto-optimized | |
Qwen: Qwen2.5 7B Instruct | $0.04 / 1Mvia Phala | auto-optimized | |
Qwen: Qwen2.5 VL 72B Instruct | $0.25 / 1Mvia Nebius | auto-optimized | |
Qwen: Qwen3 14B | $0.1 / 1Mvia NextBit | auto-optimized | |
Qwen: Qwen3 235B A22B | $0.455 / 1Mvia Alibaba | auto-optimized | |
Qwen: Qwen3 235B A22B Instruct 2507 | $0.071 / 1Mvia DeepInfra | auto-optimized | |
Qwen: Qwen3 235B A22B Thinking 2507 | $0.1 / 1Mvia WandB | auto-optimized | |
Qwen: Qwen3 30B A3B | $0.09 / 1Mvia DeepInfra | auto-optimized | |
Qwen: Qwen3 30B A3B Instruct 2507 | $0.048 / 1Mvia StreamLake | auto-optimized | |
Qwen: Qwen3 30B A3B Thinking 2507 | $0.08 / 1Mvia AtlasCloud | auto-optimized | |
Qwen: Qwen3 32B | $0.08 / 1Mvia DeepInfra | auto-optimized | |
Qwen: Qwen3 8B | $0.05 / 1Mvia AtlasCloud | auto-optimized | |
Qwen: Qwen3 Coder 30B A3B Instruct | $0.07 / 1Mvia Novita | auto-optimized | |
Qwen: Qwen3 Coder 480B A35B | $0.22 / 1Mvia Google | auto-optimized | |
Qwen: Qwen3 Coder Flash | $0.195 / 1Mvia Alibaba | auto-optimized | |
Qwen: Qwen3 Coder Next | $0.11 / 1Mvia Ionstream | auto-optimized | |
Qwen: Qwen3 Coder Plus | $0.65 / 1Mvia Alibaba | auto-optimized | |
Qwen: Qwen3 Max | $0.78 / 1Mvia Alibaba | auto-optimized | |
Qwen: Qwen3 Max Thinking | $0.78 / 1Mvia Alibaba | auto-optimized | |
Qwen: Qwen3 Next 80B A3B Instruct | $0.09 / 1Mvia DeepInfra | auto-optimized | |
Qwen: Qwen3 Next 80B A3B Thinking | $0.098 / 1Mvia Alibaba | auto-optimized | |
Qwen: Qwen3 VL 235B A22B Instruct | $0.2 / 1Mvia DeepInfra | auto-optimized | |
Qwen: Qwen3 VL 235B A22B Thinking | $0.26 / 1Mvia Alibaba | auto-optimized | |
Qwen: Qwen3 VL 30B A3B Instruct | $0.13 / 1Mvia Alibaba | auto-optimized | |
Qwen: Qwen3 VL 30B A3B Thinking | $0.13 / 1Mvia Alibaba | auto-optimized | |
Qwen: Qwen3 VL 32B Instruct | $0.104 / 1Mvia Alibaba | auto-optimized | |
Qwen: Qwen3 VL 8B Instruct | $0.08 / 1Mvia AtlasCloud | auto-optimized | |
Qwen: Qwen3 VL 8B Thinking | $0.117 / 1Mvia Alibaba | auto-optimized | |
Qwen: Qwen3.5 397B A17B | $0.39 / 1Mvia Alibaba | auto-optimized | |
Qwen: Qwen3.5 Plus 2026-02-15 | $0.26 / 1Mvia Alibaba | auto-optimized | |
Qwen: Qwen3.5 Plus 2026-04-20 | $0.3 / 1Mvia Alibaba | auto-optimized | |
Qwen: Qwen3.5-122B-A10B | $0.26 / 1Mvia Alibaba | auto-optimized | |
Qwen: Qwen3.5-27B | $0.195 / 1Mvia Alibaba | auto-optimized | |
Qwen: Qwen3.5-35B-A3B | $0.14 / 1Mvia Ambient | auto-optimized | |
Qwen: Qwen3.5-9B | $0.04 / 1Mvia DeepInfra | auto-optimized | |
Qwen: Qwen3.5-Flash | $0.065 / 1Mvia Alibaba | auto-optimized | |
Qwen: Qwen3.6 27B | $0.29 / 1Mvia Io Net | auto-optimized | |
Qwen: Qwen3.6 35B A3B | $0.14 / 1Mvia Io Net | auto-optimized | |
Qwen: Qwen3.6 Flash | $0.188 / 1Mvia Alibaba | auto-optimized | |
Qwen: Qwen3.6 Max Preview | $1.04 / 1Mvia Alibaba | auto-optimized | |
Qwen: Qwen3.6 Plus | $0.325 / 1Mvia Alibaba | auto-optimized | |
Qwen: Qwen3.7 Max | $1.25 / 1Mvia Alibaba | auto-optimized | |
Qwen: Qwen3.7 Plus | $0.4 / 1Mvia Alibaba | auto-optimized | |
Qwen2.5 72B Instruct | $0.36 / 1Mvia DeepInfra | auto-optimized | |
Qwen2.5 Coder 32B Instruct | $0.66 / 1Mvia Cloudflare | auto-optimized | |
| Cohere · 4 models | |||
Cohere: Command A | $2.50 / 1Mvia Cohere | auto-optimized | |
Cohere: Command R (08-2024) | $0.15 / 1Mvia Cohere | auto-optimized | |
Cohere: Command R+ (08-2024) | $2.50 / 1Mvia Cohere | auto-optimized | |
Cohere: Command R7B (12-2024) | $0.037 / 1Mvia Cohere | auto-optimized | |
| Amazon · 5 models | |||
Amazon: Nova 2 Lite | $0.3 / 1Mvia Amazon Bedrock | auto-optimized | |
Amazon: Nova Lite 1.0 | $0.06 / 1Mvia Amazon Bedrock | auto-optimized | |
Amazon: Nova Micro 1.0 | $0.035 / 1Mvia Amazon Bedrock | auto-optimized | |
Amazon: Nova Premier 1.0 | $2.50 / 1Mvia Amazon Bedrock | auto-optimized | |
Amazon: Nova Pro 1.0 | $0.8 / 1Mvia Amazon Bedrock | auto-optimized | |
| Microsoft · 3 models | |||
Microsoft: Phi 4 | $0.065 / 1Mvia NextBit | auto-optimized | |
Microsoft: Phi 4 Mini Instruct | $0.08 / 1Mvia WandB | auto-optimized | |
WizardLM-2 8x22B | $0.62 / 1Mvia Novita | auto-optimized | |
| NVIDIA · 5 models | |||
NVIDIA: Llama 3.3 Nemotron Super 49B V1.5 | $0.1 / 1Mvia DeepInfra | auto-optimized | |
NVIDIA: Nemotron 3 Nano 30B A3B | $0.05 / 1Mvia Ambient | auto-optimized | |
NVIDIA: Nemotron 3 Super | $0.09 / 1Mvia DekaLLM | auto-optimized | |
NVIDIA: Nemotron 3 Ultra | $0.5 / 1Mvia DeepInfra | auto-optimized | |
NVIDIA: Nemotron Nano 9B V2 | $0.04 / 1Mvia DeepInfra | auto-optimized | |
| Perplexity · 5 models | |||
Perplexity: Sonar | $1.00 / 1Mvia Perplexity | auto-optimized | |
Perplexity: Sonar Deep Research | $2.00 / 1Mvia Perplexity | auto-optimized | |
Perplexity: Sonar Pro | $3.00 / 1Mvia Perplexity | auto-optimized | |
Perplexity: Sonar Pro Search | $3.00 / 1Mvia Perplexity | auto-optimized | |
Perplexity: Sonar Reasoning Pro | $2.00 / 1Mvia Perplexity | auto-optimized | |
| AI21 · 1 model | |||
AI21: Jamba Large 1.7 | $2.00 / 1Mvia AI21 | auto-optimized | |
| Aion · 4 models | |||
AionLabs: Aion-1.0 | $4.00 / 1Mvia AionLabs | auto-optimized | |
AionLabs: Aion-1.0-Mini | $0.7 / 1Mvia AionLabs | auto-optimized | |
AionLabs: Aion-2.0 | $0.8 / 1Mvia AionLabs | auto-optimized | |
AionLabs: Aion-RP 1.0 (8B) | $0.8 / 1Mvia AionLabs | auto-optimized | |
| Allen AI · 1 model | |||
AllenAI: Olmo 3 32B Think | $0.15 / 1M | auto-optimized | |
| Anthracite · 1 model | |||
Magnum v4 72B | $3.00 / 1Mvia Mancer 2 | auto-optimized | |
| Arcee · 6 models | |||
Arcee AI: Coder Large | $0.5 / 1Mvia Together | auto-optimized | |
Arcee AI: Maestro Reasoning | $0.9 / 1Mvia Together | auto-optimized | |
Arcee AI: Spotlight | $0.18 / 1Mvia Together | auto-optimized | |
Arcee AI: Trinity Large Thinking | $0.22 / 1Mvia Parasail | auto-optimized | |
Arcee AI: Trinity Mini | $0.045 / 1Mvia Clarifai | auto-optimized | |
Arcee AI: Virtuoso Large | $0.75 / 1Mvia Together | auto-optimized | |
| Baidu · 2 models | |||
Baidu: ERNIE 4.5 VL 28B A3B | $0.14 / 1Mvia Novita | auto-optimized | |
Baidu: ERNIE 4.5 VL 424B A47B | $0.42 / 1Mvia Novita | auto-optimized | |
| ByteDance · 1 model | |||
ByteDance: UI-TARS 7B | $0.1 / 1Mvia Parasail | auto-optimized | |
| ByteDance Seed · 4 models | |||
ByteDance Seed: Seed 1.6 | $0.25 / 1Mvia Seed | auto-optimized | |
ByteDance Seed: Seed 1.6 Flash | $0.075 / 1Mvia Seed | auto-optimized | |
ByteDance Seed: Seed-2.0-Lite | $0.25 / 1Mvia Seed | auto-optimized | |
ByteDance Seed: Seed-2.0-Mini | $0.1 / 1Mvia Seed | auto-optimized | |
| DeepCogito · 1 model | |||
Deep Cogito: Cogito v2.1 671B | $1.25 / 1Mvia Together | auto-optimized | |
| Essential AI · 1 model | |||
EssentialAI: Rnj 1 Instruct | $0.15 / 1Mvia Together | auto-optimized | |
| Gryphe · 1 model | |||
MythoMax 13B | $0.06 / 1Mvia NextBit | auto-optimized | |
| IBM Granite · 2 models | |||
IBM: Granite 4.0 Micro | $0.017 / 1Mvia Cloudflare | auto-optimized | |
IBM: Granite 4.1 8B | $0.05 / 1Mvia WandB | auto-optimized | |
| Inception · 1 model | |||
Inception: Mercury 2 | $0.25 / 1Mvia Inception | auto-optimized | |
| Inclusion AI · 3 models | |||
inclusionAI: Ling-2.6-1T | $0.075 / 1Mvia Novita | auto-optimized | |
inclusionAI: Ling-2.6-flash | $0.01 / 1Mvia Novita | auto-optimized | |
inclusionAI: Ring-2.6-1T | $0.075 / 1Mvia Novita | auto-optimized | |
| Inflection · 2 models | |||
Inflection: Inflection 3 Pi | $2.50 / 1Mvia Inflection | auto-optimized | |
Inflection: Inflection 3 Productivity | $2.50 / 1Mvia Inflection | auto-optimized | |
| Kwai · 1 model | |||
Kwaipilot: KAT-Coder-Pro V2 | $0.3 / 1Mvia AtlasCloud | auto-optimized | |
| Liquid · 1 model | |||
LiquidAI: LFM2-24B-A2B | $0.03 / 1Mvia Together | auto-optimized | |
| Mancer · 1 model | |||
Mancer: Weaver (alpha) | $0.75 / 1Mvia Mancer 2 | auto-optimized | |
| MiniMax · 8 models | |||
MiniMax: MiniMax M1 | $0.4 / 1Mvia Minimax | auto-optimized | |
MiniMax: MiniMax M2 | $0.255 / 1Mvia AtlasCloud | auto-optimized | |
MiniMax: MiniMax M2-her | $0.3 / 1Mvia Minimax | auto-optimized | |
MiniMax: MiniMax M2.1 | $0.29 / 1Mvia AtlasCloud | auto-optimized | |
MiniMax: MiniMax M2.5 | $0.15 / 1Mvia AkashML | auto-optimized | |
MiniMax: MiniMax M2.7 | $0.279 / 1Mvia Morph | auto-optimized | |
MiniMax: MiniMax M3 | $0.3 / 1Mvia Minimax | auto-optimized | |
MiniMax: MiniMax-01 | $0.2 / 1Mvia Minimax | auto-optimized | |
| Moonshot · 5 models | |||
MoonshotAI: Kimi K2 0711 | $0.57 / 1Mvia Novita | auto-optimized | |
MoonshotAI: Kimi K2 0905 | $0.6 / 1Mvia AtlasCloud | auto-optimized | |
MoonshotAI: Kimi K2 Thinking | $0.6 / 1Mvia AtlasCloud | auto-optimized | |
MoonshotAI: Kimi K2.5 | $0.4 / 1Mvia ModelRun | auto-optimized | |
MoonshotAI: Kimi K2.6 | $0.684 / 1Mvia Baidu | auto-optimized | |
| Morph · 2 models | |||
Morph: Morph V3 Fast | $0.8 / 1Mvia Morph | auto-optimized | |
Morph: Morph V3 Large | $0.9 / 1Mvia Morph | auto-optimized | |
| Newmen · 1 model | |||
atlas-1 Auto-optimizes each call for the cheapest path that holds quality. | Atlas rate | auto-optimized | |
| Nex AGI · 1 model | |||
Nex AGI: DeepSeek V3.1 Nex N1 | $0.135 / 1Mvia SiliconFlow | auto-optimized | |
| Nous Research · 5 models | |||
Nous: Hermes 3 405B Instruct | $1.00 / 1Mvia DeepInfra | auto-optimized | |
Nous: Hermes 3 70B Instruct | $0.3 / 1Mvia DeepInfra | auto-optimized | |
Nous: Hermes 4 405B | $1.00 / 1Mvia Nebius | auto-optimized | |
Nous: Hermes 4 70B | $0.13 / 1Mvia Nebius | auto-optimized | |
NousResearch: Hermes 2 Pro - Llama-3 8B | $0.14 / 1Mvia Novita | auto-optimized | |
| OpenRouter · 4 models | |||
Auto Router | $-1000000 / 1M | auto-optimized | |
Body Builder (beta) | $-1000000 / 1M | auto-optimized | |
OpenRouter: Fusion | $-1000000 / 1M | auto-optimized | |
Pareto Code Router | $-1000000 / 1M | auto-optimized | |
| Perceptron · 1 model | |||
Perceptron: Perceptron Mk1 | $0.15 / 1Mvia Perceptron | auto-optimized | |
| Prime Intellect · 1 model | |||
Prime Intellect: INTELLECT-3 | $0.2 / 1Mvia Nebius | auto-optimized | |
| Reka · 2 models | |||
Reka Edge | $0.1 / 1Mvia Reka | auto-optimized | |
Reka Flash 3 | $0.1 / 1Mvia Reka | auto-optimized | |
| Relace · 2 models | |||
Relace: Relace Apply 3 | $0.85 / 1Mvia Relace | auto-optimized | |
Relace: Relace Search | $1.00 / 1Mvia Relace | auto-optimized | |
| Sao10k · 5 models | |||
Sao10K: Llama 3 8B Lunaris | $0.04 / 1Mvia DeepInfra | auto-optimized | |
Sao10k: Llama 3 Euryale 70B v2.1 | $1.48 / 1Mvia Novita | auto-optimized | |
Sao10K: Llama 3.1 70B Hanami x1 | $3.00 / 1Mvia Infermatic | auto-optimized | |
Sao10K: Llama 3.1 Euryale 70B v2.2 | $0.85 / 1Mvia DeepInfra | auto-optimized | |
Sao10K: Llama 3.3 Euryale 70B | $0.65 / 1Mvia NextBit | auto-optimized | |
| StepFun · 2 models | |||
StepFun: Step 3.5 Flash | $0.09 / 1Mvia DeepInfra | auto-optimized | |
StepFun: Step 3.7 Flash | $0.2 / 1Mvia StepFun | auto-optimized | |
| Switchpoint · 1 model | |||
Switchpoint Router | $0.85 / 1Mvia Switchpoint | auto-optimized | |
| Tencent · 2 models | |||
Tencent: Hunyuan A13B Instruct | $0.14 / 1Mvia SiliconFlow | auto-optimized | |
Tencent: Hy3 preview | $0.063 / 1Mvia GMICloud | auto-optimized | |
| TheDrummer · 4 models | |||
TheDrummer: Cydonia 24B V4.1 | $0.3 / 1Mvia Parasail | auto-optimized | |
TheDrummer: Rocinante 12B | $0.17 / 1Mvia NextBit | auto-optimized | |
TheDrummer: Skyfall 36B V2 | $0.55 / 1Mvia Parasail | auto-optimized | |
TheDrummer: UnslopNemo 12B | $0.4 / 1Mvia NextBit | auto-optimized | |
| Undi95 · 1 model | |||
ReMM SLERP 13B | $0.45 / 1Mvia NextBit | auto-optimized | |
| Upstage · 1 model | |||
Upstage: Solar Pro 3 | $0.15 / 1Mvia Upstage | auto-optimized | |
| Writer · 1 model | |||
Writer: Palmyra X5 | $0.6 / 1Mvia Amazon Bedrock | auto-optimized | |
| Xiaomi · 3 models | |||
Xiaomi: MiMo-V2-Flash | $0.1 / 1Mvia Xiaomi | auto-optimized | |
Xiaomi: MiMo-V2.5 | $0.14 / 1Mvia Xiaomi | auto-optimized | |
Xiaomi: MiMo-V2.5-Pro | $0.435 / 1Mvia Xiaomi | auto-optimized | |
| Z.AI · 12 models | |||
Z.ai: GLM 4 32B | $0.1 / 1Mvia Z.AI | auto-optimized | |
Z.ai: GLM 4.5 | $0.6 / 1Mvia Novita | auto-optimized | |
Z.ai: GLM 4.5 Air | $0.125 / 1Mvia Io Net | auto-optimized | |
Z.ai: GLM 4.5V | $0.6 / 1Mvia Novita | auto-optimized | |
Z.ai: GLM 4.6 | $0.43 / 1Mvia DeepInfra | auto-optimized | |
Z.ai: GLM 4.6V | $0.3 / 1Mvia Novita | auto-optimized | |
Z.ai: GLM 4.7 | $0.4 / 1Mvia DeepInfra | auto-optimized | |
Z.ai: GLM 4.7 Flash | $0.06 / 1Mvia DeepInfra | auto-optimized | |
Z.ai: GLM 5 | $0.6 / 1Mvia DeepInfra | auto-optimized | |
Z.ai: GLM 5 Turbo | $1.20 / 1Mvia AtlasCloud | auto-optimized | |
Z.ai: GLM 5.1 | $0.98 / 1Mvia Baidu | auto-optimized | |
Z.ai: GLM 5V Turbo | $1.20 / 1Mvia Z.AI | auto-optimized | |
Prices are input tokens per 1M. Strict is the cheapest source for pinning the exact model — same model, no substitutions. Atlas mode is the default: it auto-optimizes each call for the cheapest path that holds quality, at least 5% off from your first call and climbing as it ramps. You always see which model served each call.
On Pay as you go, what you see is the ceiling — you pay that or less, and at least 5% under direct from call one. Reliability Loop adds the platform fee for the per-operation tuning, eval-gated refund, and opt-in per-tenant tuning.
Atlas mode is the default behavior, not a separate model. It looks at the operation you tagged, the eval-gate history for that operation, and the options available for the underlying model — then serves the cheapest path that has historically held quality.
- model: "gpt-5.5",
+ model: "atlas-1", // auto-optimize: cheapest path that holds quality
+ tier: "standard", // optional; default is "standard"
metadata: { operation_id: "summarize_ticket" },What you keep
OpenAI-compatible API, your existing SDK, every supported provider model. Pin a specific model anytime.
What changes
Atlas mode serves the cheapest path that holds quality per call. The response carries a `delivery` block telling you exactly which model served it.
What protects quality
If `metadata.operation_id` resolves to an operation with a bound evaluator and a min_score, calls scoring below threshold aren't billed.
Tier is a per-call hint. Atlas mode honours it; pinned models honour it within the variants the upstream provides.
realtime
p95 < 300ms TTFT
At direct rate
Direct passthrough to the upstream provider, full precision. The only tier with a hard delivery guarantee on closed-weight models. Use when latency is non-negotiable.
standard
p95 < 8s TTFT
30–55% off
Atlas serves each call the cheapest way that holds quality on your operation's eval gates. Default tier — your bill drops without changing business logic. Best-effort: may upgrade to Realtime when no path passes; bill follows actual delivery.
batch
~24h SLA
50–70% off
Async. Routes to provider batch APIs (OpenAI / Anthropic) where supported, queued spot capacity for open-weight models. Largest discounts. Streaming not supported.
Discounts are computed against the realtime cost-leader for the same logical model. Per-task numbers are published openly on the benchmarks page; per-call savings vary with prompt mix and time of day.
A customer running 30M tokens / month of support-ticket summarization on gpt-5.5 (input $2.75 / output $11 per 1M) pays roughly $210 / month.
They flip model: "atlas-1" and bind a regex evaluator with min_score: 0.95 to the summarize_ticket operation. After two weeks of green eval gates, Atlas serves most calls the cheaper way and keeps the rest on the original model where the loop says it’s still needed.
New monthly cost: ~$84. That’s a 60% reduction. Calls scoring below the regex threshold aren’t metered, so the customer absorbs zero quality risk.
Numbers are illustrative. Run the comparison runner on your own prompts for a real per-workload estimate.
The eval-refund is the only thing standing between Atlas's cost claim and the long history of inference brokers promising the quantized version is fine. Bind a numeric threshold; calls below it never show up on your invoice. On every plan.
01 · Bind
Bind an evaluator with a min_score to your operation.
Regex, LLM-judge, structural — any evaluator that returns a numeric score. Set once per operation in /console/evaluators.
02 · Score
Atlas runs the evaluator on the served output, on the call path.
Same after() hook that meters usage. Score appears on the call record within about a second.
03 · Refund
Score below min_score → call's net metered quantity is zero.
A compensating Stripe meter event fires automatically. The call still appears in /console/calls so your team can correct it via feedback — but it never appears on your invoice.
Works on PAYG. Works on Reliability Loop. Works on Strategic. The mechanic is platform-wide — the only requirement is a bound evaluator with a numeric threshold and metadata.operation_id on the call.
Usage rates are the same across all plans — the per-call tier (Realtime / Standard / Batch) decides the bill. Plan choice gates the reliability loop, the quality refund, and Atlas Network defaults.
$0
$5 free credits · free models, no card
Always cheaper than going direct — at least 5% off from your first call, climbing as Atlas ramps on your traffic. Start free on free models with no card; add one for the full catalogue. Thumbs-down any call you don't like and we refund it in full.
$1,500 / mo
platform fee
Everything in Pay as you go plus the productized reliability workflow — operation-aware optimization, eval-gated auto-refund, evaluators and ship gates as first-class concepts, dataset promotion, and opt-in per-tenant tuning on your own data. The savings on the optimization side usually cover the platform fee.
Custom
annual commitment
For teams running significant volume. Negotiated rates per quantization tier, dedicated capacity, and a solutions engineer who knows your evals.
On Pay as you go, thumbs-down any call and we credit it back automatically (fair-use limits in the terms). On the Reliability Loop tier this stacks with the eval-gated auto-refund: bind an evaluator with a numeric min_score to any operation, and calls below threshold are refunded synchronously without anyone hitting thumbs-down. Strategic customers negotiate per-tier rates and dedicated capacity.
Atlas is sold to teams who commit to meaningful production volume. That commitment unlocks the reliability loop.