Meta: Llama 3.3 70B Instruct
meta-llama/llama-3.3-70b-instruct
Cheapest provider
$0.10 / 1M
DeepInfra
Fastest provider (p95)
435 tok/s
Groq
Intelligence (composite)
58.4
MMLU-Pro · HumanEval · math · GPQA
Per-provider performance
Latency / throughput / uptime / price measured across providers over the last 30 minutes of live traffic. This is what proves “sourced cheapest” — Atlas mode draws on these per call to serve the cheapest path that holds quality.
| Provider | Quant | Input $/1M | Output $/1M | Latency p50 / p95 | Throughput p50 / p95 | Uptime 30m | Success |
|---|---|---|---|---|---|---|---|
| DeepInfra | q8· fp8 | $0.1000 | $0.3200 | 436ms / 2748ms | 16 / 30 tok/s | 98.72% | 98.7% |
| Inceptron | q8· fp8 | $0.1200 | $0.3800 | 543ms / 3056ms | 28 / 43 tok/s | 98.51% | 96.9% |
| AkashML | q8· fp8 | $0.1300 | $0.4000 | 463ms / 1439ms | 36 / 69 tok/s | 99.49% | 99.5% |
| Nebius | q8· fp8 | $0.1300 | $0.4000 | 367ms / 1737ms | 27 / 42 tok/s | 90.74% | 90.7% |
| Novita | full· bf16 | $0.1350 | $0.4000 | 613ms / 1160ms | 32 / 50 tok/s | 99.71% | 90.0% |
| Parasail | q8· int8 | $0.2200 | $0.5000 | 679ms / 2107ms | 27 / 39 tok/s | 97.47% | 97.1% |
| Cloudflare | q8· fp8 | $0.2930 | $2.2530 | 513ms / 1723ms | 28 / 43 tok/s | 98.87% | 98.9% |
| SambaNova | full· bf16 | $0.4500 | $0.9000 | 614ms / 2904ms | 60 / 260.4 tok/s | 96.70% | 98.4% |
| Groq | undisclosed | $0.5900 | $0.7900 | 242ms / 666ms | 190 / 434.8 tok/s | 99.69% | 98.7% |
| SambaNova | full· bf16 | $0.6000 | $1.2000 | 614ms / 2904ms | 60 / 260.4 tok/s | 98.41% | 98.4% |
| WandB | full· fp16 | $0.7100 | $0.7100 | 314ms / 794ms | 47 / 89 tok/s | 100.00% | 100.0% |
| undisclosed | $0.7200 | $0.7200 | 334ms / 1078ms | 22 / 99.1 tok/s | 90.04% | 90.0% | |
| undisclosed | $0.7200 | $0.7200 | 334ms / 1078ms | 22 / 99.1 tok/s | 100.00% | 90.0% | |
| Together | q8· fp8 | $1.0400 | $1.0400 | 1125ms / 5929ms | 22 / 81.8 tok/s | 99.30% | 99.3% |
“—” means live telemetry hasn’t accumulated enough recent traffic for that endpoint. “undisclosed” means the provider serves the model but doesn’t expose the quantization label (typically running fp8 / int8 internally).
Intelligence breakdown
Composite score is a weighted average of public benchmarks (30% MMLU-Pro, 25% code pass@1, 25% math, 20% GPQA). Numbers come from model cards and the Artificial Analysis intelligence harness; missing components are renormalised over what’s present.
MMLU-Pro
68.9
broad reasoning
Code
33.3
pass@1 (HumanEval / LiveCodeBench)
MATH
77.0
math accuracy
GPQA Diamond
50.5
hard reasoning
Source: Meta Llama 3.x cards — Llama 3.3 70B (MMLU-Pro 68.9, GPQA-Diamond 50.5, LiveCodeBench 33.3); llama.com
How Atlas mode sources Meta: Llama 3.3 70B Instruct
- Strict mode — pin Meta: Llama 3.3 70B Instruct exactly and we pass it straight through, sourced from the cheapest provider above. The same model, no substitutions — currently DeepInfra at $0.10/1M.
- Atlas mode — the default. Each call is auto-optimized for the cheapest path that holds quality, at least 5% off going direct from call one and climbing as it ramps. You always see which model served the call and exactly what you saved — thumbs-down anything you don’t like for a full refund.