How it works
Two env vars. Then it just gets cheaper.
Atlas is the optimization layer between your tools and the models. Point your coding agent at it, keep the models you use, and pay less than going direct from your first call. You see the model that served every call — and if it isn’t as good, it’s free.
Start here
Change two env vars.
The whole migration. Keep your client, your SDK, and your model ids — OpenAI- and Anthropic-compatible. Live on day one.
export ANTHROPIC_BASE_URL="https://api.newmen.ai"
export ANTHROPIC_API_KEY="nm_live_..."Claude Code calls https://api.newmen.ai/v1/messages
Prefer the full walkthrough? See the migration guide.
What happens over time
Quality holds. Cost and latency fall.
As Atlas ramps on your traffic, it learns which paths hold quality on your workloads and serves more calls the cheaper, faster way. Quality is verified per call — anything that isn't as good is free.
Illustrative. Your real numbers are measured on your own traffic.
Two modes
Optimize by default. Pin when it matters.
One model field, two behaviors. You're never locked out of control.
Atlas mode (default)
Pass model: "atlas-1" (or just let it default) and Atlas serves each call the cheapest way that holds quality — at least 5% under direct from call one, climbing as it ramps. You always see which model served the call.
Strict mode
Pin an exact model — model: "openai/gpt-5.5" — and Atlas delivers that exact model, sourced cheaper. Verbatim prompt, full precision, no substitutions, for workloads where the model must not change.
Go deeper in the methodology, or see the numbers in the proof.
Talk to a solutions engineer
Atlas is sold to teams who commit to meaningful production volume. That commitment unlocks the reliability loop.