Thank you NVIDIA and AMD. DGX Spark gifted by NVIDIA, Strix Halo gifted by AMD, and now a Strix Halo desktop from AMD too. That desktop is the box we ran our MLPerf submission on. See the post ↗
Open source. Pure Rust and CUDA. Verified on GB10.

The inference engine for the machine on your desk.

Atlas is an open source LLM engine hand tuned for NVIDIA DGX Spark and AMD Strix Halo. One ~75 MB binary, no Python, no PyTorch. Tuned for Qwen3.8 and Nemotron, and verified every single release.

$ curl -fsSL https://atlasinference.dev/quickstart.sh | sh

Do not take our word for it. First token in under 90 seconds on a DGX Spark. Validated in MLPerf Inference v6.1: 20.1 tok/s across 1,007 agentic turns on DGX Spark, 19.63 tok/s on AMD Strix Halo.

// 01 · news

Peer-reviewed receipts.

MLPerf Inference v6.1 results, independent coverage in StorageReview, and working group statements—every card links straight to the primary source.

MLCommons September 2026

MLCommons Chairs Report & Atlas Supplemental Statement

MLPerf Inference Working Group chairs Miro Hodak (AMD) and Frank Han (Dell) broke down v6.1’s milestone shift to agentic evaluation. In the official MLCommons Supplemental Statement, Atlas declared: "For the past few years, serious agentic work meant a datacenter round trip. That assumption is what this submission is meant to retire. A developer choosing an agentic stack should not be choosing an accelerator vendor for the next three years."

Read the Chairs analysis & statement →
Edge Agentic September 2026

MLPerf v6.1 Edge Results Published

Our MLPerf Inference v6.1 numbers are officially published in the closed edge division across NVIDIA DGX Spark (GB10) and AMD Strix Halo (gfx1151). Compiling identical pure Rust and CUDA source across both hardware architectures with SCALE, Atlas completed all 1,007 SWE-bench Verified coding turns with deterministic accuracy, setting the benchmark for local agentic AI.

Download the MLCommons supplemental PDF →
// 02 · momentum

Built in the open, starred in the open.

Atlas went from one Reddit post to a whole crew of builders running it on their own Sparks. The curve below is live, regenerated from the GitHub API on every deploy.

675★
GitHub stars and climbing, live from the API.
0350700675 ★MayJunAug
“The Ryzen AI Max+ 395 appears through Atlas Inference, a first-time submitter that ran the new Edge Agentic workload on both an NVIDIA DGX Spark and an AMD Strix Halo desktop with the same engine and quantization recipe... That’s a narrower gap than we measured with off-the-shelf runtimes.”
StorageReview, MLPerf v6.1 Review
“For the past few years, serious agentic work meant a datacenter round trip. That assumption is what this submission is meant to retire. Local model capabilities and accessible compute have reached a level that everyday users can experience for themselves without the cloud.”
Atlas Inference, MLCommons v6.1 Statement
“Testing Atlas on a DGX Spark in an agentic workflow for over an hour. Super impressed. Spark is actually awesome with Atlas.”
PersonWhoThinks, r/LocalLLaMA
“Night and day compared to the 10 minute torch.compile cycle. Startup in about 15 seconds and it just stays coherent in an agentic loop.”
ronald_15496, #general
// come build with us

The action is in Discord.

Hundreds of builders are running Atlas on their own Sparks right now. We are in there every single day, shipping fixes, taking model requests, and tuning kernels live. Your machine is the test fleet and your voice sets the roadmap. Pull up.

Join the Discord Active every day. Bring your Spark.
// 03 · verified

Every number is a receipt.

The website is a build artifact of the repo. Models come from recipes, performance comes from committed gate enforced baselines, stamped with commit and date. If a number is not in the repo, it is not on this page.

What the gate checks

An Atlas image ships only after the serve matrix passes: every model boots, stays coherent (greedy determinism, no token leakage, tool reliability), and holds throughput within 10% of its committed baseline. What “verified” means · gate_results.py

A release that ships slower than the committed baseline fails our gate. That one sentence is the whole positioning.

Published in MLPerf Inference v6.1: 20.1 tok/s on NVIDIA DGX Spark (GB10) and 19.63 tok/s on AMD Strix Halo (gfx1151) across 1,007 agentic coding turns. Featured in StorageReview and the MLCommons Chairs Report.

Atlas is a member of MLCommons and co-contributor to the Edge Agentic taskforce. In the official v6.1 Supplemental Statement and Working Group analysis, Atlas demonstrated cross-architecture edge parity without cloud dependency. read the MLCommons report.

The MLPerf name and logo are registered and unregistered trademarks of MLCommons Association in the United States and other countries. All rights reserved. Unauthorized use strictly prohibited. See mlcommons.org for more information.

python3 tests/run_all_models.py && python3 tests/gate_results.py --update-baselines

Beat these numbers or catch a regression, open an issue and we will feature it.

Every model card comes from a recipe in atlas-recipes.

serve matrix DGX Spark · GB10
✓ MLPerf Inference v6.1 Published

Atlas Inference v6.1 results are officially published by MLCommons in the closed edge division across both NVIDIA DGX Spark (20.1 tok/s) and AMD Strix Halo (19.63 tok/s), completing all 1,007 agentic turns with identical pure Rust and CUDA source code.

✓DGX Spark · GB1020.10 tok/s
✓Strix Halo · gfx115119.63 tok/s
✓1,007 agentic turns<64 min
atlas ee15431af 2026-09-10
// 04 · hardware

Prosumer first. Desk machines, not clusters.

NVIDIA DGX Spark🎁

GB10 · SM121
MLPerf v6.1: 20.1 tok/s

One multi model binary serves a full matrix of hand tuned targets on a single GB10. NVFP4 and FP8, Qwen3.8-27B, Qwen 3.8 Flash, Nemotron 3.5 Lightning with DSpark, EP=2 across two Sparks. Validated in MLPerf v6.1 completing all 1,007 agentic turns in under 64 minutes with a 40-second cold-to-serve time.

AMD Strix Halo🎁

gfx1151 · RDNA 3.5
MLPerf v6.1: 19.63 tok/s

One codebase, both camps. Our CUDA kernels compile straight for AMD gfx1151 with SCALE by Spectral Compute. As reported by StorageReview, Atlas delivered 19.63 tokens/s across the 1,007-turn SWE-bench Verified agentic workload, narrowing the gap with DGX Spark.

// 05 · models

Every model here has a recipe.

Pick a vendor, then a family. Every card maps to one recipe in atlas-recipes, so the site cannot list a model we do not ship. Copy the command and run it as is. Qwen3.8 leads because it is our flagship.

Every recipe is the single source of truth in atlas-recipes, so the site cannot list a model we do not ship. EP=2 is Expert Parallelism across two GB10 nodes.
Upstream Hugging Face Integration huggingface/transformers #46423 ↗

Fused Qwen3.6 Gated DeltaNet kernel merged directly into Hugging Face Transformers

Our fused Qwen3.6 Gated DeltaNet kernel ships in Hugging Face Transformers. transformers #46423 · kernel repo on the Hub. We are Qwen Dev Ambassadors and we ship a recipe for every Qwen release. Qwen ambassadors ↗

0.999970 Cosine parity vs FP32 torch fallback across 27B & 35B-A3B
100% Bit-Identical Full-model greedy generate token IDs matched on GB10
SM121 Native Auto-dispatched on DGX Spark via @use_kernel_forward_from_hub
Upstream Main Merged into transformers core, no custom torch fork required
Decorates Qwen3_5GatedDeltaNet and Qwen3_5MoeGatedDeltaNet with compute-capability-gated dispatch in integrations/hub_kernels.py. Standard pip install transformers picks up Atlas's hand-tuned fused kernel automatically on NVIDIA DGX Spark.
// 06 · run

One binary. Up in ninety seconds.

bash
$ curl -fsSL https://atlasinference.dev/quickstart.sh | sh
# checks for sparkrun, installs via uvx if missing, runs the flagship recipe

The script checks for sparkrun, installs it with uvx if missing, then runs the flagship Qwen3.8-27B recipe.

Prefer to inspect first?

Rather not pipe curl to a shell. Install sparkrun, then run the flagship recipe direct.

$ uvx sparkrun setup install
$ sparkrun run @atlas/qwen3.8-27b-nvfp4 --hosts localhost

The first 60 seconds live here. Everything after, per model recipes, EP=2, tuning, lives in the docs. Read the deployment guide · README

// 07 · mission

Why pure Rust.

The python AI stack was built for research. It is dynamic, flexible, and twenty gigabytes of dependencies that break when a minor version bumps. That is fine for a notebook, but it is the wrong foundation for an engine that lives on your desk and runs your day.

Atlas is written in pure Rust and CUDA. Every kernel is compiled straight to machine code, every memory allocation is accounted for, and the whole engine is a single binary under eighty megabytes. When you boot it, it is ready in seconds, not minutes.

We believe the future of AI is personal, private, and local. You should own the hardware, you should own the weights, and the software that runs them should be as reliable as your operating system. That is what we are building.

// 08 · contribute

Built in the open. PRs welcome.

Atlas is open source under the AGPLv3. Everything we build is public from day one.

Pick up a job 3 open

Small, concrete pieces of work. A new model, or a feature for one we already run. Grab one, or post one you want to see.

from the last deploy
Open Feature

DFlash2: batched concurrency on main

nvidia/Qwen3.8-27B-NVFP4 + incoai/Qwen3.8-27B-DFlash2

Finish DFlash2 concurrency on main. #52 made batched verify/propose work: aggregate went from 24 to 42 tok/s at C16 on a GB10 (γ=8, Option B). Its known follow-ups are what's left: a gdn_decode_wy8 kernel, so the batched WY fast path covers k=8 (today it falls back per-sequence, correctly) owner-in-key CUDA-graph replay for batched verify (currently vetoed, so it runs eager) report the multi-lane propose validation (ATLAS_DFLASH_PROPOSE_LANES=4) Done looks like: a C1/C4/C8/C16 sweep on GB10 that beats #52's table, with ST-995 unchanged.

Shared code, other models could change · the scheduler's batched verify path, GDN decode kernels, and the CUDA graph cache key
Testing @AzeezIsh (GB10)
Open Feature

Qwen3.8-Flash-Next: land the Strix Halo Windows port

nvidia/Qwen3.8-Flash-Next-NVFP4

Land the Strix Halo Windows port of Qwen3.8-Flash-Next (native HIP, gfx1151). The work is in #41 and its refresh #54, stacked on #35 (the nvidia-pack loader). Done looks like: #35 and then #41/#54 merged to main, the agentic-webserver leg on #54 reported green, and win_serve_flashnext_nvfp4.ps1 booting from a clean checkout. Note: first_run.ps1 defaults ATLAS_TARGET_MODEL=qwen3.8-27b, so it has to be overridden to qwen3.8-flash-next.

Shared code, other models could change · the cuBLASLt BF16 fallback (installed at backend init for every model), strix-hip/common/rms_norm.cu becoming a symlink to gb10's
Testing @AzeezIsh (Strix Halo, Windows 11)
Open Feature

Qwen3.8-Flash-Next: run on Strix Halo Linux

nvidia/Qwen3.8-Flash-Next-NVFP4

Serve Qwen3.8-Flash-Next on Strix Halo Linux (gfx1151, ROCm). The Windows port in #41 / #54 brings up kernels/strix-hip/qwen3.8-flash-next/; nothing has run it on Linux yet. Done looks like: builds with ATLAS_TARGET_MODEL=qwen3.8-flash-next ./build-amd.sh, boots with a clean kernel audit (no unresolved-kernel flag), and posts an ST-995 run plus one agentic leg on a 128 GB Strix Halo. Heads-up: the weights are 92.6 GB, so --gpu-memory-utilization 0.90 is the floor. 0.80 will not boot.

Shared code, other models could change · strix-hip/common kernels and the cuBLASLt BF16 boundary fallback in spark-runtime that #41 adds
Testing @AzeezIsh (Strix Halo)

// 09 · roadmap

What we are building next.

Everything real links to an issue, a PR, or the Discord where the work happens. The teasers are teasers, and we say so.

🎁Nemotron-3.5-Lightning + DSpark

PR recipe live

Verified on GB10 silicon. Double-buffered Mamba-2 chunked scans paired with the 6-layer DSpark speculative drafter.

Recipe in sparkrun →

Trifecta, three Sparks

Next up

Scale beyond two nodes. EP=3 across three DGX Sparks with ring all-to-all.

Track in Discord →

FP4 KV cache

In development

Cut KV cache memory in half again with native FP4 quantization, doubling effective context length on GB10.

Follow PR →

Community MoE ports

Tracking

Large MoE NVFP4 ports across EP topologies, DeepSeek and Kimi class, tracked in the open.

Open issues →
// 10 · contact

Talk to the team.

Whether you are deploying Atlas in production, benchmarking on your own hardware, or just want to chat about local inference, we would love to hear from you.