Skip to content

fix(local-llm): CUDA build that ships its runtime on Windows, and Windows on ARM engine (ATO-244, ATO-252) - #613

Merged
sosidudku1 merged 2 commits into
mainfrom
fix/local-gpu-not-cpu-ato244
Oct 6, 2026
Merged

sosidudku1 merged 2 commits into
mainfrom
fix/local-gpu-not-cpu-ato244

Conversation

@sosidudku1

@sosidudku1 sosidudku1 commented Oct 6, 2026 •

Copy link
Copy Markdown
Collaborator

On Windows the agent picked llama-turboquant-windows-x64-cuda-13.3.zip for any NVIDIA driver with CUDA 13.x. In release turboquant-6df272c that zip came without the CUDA runtime (cudart64_13.dll, cublas64_13.dll, cublasLt64_13.dll). So on a machine without the CUDA Toolkit, ggml-cuda.dll could not load and llama-server silently ran on the CPU. On an RTX 3080 Ti with driver 591.44, prompt processing ran at about 16 tok/s and the first turn took about 7 minutes.

The engine packaging was fixed separately (AtomicBot-ai/atomic-llama-cpp-turboquant-nightly#4, release turboquant-ad5ad5f). This PR makes the agent safe even when a zip arrives incomplete.

What changes

  • Selection. With driver CUDA ≥ 12.4 the agent picks cuda-12.4 (newer drivers run older runtimes). On Blackwell (compute capability ≥ 10, read with nvidia-smi --query-gpu=compute_cap) it picks Vulkan, because the 12.4 build has no code for it. Below 12.4, or with no NVIDIA driver, it picks Vulkan, as before. cuda-13.3 is no longer auto-picked but can still be pinned with localModels.managed.backendVariant.
  • Completeness guard. Under auto, an installed CUDA build without cudart64_*.dll next to ggml-cuda.dll counts as stale and is replaced through the existing auto-update at managed start. Existing broken installs are fixed without manual steps.
  • No re-download loop. If a CUDA zip itself arrives without its runtime, downloadBackend refuses it and installs the same release's Vulkan zip. It records refusedCudaAsset in backend-version.json and tries the CUDA zip again only on a newer release.
  • Tests cover the selection table, the stale checks, and the refuse-and-fall-back path. README and the backendVariant comment are updated.

The CPU zip is still never picked by detection; cpu-backend-fallback.ts only learns to never fall back on win32-arm64.

Windows on ARM (ATO-252)

  • resolvePlatformAsset no longer throws on win32-arm64. It resolves to llama-turboquant-windows-arm64-cpu.zip (engine release turboquant-9ca222f, CPU only); backendVariant and the NVIDIA detection are ignored there, and there is no CPU fallback loop.
  • When no release carries that asset, downloadBackend and models update say plainly: "Local models are not available yet for Windows on ARM. Use a cloud model instead." models update exits 1 instead of reporting "backend unchanged".

Checks

  • On the MacBook Pro: agent lint and build clean, the whole vitest suite 12 925 passed with 0 failed (with the ARM part).
  • The same change runs in the desktop batch (Desktop 06.10 batch with Windows on ARM (ATO-252) #614): full desktop smoke and e2e (8/8) on macOS, plus the smoke suite on Windows x64 and Windows on ARM runners (738 passed on each, no unexpected failures).
  • On the reporter's PC (RTX 3080 Ti) the test build loads the model on the GPU.

… (ATO-244)

The cuda-13.3 Windows zip of turboquant-6df272c ships ggml-cuda.dll
without cudart64_13 / cublas64_13 / cublasLt64_13. On a machine without
the CUDA Toolkit the CUDA backend fails to load silently, llama-server
lists no devices and the model runs on the CPU. Detection sent every
driver reporting CUDA >= 13.0 (r580+) there.

Selection: a driver with CUDA >= 12.4 now gets the cuda-12.4 build,
which bundles its runtime and has native code for Ampere, Ada and
Hopper. Blackwell (compute capability >= 10, RTX 50-series), which the
12.4 toolkit has no code for, gets Vulkan. The capability comes from
`nvidia-smi --query-gpu=compute_cap`; a failed query counts as
pre-Blackwell. cuda-13.3 stays available as a pin.

Guard: a Windows CUDA install with ggml-cuda.dll but no cudart64_*.dll
is stale for the variant check (auto-detection only, pins are left
alone), and the installer refuses such a zip, installing the release's
Vulkan build and recording the refusal so the next start does not pull
the same broken zip again. Existing cuda-13.3 installs are replaced on
the next managed start through the existing auto-update variant check.
@sosidudku1 sosidudku1 changed the title fix(local-llm): pick the CUDA build that ships its runtime on Windows (ATO-244) fix(local-llm): CUDA build that ships its runtime on Windows, and Windows on ARM engine (ATO-244, ATO-252) Oct 6, 2026
@sosidudku1
sosidudku1 merged commit a8afe96 into main Oct 6, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant