Repository navigation
fix(local-llm): CUDA build that ships its runtime on Windows, and Windows on ARM engine (ATO-244, ATO-252) - #613
Merged
Merged
Conversation
… (ATO-244) The cuda-13.3 Windows zip of turboquant-6df272c ships ggml-cuda.dll without cudart64_13 / cublas64_13 / cublasLt64_13. On a machine without the CUDA Toolkit the CUDA backend fails to load silently, llama-server lists no devices and the model runs on the CPU. Detection sent every driver reporting CUDA >= 13.0 (r580+) there. Selection: a driver with CUDA >= 12.4 now gets the cuda-12.4 build, which bundles its runtime and has native code for Ampere, Ada and Hopper. Blackwell (compute capability >= 10, RTX 50-series), which the 12.4 toolkit has no code for, gets Vulkan. The capability comes from `nvidia-smi --query-gpu=compute_cap`; a failed query counts as pre-Blackwell. cuda-13.3 stays available as a pin. Guard: a Windows CUDA install with ggml-cuda.dll but no cudart64_*.dll is stale for the variant check (auto-detection only, pins are left alone), and the installer refuses such a zip, installing the release's Vulkan build and recording the refusal so the next start does not pull the same broken zip again. Existing cuda-13.3 installs are replaced on the next managed start through the existing auto-update variant check.
…when none is published (ATO-252)
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
On Windows the agent picked
llama-turboquant-windows-x64-cuda-13.3.zipfor any NVIDIA driver with CUDA 13.x. In releaseturboquant-6df272cthat zip came without the CUDA runtime (cudart64_13.dll,cublas64_13.dll,cublasLt64_13.dll). So on a machine without the CUDA Toolkit,ggml-cuda.dllcould not load and llama-server silently ran on the CPU. On an RTX 3080 Ti with driver 591.44, prompt processing ran at about 16 tok/s and the first turn took about 7 minutes.The engine packaging was fixed separately (AtomicBot-ai/atomic-llama-cpp-turboquant-nightly#4, release
turboquant-ad5ad5f). This PR makes the agent safe even when a zip arrives incomplete.What changes
cuda-12.4(newer drivers run older runtimes). On Blackwell (compute capability ≥ 10, read withnvidia-smi --query-gpu=compute_cap) it picks Vulkan, because the 12.4 build has no code for it. Below 12.4, or with no NVIDIA driver, it picks Vulkan, as before.cuda-13.3is no longer auto-picked but can still be pinned withlocalModels.managed.backendVariant.auto, an installed CUDA build withoutcudart64_*.dllnext toggml-cuda.dllcounts as stale and is replaced through the existing auto-update at managed start. Existing broken installs are fixed without manual steps.downloadBackendrefuses it and installs the same release's Vulkan zip. It recordsrefusedCudaAssetinbackend-version.jsonand tries the CUDA zip again only on a newer release.backendVariantcomment are updated.The CPU zip is still never picked by detection;
cpu-backend-fallback.tsonly learns to never fall back on win32-arm64.Windows on ARM (ATO-252)
resolvePlatformAssetno longer throws on win32-arm64. It resolves tollama-turboquant-windows-arm64-cpu.zip(engine releaseturboquant-9ca222f, CPU only);backendVariantand the NVIDIA detection are ignored there, and there is no CPU fallback loop.downloadBackendandmodels updatesay plainly: "Local models are not available yet for Windows on ARM. Use a cloud model instead."models updateexits 1 instead of reporting "backend unchanged".Checks