Skip to content

Expose llama_model_params.devices as IModelParams.Devices - #1447

Merged
martindevans merged 3 commits into
SciSharp:masterfrom
ipevo-brooks:model-params-devices
Oct 5, 2026
Merged

martindevans merged 3 commits into
SciSharp:masterfrom
ipevo-brooks:model-params-devices

Conversation

@ipevo-brooks

Copy link
Copy Markdown
Contributor

Problem

llama_model_params.devices has been declared in LLamaModelParams for a while but is private with a todo comment, so there is no way to set it from managed code. The only knob exposed is MainGpu.

That matters because of how llama.cpp picks devices when devices is NULL (llama_prepare_model_devices in src/llama.cpp):

  • all discrete GPUs (GGML_BACKEND_DEVICE_TYPE_GPU) are candidates;
  • integrated GPUs (GGML_BACKEND_DEVICE_TYPE_IGPU) are only added when no discrete GPU exists;
  • main_gpu indexes into that candidate list.

So on a laptop with e.g. an NVIDIA dGPU plus an Intel Arc iGPU, the iGPU can never be selected through LLamaSharp, whatever MainGpu is set to. TensorBufferOverrides cannot work around it either: the KV cache still follows the candidate list, so weights and KV end up on different devices and the load crashes.

llama.cpp's own escape hatch for this is devices (--device Vulkan1 on the CLI). This PR wires it through.

Changes

  • IModelParams.Devices / ModelParams.Devices (List<string>, default empty): device names as returned by ggml_backend_dev_name, e.g. "Vulkan1", "CUDA0", "CPU". Empty keeps the current behaviour (llama.cpp chooses).
  • IModelParamsExtensions.ToLlamaModelParams: resolves the names against ggml_backend_dev_count/get/name, builds a NULL-terminated ggml_backend_dev_t[], pins it in the existing GroupDisposable, and assigns devices. Unknown names are ignored; if nothing matches, devices stays NULL. Same pattern as ConvertOverrides.
  • LLamaModelParams.devices becomes public IntPtr* (same size and offset as before, so the native layout is unchanged).
  • NativeApi.ggml_backend_dev_name P/Invoke added (ggml-base).
  • LLama.Web ModelOptions implements the new member.
  • ModelsParamsTests.SerializeRoundTripSystemTextJson covers the new property.

Breaking change

Adding a member to IModelParams breaks external implementers of that interface (default interface members are not available on netstandard2.0). ModelParams and LLama.Web's ModelOptions are updated here; third-party implementations need a one-line List<string> Devices { get; } = new();.

Testing

  • LLama, LLama.Web and LLama.Unittest build; ModelsParamsTests passes.
  • The same change has been running since 2026-05 in a vendored copy of LLamaSharp (0.27.0, then 0.29.0) on Windows with the Vulkan backend, selecting an Intel Arc 140T iGPU next to an RTX 4050 via Devices = { "Vulkan1" }, and selecting the iGPU / CPU explicitly on an Intel N100.

LLamaModelParams.devices was private with a todo, so the only way to pick a
device from managed code was main_gpu. llama.cpp's default device selection
(llama_prepare_model_devices) only considers integrated GPUs when no discrete
GPU is present, and main_gpu indexes into that filtered list, so an iGPU next
to a dGPU could never be selected. llama.cpp's escape hatch for this is the
devices list (--device on the CLI).

Add IModelParams.Devices (device names as returned by ggml_backend_dev_name,
e.g. "Vulkan1"). ToLlamaModelParams resolves the names against the available
ggml backend devices, builds a NULL-terminated array pinned in the existing
GroupDisposable and assigns it to devices. Empty list or no matching names
keeps devices NULL, i.e. the current behaviour.

Also adds the ggml_backend_dev_name P/Invoke, implements the member in
LLama.Web's ModelOptions, and covers the property in the ModelParams JSON
round-trip test.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟢 Approval recommended

The only finding is a non-blocking documentation nit.

Review effort: Lite
Findings: None

What changed in this PR

Adds managed llama.cpp device selection via IModelParams.Devices, including native device resolution and serialization support.

Changes:

  • Exposes configurable device names through model parameters and web options.
  • Resolves and pins native device handles during model loading.
  • Adds interop and JSON round-trip test coverage.
File Summary
LLama/​Native/​NativeApi.cs Adds device-name interop.
LLama/​Native/​LLamaModelParams.cs Exposes the native devices pointer.
LLama/​Extensions/​IModelParamsExtensions.cs Resolves and pins configured devices.
LLama/​Common/​ModelParams.cs Adds managed device configuration.
LLama/​Abstractions/​IModelParams.cs Defines the new API member; documentation nit noted for integrated-device fallback wording.
LLama.Web/​Common/​ModelOptions.cs Supports device configuration.
LLama.Unittest/​ModelsParamsTests.cs Tests device-list serialization.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

@martindevans martindevans left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Just one minor issue with error handling, otherwise this looks pretty good. Thanks for working on it.

Comment thread LLama/Extensions/IModelParamsExtensions.cs Outdated
Review feedback: a device name that does not match any available device
was silently dropped. ConvertDevices now resolves each name with an exact
match first, then a case insensitive match, and throws
UnknownDeviceException (new, in LLama/Exceptions) if neither matches. The
exception carries the requested name and the list of available device
names, and both appear in the message.

Adds unit tests for the unknown-name and case-insensitive cases.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@ipevo-brooks

Copy link
Copy Markdown
Contributor Author

Thanks for the review. Pushed d704751 addressing both points:

  • ConvertDevices now resolves each name with an exact match first, then a case insensitive match (OrdinalIgnoreCase, equivalent to the ToUpperInvariant suggestion), and throws if neither matches. The "no match -> leave devices NULL" fallback is gone.
  • New UnknownDeviceException in LLama/Exceptions, carrying RequestedDevice and AvailableDevices; both are included in the message, e.g. Unknown device 'Vulkan5'. Available devices: CPU, Vulkan0, Vulkan1.
  • IModelParams.Devices docs and ToLlamaModelParams <exception> tag updated.
  • Added ModelsParamsTests.UnknownDeviceThrows and DeviceNameMatchIsCaseInsensitive.

…loaded

The DllImport resolver only handled "llama" and "mtmd" and returned
IntPtr.Zero for "ggml" and "ggml-base", leaving them to the default
runtime probing. NativeLibraryUtils loads both by full path as
dependencies of llama, which the default probing finds again by name on
Windows and Linux but not on macOS (dlopen of a bare "libggml.dylib" does
not match the @rpath install name of the already-loaded library), so any
P/Invoke into ggml failed there with DllNotFoundException. This broke the
new Devices tests on the macOS CI job and also affects
TensorBufferOverrides, which uses the same imports.

Keep the ggml and ggml-base handles when they are loaded as dependencies,
and have the resolver return them (loading llama first if needed).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@ipevo-brooks

Copy link
Copy Markdown
Contributor Author

The new tests failed on the macOS job with DllNotFoundException: Unable to load shared library 'ggml'. Root cause is pre-existing and not specific to this PR: the DllImportResolver only handles llama and mtmd, so P/Invokes into ggml / ggml-base fall back to default runtime probing. NativeLibraryUtils loads both by full path as dependencies of llama, and default probing finds them again by name on Windows/Linux but not on macOS (dlopen of a bare libggml.dylib does not match the @rpath install name of the already-loaded library). TensorBufferOverrides uses the same imports and would hit the same failure on macOS.

Pushed 6789992: NativeLibraryUtils keeps the ggml / ggml-base handles when it loads them as dependencies, and the resolver returns those handles for the two names (loading llama first if it has not been loaded yet). macOS job is green now. Happy to split this into its own PR if you would rather keep it separate.

@martindevans martindevans left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for looking into that additional MacOS issue. This looks good to go to me 👍

@martindevans
martindevans merged commit 0dd3682 into SciSharp:master Oct 5, 2026
8 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants