Device configuration (sbd.device_config)

Device configuration helper for SBD Python bindings.

This module provides utilities to easily switch between CPU and GPU execution without changing user code.

class DeviceConfig(device='cpu', use_precalculated_dets=True, max_memory_gb=-1, use_gpu=None)[source]

Bases: object

Helper class to configure CPU vs GPU execution for SBD calculations.

Usage:

# Auto-detect (uses GPU if available)
config = DeviceConfig.auto()

# Force CPU
config = DeviceConfig.cpu()

# Force GPU with specific settings
config = DeviceConfig.gpu(max_memory_gb=16)

# Apply to SBD configuration
sbd_config = sbd.TPB_SBD()
config.apply(sbd_config)

Initialize device configuration.

Parameters:
  • device (str) – Backend device key — ‘cpu’, ‘gpu’ (NVHPC Thrust, NVIDIA-only), ‘gpu-omp’ (OpenMP target offload, NVIDIA or AMD), or any alias known to sbd._device_aliases. Default ‘cpu’.

  • use_precalculated_dets (bool) – Use precalculated determinants (GPU only)

  • max_memory_gb (int) – Maximum GPU memory in GB (-1 = auto)

  • use_gpu (bool | None) – Deprecated boolean. If supplied without device, True maps to ‘gpu’ (Thrust) and False to ‘cpu’ for backward compatibility with pre-OMP-offload code.

apply(sbd_config)[source]

Apply device configuration to an SBD TPB_SBD configuration object.

Parameters:

sbd_config – sbd.TPB_SBD configuration object

Return type:

None

classmethod auto(max_memory_gb=-1)[source]

Auto-detect the best available backend and use it.

Resolves against the backends that were actually COMPILED, not just the hardware that is present. Detecting a GPU and returning ‘gpu’ unconditionally was wrong in two ways: on an AMD host it selected the Thrust/CUDA backend, which cannot exist there (upstream wires Thrust to nvc++ -cuda), so a machine that reported “GPU detected (HIP)” then failed to load a CUDA module; and on a CPU-only build with a GPU present it picked a backend that was never built.

Preference order matches sbd._resolve_device(): Thrust (‘gpu’) first where it exists, since it is the long-validated NVIDIA default and keeps more phases on the device, then OpenMP offload (‘gpu-omp’), then CPU.

Parameters:

max_memory_gb (int) – Maximum GPU memory in GB (-1 = auto)

Returns:

DeviceConfig for the best backend that is both built and runnable

Return type:

DeviceConfig

classmethod cpu()[source]

Force CPU execution.

Return type:

DeviceConfig

classmethod gpu(use_precalculated_dets=True, max_memory_gb=-1)[source]

Force NVHPC Thrust GPU execution. NVIDIA only.

Requires SBD compiled with THRUST (the _core_gpu_thrust extension, i.e. SBD_BUILD_BACKEND=gpu, or the default auto).

There is no AMD equivalent: upstream SBD wires the Thrust path to nvc++ -cuda, so no rocThrust configuration exists to build. On an AMD host use gpu_omp() instead.

Parameters:
  • use_precalculated_dets (bool)

  • max_memory_gb (int)

Return type:

DeviceConfig

classmethod gpu_nvidia_omp(max_memory_gb=-1)[source]

Deprecated alias for gpu_omp().

The old name dates from when SBD shipped a LLVM-with-NVPTX offload backend distinct from the nvc++ path. The LLVM backend was removed in v1.6 (see tag v1.5.0-llvm for that history); gpu_omp is the single OpenMP-offload path now.

Parameters:

max_memory_gb (int)

Return type:

DeviceConfig

classmethod gpu_omp(max_memory_gb=-1)[source]

Force OpenMP target-offload GPU execution. Works on NVIDIA and AMD.

Requires SBD compiled with the OMP-offload backend (the _core_gpu_omp_offload extension), which the default SBD_BUILD_BACKEND=auto builds whenever a GPU compiler is present; narrow it to gpu_omp_offload to build only this one.

One module and one device string serve both vendors – the same source and macros, compiled by nvc++ -mp=gpu with NVHPC’s libnvomp, or by amdclang++ --offload-arch=gfx* with LLVM’s libomp/libomptarget. sbd.get_backend('gpu-omp').__sbd_offload_target__ reports which, e.g. 'amdgcn-amd-amdhsa:gfx90a'.

It installs alongside the CPU backend (and Thrust, on NVIDIA) – backends are imported lazily, one per process, which is what keeps them apart. The one combination to avoid in a single process is this backend together with the CPU one: they share an OpenMP runtime (libnvomp on NVIDIA, libomp on AMD), and loading _core_cpu first leaves it initialised host-only, after which offload regions run on the host. See sbd.has_backend_conflict().

Parameters:

max_memory_gb (int)

Return type:

DeviceConfig

get_device_info()[source]

Get information about available compute devices.

Returns:

Dictionary with device information

Return type:

dict

print_device_info()[source]

Print the compiled backends, then the hardware they could run on.

Backends come first because they are what actually constrains a run: the hardware being present says nothing about whether a backend was compiled for it. An earlier version printed “CPU Available: Always” unconditionally while sbd.available_backends() returned [] – reported as issue #9.