unsloth/studio/backend/tests/test_gpu_selection.py
Daniel Han 44a05f3496
Studio: stop the /api/system poll from pinning a CUDA/HIP primary context (#8677)
* Stop the /api/system poll from pinning a CUDA/HIP primary context

The frontend polls GET /api/system every 5s from the root-mounted floating
monitor and every 3s from Settings -> Resources. On any host where the smi
shortcut returns None (ROCm, or no nvidia-smi on PATH) both handlers reached
torch.cuda.mem_get_info, which attaches a primary context worth ~612 MiB on a
B200. A context is never released while the process lives, so an otherwise idle
Studio lost that memory to telemetry alone.

get_backend_visible_gpu_info only reads name and total and discards used, so it
now goes through _torch_get_device_inventory, which answers from
get_device_properties. props.total_memory is byte-identical to mem_get_info's
total half (191505498112 measured both ways here), so no displayed number moves.

On Linux ROCm get_visible_gpu_utilization asked torch for occupancy and then
overwrote it with sysfs figures in _overlay_system_wide_vram, spending a context
on numbers that were discarded. The overlay's decision is now split out from its
mutation, so the sysfs answer is taken first when it covers every visible device
and torch is consulted only when the overlay declines, which keeps a device
sysfs cannot speak for on its old path.

clear_gpu_cache called torch.cuda.synchronize() unconditionally on the CUDA
branch. With nothing reserved there is no work in flight to wait on, so the
synchronize is skipped when memory_reserved is zero across the visible devices;
memory_reserved needs no context itself, while is_initialized is useless as a
guard because reading device properties already flips it. empty_cache and
ipc_collect still run unconditionally, and the branch is deliberately left
unwrapped because the diffusion and video unload paths document that a sticky
CUDA fault has to propagate out of it.

Measured per fresh process with per-PID nvidia-smi: the visibility path went
612 MiB -> 0 MiB with name, total_gb, index and visible_ordinal unchanged.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Keep the GTT total on a ROCm APU when the poll skips mem_get_info

get_backend_visible_gpu_info moved to _torch_get_device_inventory on the
strength of props.total_memory being byte-identical to mem_get_info's total
half. That holds on NVIDIA and on discrete ROCm, but not on a unified-memory
APU, where this repo already documents the two as different numbers:
_allocator_divides_by_props_total records that c10 caps at fraction *
props.total_memory from torch 2.10 and at fraction * hipMemGetInfo total
through 2.9 precisely because they differ there, and
_apply_unified_memory_correction adopts the mem_get_info total over amd-smi's
because props reports the dedicated carve-out while hipMemGetInfo spans the
GTT pool.

So on a Strix Halo the visibility endpoint went from the ~100 GiB pool torch
can actually allocate into to the ~8 GiB carve-out, and memory_total_gb feeds
use-gpu-info's maxDeviceMemoryGb and loadDeviceMemoryGb, which gate whether a
model is offered at all. The same substitution also fed the new sysfs-first
probe: with the carve-out in hand the overlay's 10 percent total check starts
agreeing with the sysfs figure, so it adopts it instead of declining, which is
the exact APU case its docstring says it keeps torch's number for.

The inventory now asks mem_get_info for the total on unified devices only,
using the same classifier the training worker and the llama.cpp backend use so
all three agree on what an APU is. There is no context-free source for the GTT
total, so an APU pays for a context here; every discrete card and every NVIDIA
host keeps the free path, which is where the 612 MiB was being lost. A failed
probe keeps the carve-out rather than dropping the device.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Tighten the comments added by this PR

* Keep the driver total for a ROCm APU the classifier cannot place

_rocm_props_total_is_carve_out delegated the whole decision to
_rocm_classify_unified_memory, which answers unified or not-unified and never
"cannot tell". It knows an APU by hipDeviceProp_t.integrated or by a hardcoded
gfx1150/1151/1152 set, so a gfx1103 Phoenix iGPU whose runtime leaves that flag
at 0 came back not-unified, the inventory kept props.total_memory, and the
endpoint published the dedicated carve-out as the whole device. On such a part
that is an 8 GiB-class number standing in for the GTT pool, and it gates which
models are offered.

clr only fills the flag in from the HSA "agent is APU" property from ROCm 6.1.2
on; before that an APU answers 0 exactly like a discrete card, and the last
field of torch.version.hip is a build number, so 6.1.2 cannot be told from
6.1.0. The flag is therefore trusted from 6.2 on, and on anything older, on a
wheel that omits the field, and when the classifier itself raises, the driver
total is worth a context. A total that is too small hides models; the context
does not.

Discrete cards on a current ROCm runtime keep the context-free inventory, which
is the case this path was added for.

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-08-13 06:33:40 -07:00

2297 lines
90 KiB
Python

# SPDX-License-Identifier: AGPL-3.0-only
# Copyright 2026-present the Unsloth AI Inc. team. All rights reserved. See /studio/LICENSE.AGPL-3.0
import asyncio
import importlib.util
import os
import re
import sys
import unittest
from contextlib import nullcontext
from pathlib import Path
from types import ModuleType, SimpleNamespace
from unittest.mock import patch
from fastapi import HTTPException
from core.training.training import TrainingBackend
from models.inference import LoadRequest
from models.training import TrainingStartRequest
from utils.hardware import (
apply_gpu_ids,
DeviceType,
auto_select_gpu_ids,
estimate_required_model_memory_gb,
get_backend_visible_gpu_info,
get_device_map,
get_gpu_utilization,
get_offloaded_device_map_entries,
get_parent_visible_gpu_ids,
get_visible_gpu_utilization,
get_vulkan_inference_gpu_info,
prepare_gpu_selection,
resolve_requested_gpu_ids,
)
import utils.hardware.hardware as _hw_module
_BACKEND_ROOT = Path(__file__).resolve().parent.parent
async def _inline_to_thread(func, /, *args, **kwargs):
return func(*args, **kwargs)
def _fake_unsloth_attention_modules(resolver):
unsloth_module = ModuleType("unsloth")
models_module = ModuleType("unsloth.models")
utils_module = ModuleType("unsloth.models._utils")
utils_module.resolve_attention_implementation = resolver
models_module._utils = utils_module
unsloth_module.models = models_module
return {
"unsloth": unsloth_module,
"unsloth.models": models_module,
"unsloth.models._utils": utils_module,
}
def _load_route_module(name: str, relative_path: str):
spec = importlib.util.spec_from_file_location(name, _BACKEND_ROOT / relative_path)
module = importlib.util.module_from_spec(spec)
spec.loader.exec_module(module)
return module
class _GpuCacheResetMixin:
"""Reset module-level GPU caches between tests to prevent state leaks."""
def tearDown(self):
_hw_module._physical_gpu_count = None
_hw_module._visible_gpu_count = None
class TestResolveRequestedGpuIds(_GpuCacheResetMixin, unittest.TestCase):
def test_parent_visibility_defaults_to_physical_enumeration(self):
with (
patch.dict(os.environ, {}, clear = True),
patch("utils.hardware.hardware.get_physical_gpu_count", return_value = 4),
):
self.assertEqual(get_parent_visible_gpu_ids(), [0, 1, 2, 3])
self.assertEqual(resolve_requested_gpu_ids(None), [0, 1, 2, 3])
def test_parent_visibility_uses_cuda_visible_devices(self):
with patch.dict(os.environ, {"CUDA_VISIBLE_DEVICES": "1,3"}, clear = True):
self.assertEqual(get_parent_visible_gpu_ids(), [1, 3])
self.assertEqual(resolve_requested_gpu_ids(None), [1, 3])
def test_parent_visibility_uses_empty_numeric_ids_for_uuid_masks(self):
with (
patch.dict(os.environ, {"CUDA_VISIBLE_DEVICES": "GPU-aaa,GPU-bbb"}, clear = True),
patch("utils.hardware.hardware.get_physical_gpu_count", return_value = 8),
):
self.assertEqual(get_parent_visible_gpu_ids(), [])
def test_invalid_requests_raise_clear_value_errors(self):
cases = [
([1, 1], "duplicate GPU IDs"),
([-1], "Rejected IDs: [-1]"),
([99], "Rejected IDs: [99]"),
([0], "outside the parent-visible set [1, 3]"),
]
with (
patch.dict(os.environ, {"CUDA_VISIBLE_DEVICES": "1,3"}, clear = True),
patch("utils.hardware.hardware.get_physical_gpu_count", return_value = 8),
):
for gpu_ids, message in cases:
with self.subTest(gpu_ids = gpu_ids):
with self.assertRaisesRegex(ValueError, re.escape(message)):
resolve_requested_gpu_ids(gpu_ids)
def test_explicit_ids_must_be_physical_not_relative(self):
with (
patch.dict(os.environ, {"CUDA_VISIBLE_DEVICES": "1,3"}, clear = True),
patch("utils.hardware.hardware.get_physical_gpu_count", return_value = 8),
):
self.assertEqual(resolve_requested_gpu_ids([1, 3]), [1, 3])
def test_explicit_ids_are_rejected_for_uuid_parent_visibility(self):
with (
patch.dict(os.environ, {"CUDA_VISIBLE_DEVICES": "GPU-aaa,GPU-bbb"}, clear = True),
patch("utils.hardware.hardware.get_physical_gpu_count", return_value = 8),
):
with self.assertRaisesRegex(
ValueError,
"unsupported when CUDA_VISIBLE_DEVICES uses non-numeric or subdevice",
):
resolve_requested_gpu_ids([1])
def test_empty_list_is_treated_as_auto(self):
with (
patch.dict(os.environ, {"CUDA_VISIBLE_DEVICES": "1,3"}, clear = True),
patch("utils.hardware.hardware.get_physical_gpu_count", return_value = 8),
):
self.assertEqual(resolve_requested_gpu_ids([]), [1, 3])
def test_vulkan_ordinals_bypass_cuda_parent_visible_validation(self):
# Vulkan build on a CPU-only torch host: no CUDA parent-visible set and a zero physical count,
# yet a valid Vulkan ordinal must not be rejected as a CUDA physical id (issue #7239).
with (
patch.dict(os.environ, {}, clear = True),
patch("utils.hardware.hardware.get_physical_gpu_count", return_value = 0),
):
# As a CUDA physical id, [0] is outside the empty parent-visible set.
with self.assertRaises(ValueError):
resolve_requested_gpu_ids([0])
# As Vulkan ordinals, [0] and [0, 1] pass through unchanged.
self.assertEqual(resolve_requested_gpu_ids([0], is_vulkan = True), [0])
self.assertEqual(resolve_requested_gpu_ids([0, 1], is_vulkan = True), [0, 1])
# Malformed ordinals are still rejected.
with self.assertRaisesRegex(ValueError, "duplicate GPU IDs"):
resolve_requested_gpu_ids([0, 0], is_vulkan = True)
with self.assertRaisesRegex(ValueError, "non-negative"):
resolve_requested_gpu_ids([-1], is_vulkan = True)
def test_apply_gpu_ids_only_updates_cuda_visible_devices(self):
with patch.dict(
os.environ,
{"CUDA_VISIBLE_DEVICES": "1,3", "TEST_PARENT_ENV": "keep-me"},
clear = True,
):
apply_gpu_ids([5, 6])
self.assertEqual(os.environ["CUDA_VISIBLE_DEVICES"], "5,6")
self.assertEqual(os.environ["TEST_PARENT_ENV"], "keep-me")
class TestVisibleGpuUtilization(_GpuCacheResetMixin, unittest.TestCase):
def test_gpu_utilization_preserves_primary_shape_with_devices(self):
devices = [
{
"index": 5,
"visible_ordinal": 0,
"gpu_utilization_pct": 11.0,
"temperature_c": 40.0,
"vram_used_gb": 4.0,
"vram_total_gb": 24.0,
"vram_utilization_pct": 16.7,
"power_draw_w": 80.0,
"power_limit_w": 300.0,
"power_utilization_pct": 26.7,
},
{
"index": 3,
"visible_ordinal": 1,
"gpu_utilization_pct": 22.0,
"temperature_c": 50.0,
"vram_used_gb": 8.0,
"vram_total_gb": 24.0,
"vram_utilization_pct": 33.3,
"power_draw_w": 120.0,
"power_limit_w": 300.0,
"power_utilization_pct": 40.0,
},
]
with (
patch("utils.hardware.hardware.get_device", return_value = DeviceType.CUDA),
patch.object(_hw_module, "IS_ROCM", False),
patch(
"utils.hardware.hardware._get_parent_visible_gpu_spec",
return_value = {"raw": "5,3", "numeric_ids": [5, 3]},
),
patch(
"utils.hardware.hardware._smi_query",
return_value = {
"available": True,
"devices": devices,
"backend_cuda_visible_devices": "5,3",
"parent_visible_gpu_ids": [5, 3],
"index_kind": "physical",
},
),
):
result = get_gpu_utilization()
self.assertIsInstance(result, dict)
self.assertTrue(result["available"])
self.assertEqual(result["backend"], "cuda")
self.assertEqual(result["index"], 5)
self.assertEqual(result["visible_ordinal"], 0)
self.assertEqual(result["vram_total_gb"], 24.0)
self.assertEqual(result["parent_visible_gpu_ids"], [5, 3])
self.assertEqual([device["index"] for device in result["devices"]], [5, 3])
def test_gpu_utilization_cpu_returns_legacy_unavailable_object(self):
with patch("utils.hardware.hardware.get_device", return_value = DeviceType.CPU):
result = get_gpu_utilization()
self.assertEqual(result, {"available": False, "backend": "cpu", "devices": []})
def test_gpu_utilization_mlx_stays_available_without_agx_stats(self):
fake_psutil = ModuleType("psutil")
fake_psutil.virtual_memory = lambda: SimpleNamespace(total = 64 * 1024**3)
with (
patch.dict(sys.modules, {"psutil": fake_psutil}),
patch("utils.hardware.hardware.get_device", return_value = DeviceType.MLX),
patch("utils.hardware.hardware._read_apple_gpu_stats", return_value = {}),
patch(
"core.training.get_training_backend",
return_value = SimpleNamespace(_progress = None),
),
patch("utils.hardware.apple.read_gpu_temperature_c", return_value = None),
patch("utils.hardware.apple.read_gpu_power_w", return_value = None),
):
result = get_gpu_utilization()
self.assertTrue(result["available"])
self.assertEqual(result["backend"], "mlx")
self.assertIsNone(result["gpu_utilization_pct"])
self.assertEqual(result["vram_used_gb"], 0)
self.assertEqual(result["vram_total_gb"], 64.0)
self.assertEqual(len(result["devices"]), 1)
def test_gpu_utilization_xpu_uses_visible_devices(self):
with (
patch("utils.hardware.hardware.get_device", return_value = DeviceType.XPU),
patch(
"utils.hardware.hardware.get_visible_gpu_utilization",
return_value = {
"available": True,
"backend": "xpu",
"parent_visible_gpu_ids": [2, 0],
"index_kind": "physical",
"devices": [
{
"index": 2,
"visible_ordinal": 1,
"gpu_utilization_pct": None,
"temperature_c": None,
"vram_used_gb": 3.0,
"vram_total_gb": 16.0,
"vram_utilization_pct": 18.8,
"power_draw_w": None,
"power_limit_w": None,
"power_utilization_pct": None,
},
{
"index": 0,
"visible_ordinal": 0,
"gpu_utilization_pct": None,
"temperature_c": None,
"vram_used_gb": 1.0,
"vram_total_gb": 16.0,
"vram_utilization_pct": 6.3,
"power_draw_w": None,
"power_limit_w": None,
"power_utilization_pct": None,
},
],
},
),
):
result = get_gpu_utilization()
self.assertEqual(result["backend"], "xpu")
self.assertEqual(result["index"], 0)
self.assertEqual(result["visible_ordinal"], 0)
self.assertEqual([device["index"] for device in result["devices"]], [0, 2])
def test_visible_gpu_utilization_filters_to_parent_visible_ids(self):
smi_output = "\n".join(
[
"0, 10, 30, 1000, 10000, 50, 100",
"1, 20, 40, 2000, 10000, 60, 120",
"3, 30, 50, 3000, 10000, 70, 140",
]
)
with (
patch.dict(os.environ, {"CUDA_VISIBLE_DEVICES": "1,3"}, clear = True),
patch("utils.hardware.hardware.get_device", return_value = DeviceType.CUDA),
patch("utils.hardware.nvidia.subprocess.run") as mock_run,
):
mock_run.return_value = SimpleNamespace(
returncode = 0,
stdout = smi_output,
)
result = get_visible_gpu_utilization()
self.assertTrue(result["available"])
self.assertEqual(result["parent_visible_gpu_ids"], [1, 3])
self.assertEqual(result["index_kind"], "physical")
self.assertEqual([device["index"] for device in result["devices"]], [1, 3])
self.assertEqual(result["devices"][0]["visible_ordinal"], 0)
self.assertEqual(result["devices"][1]["visible_ordinal"], 1)
self.assertEqual(result["devices"][0]["gpu_utilization_pct"], 20.0)
self.assertEqual(result["devices"][1]["power_utilization_pct"], 50.0)
def test_backend_visible_gpu_info_preserves_physical_indices(self):
smi_output = "\n".join(
[
"0, GPU Zero, 10000",
"1, GPU One, 20000",
"3, GPU Three, 30000",
]
)
with (
patch.dict(os.environ, {"CUDA_VISIBLE_DEVICES": "1,3"}, clear = True),
patch("utils.hardware.hardware.get_device", return_value = DeviceType.CUDA),
patch("utils.hardware.nvidia.subprocess.run") as mock_run,
):
mock_run.return_value = SimpleNamespace(
returncode = 0,
stdout = smi_output,
)
result = get_backend_visible_gpu_info()
self.assertTrue(result["available"])
self.assertEqual(result["parent_visible_gpu_ids"], [1, 3])
self.assertEqual(result["index_kind"], "physical")
self.assertEqual([device["index"] for device in result["devices"]], [1, 3])
self.assertEqual(result["devices"][0]["visible_ordinal"], 0)
self.assertEqual(result["devices"][1]["visible_ordinal"], 1)
self.assertEqual(result["devices"][0]["name"], "GPU One")
self.assertAlmostEqual(result["devices"][1]["memory_total_gb"], 29.3, places = 1)
def test_uuid_parent_visibility_falls_back_to_torch(self):
"""UUID/MIG masks fall through nvidia to the torch fallback and
still report visible devices using relative ordinals."""
# Inventory shape: this endpoint reads name and total and discards used, so
# it asks for the context-free helper.
fake_torch_devices = [
{
"index": 0,
"visible_ordinal": 0,
"name": "GPU-A",
"total_gb": 24.0,
"used_gb": None,
},
{
"index": 1,
"visible_ordinal": 1,
"name": "GPU-B",
"total_gb": 24.0,
"used_gb": None,
},
]
with (
patch.dict(os.environ, {"CUDA_VISIBLE_DEVICES": "GPU-aaa,GPU-bbb"}, clear = True),
patch("utils.hardware.hardware.get_device", return_value = DeviceType.CUDA),
patch("utils.hardware.hardware._torch_get_physical_gpu_count", return_value = 2),
patch(
"utils.hardware.hardware._torch_get_device_inventory",
return_value = fake_torch_devices,
),
):
result = get_backend_visible_gpu_info()
self.assertTrue(result["available"])
self.assertEqual(result["parent_visible_gpu_ids"], [])
self.assertEqual(len(result["devices"]), 2)
self.assertEqual(result["index_kind"], "relative")
def test_mlx_visible_gpu_info_is_best_effort_relative(self):
with (
patch("utils.hardware.hardware.get_device", return_value = DeviceType.MLX),
patch(
"utils.hardware.hardware.get_gpu_memory_info",
return_value = {
"available": True,
"device_name": "Apple Silicon",
"total_gb": 64.0,
"allocated_gb": 8.0,
"utilization_pct": 12.5,
},
),
):
result = get_backend_visible_gpu_info()
self.assertTrue(result["available"])
self.assertEqual(result["index_kind"], "relative")
self.assertEqual(result["devices"][0]["index"], 0)
self.assertEqual(result["devices"][0]["visible_ordinal"], 0)
def test_discrete_vulkan_inference_gpu_info(self):
with (
patch(
"core.inference.llama_cpp.LlamaCppBackend._is_vulkan_backend",
return_value = True,
),
patch(
"core.inference.llama_cpp.LlamaCppBackend.vulkan_device_inventory",
return_value = [
{
"index": 0,
"name": "Vulkan0",
"free_mib": 7402,
"total_mib": 8192,
"is_igpu": False,
}
],
),
):
result = get_vulkan_inference_gpu_info()
self.assertTrue(result["available"])
self.assertEqual(result["backend"], "vulkan")
# ggml Vulkan ordinals are the space `--device Vulkan<i>` pins, so they are selectable.
self.assertEqual(result["index_kind"], "vulkan")
self.assertEqual(result["parent_visible_gpu_ids"], [])
self.assertEqual(
result["devices"],
[
{
"index": 0,
"index_kind": "vulkan",
"visible_ordinal": 0,
"name": "Vulkan0",
"memory_total_gb": 8.0,
"vram_used_gb": 0.77,
"vram_free_gb": 7.23,
"vram_utilization_pct": 9.6,
"shared_memory": False,
}
],
)
def test_vulkan_igpu_info_uses_capped_free_budget(self):
with (
patch(
"core.inference.llama_cpp.LlamaCppBackend._is_vulkan_backend",
return_value = True,
),
patch(
"core.inference.llama_cpp.LlamaCppBackend.vulkan_device_inventory",
return_value = [
{
"index": 0,
"name": "Vulkan0",
"free_mib": 12288,
"total_mib": 32768,
"is_igpu": True,
}
],
),
patch(
"core.inference.llama_cpp._apply_igpu_host_reserve_mib",
return_value = 12288,
),
):
result = get_vulkan_inference_gpu_info()
device = result["devices"][0]
self.assertEqual(device["memory_total_gb"], 12.0)
self.assertEqual(device["vram_free_gb"], 12.0)
self.assertIsNone(device["vram_used_gb"])
self.assertIsNone(device["vram_utilization_pct"])
self.assertTrue(device["shared_memory"])
def test_forced_vulkan_overrides_torch_gpu_visibility_for_inference(self):
with (
patch("utils.hardware.hardware.get_device", return_value = DeviceType.CUDA),
patch(
"core.inference.llama_cpp.LlamaCppBackend._is_vulkan_backend",
return_value = True,
),
patch(
"core.inference.llama_cpp.LlamaCppBackend.vulkan_device_inventory",
return_value = [
{
"index": 1,
"name": "Vulkan1",
"free_mib": 6144,
"total_mib": 8192,
"is_igpu": False,
}
],
),
patch(
"utils.hardware.nvidia.get_backend_visible_gpu_info",
return_value = {
"available": True,
"backend": "cuda",
"devices": [{"index": 0, "name": "CUDA0", "memory_total_gb": 24.0}],
},
),
patch(
"utils.hardware.hardware._get_parent_visible_gpu_spec",
return_value = {"raw": None, "numeric_ids": None},
),
):
training_result = get_backend_visible_gpu_info()
inference_result = get_vulkan_inference_gpu_info()
self.assertEqual(training_result["backend"], "cuda")
self.assertEqual(inference_result["backend"], "vulkan")
self.assertEqual(inference_result["devices"][0]["index"], 1)
def test_vulkan_install_without_devices_reports_unavailable(self):
with (
patch(
"core.inference.llama_cpp.LlamaCppBackend._is_vulkan_backend",
return_value = True,
),
patch(
"core.inference.llama_cpp.LlamaCppBackend.vulkan_device_inventory",
return_value = [],
),
):
result = get_vulkan_inference_gpu_info()
self.assertFalse(result["available"])
self.assertEqual(result["backend"], "vulkan")
self.assertEqual(result["devices"], [])
class TestGpuAutoSelection(_GpuCacheResetMixin, unittest.TestCase):
def test_get_device_map_uses_explicit_gpu_selection(self):
with patch("utils.hardware.hardware.get_device", return_value = DeviceType.CUDA):
self.assertEqual(get_device_map(None), "sequential")
self.assertEqual(get_device_map([0]), "sequential")
self.assertEqual(get_device_map([0, 1]), "balanced")
def test_get_device_map_uses_all_inherited_visible_gpus_for_uuid_masks(self):
with (
patch.dict(os.environ, {"CUDA_VISIBLE_DEVICES": "GPU-aaa,GPU-bbb"}, clear = True),
patch("utils.hardware.hardware.get_device", return_value = DeviceType.CUDA),
):
self.assertEqual(get_device_map(None), "balanced")
def test_get_offloaded_device_map_entries_returns_only_cpu_and_disk(self):
model = SimpleNamespace(
hf_device_map = {
"model.embed_tokens": 0,
"model.layers.0": 1,
"model.layers.1": "cpu",
"lm_head": "disk",
}
)
self.assertEqual(
get_offloaded_device_map_entries(model),
{
"model.layers.1": "cpu",
"lm_head": "disk",
},
)
def test_get_offloaded_device_map_entries_handles_models_without_device_map(self):
self.assertEqual(get_offloaded_device_map_entries(SimpleNamespace()), {})
@patch(
"utils.hardware.hardware._resolve_model_identifier_for_gpu_estimate",
new = lambda model_name, **_: model_name,
)
@patch(
"utils.hardware.hardware._load_config_for_gpu_estimate",
new = lambda *_args, **_kwargs: None,
)
def test_estimate_required_memory_formulas(self):
eight_gb = 8 * (1024**3)
with patch(
"utils.hardware.hardware.estimate_fp16_model_size_bytes",
return_value = (eight_gb, "config"),
):
# FP16 inference: 8GB * 1.3 = 10.4GB
required_gb, metadata = estimate_required_model_memory_gb(
"unsloth/test",
load_in_4bit = False,
)
self.assertAlmostEqual(required_gb, 10.4, places = 3)
self.assertEqual(metadata["model_size_source"], "config")
# 4bit inference: base_4bit = 8/3.2 = 2.5GB
# required = 2.5 + max(2.5*0.3, 2.0) = 2.5 + 2.0 = 4.5GB
required_gb, _ = estimate_required_model_memory_gb(
"unsloth/test",
load_in_4bit = True,
)
self.assertAlmostEqual(required_gb, 4.5, places = 2)
# Full FT fallback: model_size * 3.5 + overhead
required_gb, metadata = estimate_required_model_memory_gb(
"unsloth/test", training_type = "Full Finetuning"
)
self.assertEqual(metadata.get("estimation_mode"), "fallback")
self.assertGreater(required_gb, 25.0)
self.assertLess(required_gb, 40.0)
# LoRA fp16 fallback: model_size + lora_overhead + activations + overhead
required_gb, metadata = estimate_required_model_memory_gb(
"unsloth/test",
training_type = "LoRA/QLoRA",
load_in_4bit = False,
)
self.assertEqual(metadata.get("estimation_mode"), "fallback")
self.assertGreater(required_gb, 8.0)
self.assertLess(required_gb, 15.0)
# QLoRA 4-bit fallback: compressed weights + lora overhead + activations + overhead
required_gb, metadata = estimate_required_model_memory_gb(
"unsloth/test",
training_type = "LoRA/QLoRA",
load_in_4bit = True,
)
self.assertEqual(metadata.get("estimation_mode"), "fallback")
self.assertGreater(required_gb, 3.0)
self.assertLess(required_gb, 8.0)
# Larger model: 16GB fp16
sixteen_gb = 16 * (1024**3)
with patch(
"utils.hardware.hardware.estimate_fp16_model_size_bytes",
return_value = (sixteen_gb, "config"),
):
required_gb, _ = estimate_required_model_memory_gb(
"unsloth/test",
training_type = "LoRA/QLoRA",
load_in_4bit = True,
)
# QLoRA for 16GB model should be < 12 GB
self.assertGreater(required_gb, 5.0)
self.assertLess(required_gb, 12.0)
def test_estimate_fp16_model_size_bytes_uses_vllm_fallback_last(self):
config = object()
with (
patch(
"utils.hardware.hardware._resolve_model_identifier_for_gpu_estimate",
return_value = "unsloth/test",
),
patch(
"utils.hardware.hardware._get_hf_safetensors_total_params",
return_value = None,
),
patch(
"utils.hardware.hardware._load_config_for_gpu_estimate",
return_value = config,
),
patch(
"utils.hardware.hardware._estimate_fp16_model_size_bytes_from_config",
return_value = None,
),
patch(
"utils.hardware.hardware._get_local_weight_size_bytes",
return_value = None,
),
patch(
"utils.hardware.hardware._estimate_fp16_model_size_bytes_from_vllm_utils",
return_value = 1234,
),
):
model_size_bytes, source = _hw_module.estimate_fp16_model_size_bytes("unsloth/test")
self.assertEqual(model_size_bytes, 1234)
self.assertEqual(source, "vllm_utils")
def test_offline_safetensors_probe_uses_config_without_hub_access(self):
config = object()
for offline_variable in ("HF_HUB_OFFLINE", "TRANSFORMERS_OFFLINE"):
with self.subTest(offline_variable = offline_variable):
with (
patch.dict(os.environ, {offline_variable: "true"}, clear = True),
patch("huggingface_hub.model_info") as hub_info,
patch(
"utils.hardware.hardware._resolve_model_identifier_for_gpu_estimate",
return_value = "unsloth/test",
),
patch(
"utils.hardware.hardware._load_config_for_gpu_estimate",
return_value = config,
),
patch(
"utils.hardware.hardware._estimate_fp16_model_size_bytes_from_config",
return_value = 1234,
),
patch(
"utils.hardware.hardware._get_local_weight_size_bytes",
return_value = None,
),
):
model_size_bytes, source = _hw_module.estimate_fp16_model_size_bytes(
"unsloth/test"
)
self.assertEqual((model_size_bytes, source), (1234, "config"))
hub_info.assert_not_called()
def test_auto_select_gpu_ids_chooses_smallest_fitting_subset(self):
fake_devices = {
"devices": [
{"index": 0, "vram_total_gb": 16.0, "vram_used_gb": 4.0},
{"index": 1, "vram_total_gb": 16.0, "vram_used_gb": 6.0},
{"index": 2, "vram_total_gb": 16.0, "vram_used_gb": 7.0},
]
}
with (
patch("utils.hardware.hardware.get_device", return_value = DeviceType.CUDA),
patch(
"utils.hardware.hardware.estimate_required_model_memory_gb",
return_value = (
14.0,
{"required_gb": 14.0, "model_size_source": "config"},
),
),
patch(
"utils.hardware.hardware.get_visible_gpu_utilization",
return_value = fake_devices,
),
):
selected, metadata = auto_select_gpu_ids("unsloth/test")
self.assertEqual(selected, [0, 1])
self.assertEqual(metadata["selection_mode"], "auto")
# First GPU full (12GB) + second GPU with overhead (10*0.85=8.5) = 20.5GB
self.assertAlmostEqual(metadata["usable_gb"], 20.5, places = 3)
def test_auto_select_gpu_ids_falls_back_to_all_visible(self):
fake_devices = {
"devices": [
{"index": 0, "vram_total_gb": 12.0, "vram_used_gb": 2.0},
{"index": 1, "vram_total_gb": 12.0, "vram_used_gb": 2.0},
]
}
with (
patch("utils.hardware.hardware.get_device", return_value = DeviceType.CUDA),
patch(
"utils.hardware.hardware.estimate_required_model_memory_gb",
return_value = (
30.0,
{"required_gb": 30.0, "model_size_source": "config"},
),
),
patch(
"utils.hardware.hardware.get_visible_gpu_utilization",
return_value = fake_devices,
),
):
selected, metadata = auto_select_gpu_ids("unsloth/test")
self.assertEqual(selected, [0, 1])
self.assertEqual(metadata["selection_mode"], "fallback_all")
# First GPU full (10GB) + second GPU with overhead (10*0.85=8.5) = 18.5GB
self.assertAlmostEqual(metadata["usable_gb"], 18.5, places = 3)
def test_prepare_gpu_selection_preserves_explicit_ids_without_auto_selection(self):
with (
patch("utils.hardware.hardware.get_device", return_value = DeviceType.CUDA),
patch(
"utils.hardware.hardware.resolve_requested_gpu_ids",
return_value = [2, 3],
),
patch("utils.hardware.hardware.auto_select_gpu_ids") as mock_auto_select,
):
selected, metadata = prepare_gpu_selection(
[2, 3],
model_name = "unsloth/test",
)
self.assertEqual(selected, [2, 3])
self.assertEqual(metadata["selection_mode"], "explicit")
mock_auto_select.assert_not_called()
def test_prepare_gpu_selection_treats_empty_list_as_auto(self):
with patch(
"utils.hardware.hardware.auto_select_gpu_ids",
return_value = ([0, 1], {"selection_mode": "auto"}),
) as mock_auto_select:
selected, metadata = prepare_gpu_selection(
[],
model_name = "unsloth/test",
)
self.assertEqual(selected, [0, 1])
self.assertEqual(metadata["selection_mode"], "auto")
mock_auto_select.assert_called_once()
def test_prepare_gpu_selection_preserves_uuid_parent_visibility_in_auto_mode(self):
with (
patch.dict(os.environ, {"CUDA_VISIBLE_DEVICES": "GPU-aaa,GPU-bbb"}, clear = True),
patch("utils.hardware.hardware.get_device", return_value = DeviceType.CUDA),
patch(
"utils.hardware.hardware.estimate_required_model_memory_gb",
return_value = (
14.0,
{"required_gb": 14.0, "model_size_source": "config"},
),
),
):
selected, metadata = prepare_gpu_selection(
None,
model_name = "unsloth/test",
)
self.assertIsNone(selected)
self.assertEqual(metadata["selection_mode"], "inherit_parent_visible")
self.assertIsNone(metadata["selected_gpu_ids"])
class TestPreSpawnGpuResolution(_GpuCacheResetMixin, unittest.TestCase):
def test_training_backend_resolves_explicit_gpu_ids_before_spawn(self):
backend = TrainingBackend()
class DummyProcess:
pid = 12345
def start(self):
return None
class DummyThread:
def start(self):
return None
dummy_queue = object()
with (
patch(
"core.training.training.prepare_gpu_selection",
return_value = ([1, 2], {"selection_mode": "explicit"}),
),
patch(
"core.training.training._CTX.Queue",
side_effect = [dummy_queue, dummy_queue],
),
patch(
"core.training.training._CTX.Process", return_value = DummyProcess()
) as mock_process,
patch("core.training.training.threading.Thread", return_value = DummyThread()),
):
backend.start_training(
job_id = "test-job-1",
model_name = "unsloth/test",
training_type = "LoRA/QLoRA",
gpu_ids = [1, 2],
)
config = mock_process.call_args.kwargs["kwargs"]["config"]
self.assertEqual(config["gpu_ids"], [1, 2])
self.assertEqual(config["resolved_gpu_ids"], [1, 2])
self.assertEqual(config["gpu_selection"]["selection_mode"], "explicit")
def test_training_backend_auto_selects_gpu_ids_when_omitted(self):
backend = TrainingBackend()
class DummyProcess:
pid = 12345
def start(self):
return None
class DummyThread:
def start(self):
return None
dummy_queue = object()
with (
patch(
"core.training.training.prepare_gpu_selection",
return_value = ([0, 1], {"selection_mode": "auto"}),
),
patch(
"core.training.training._CTX.Queue",
side_effect = [dummy_queue, dummy_queue],
),
patch(
"core.training.training._CTX.Process", return_value = DummyProcess()
) as mock_process,
patch("core.training.training.threading.Thread", return_value = DummyThread()),
):
backend.start_training(
job_id = "test-job-2",
model_name = "unsloth/test",
training_type = "LoRA/QLoRA",
gpu_ids = None,
)
config = mock_process.call_args.kwargs["kwargs"]["config"]
self.assertIsNone(config["gpu_ids"])
self.assertEqual(config["resolved_gpu_ids"], [0, 1])
self.assertEqual(config["gpu_selection"]["selection_mode"], "auto")
def test_training_backend_preserves_uuid_parent_visibility_in_auto_mode(self):
backend = TrainingBackend()
class DummyProcess:
pid = 12345
def start(self):
return None
class DummyThread:
def start(self):
return None
dummy_queue = object()
with (
patch.dict(os.environ, {"CUDA_VISIBLE_DEVICES": "GPU-aaa,GPU-bbb"}, clear = True),
patch("utils.hardware.hardware.get_device", return_value = DeviceType.CUDA),
patch(
"core.training.training._CTX.Queue",
side_effect = [dummy_queue, dummy_queue],
),
patch(
"core.training.training._CTX.Process", return_value = DummyProcess()
) as mock_process,
patch("core.training.training.threading.Thread", return_value = DummyThread()),
patch(
"utils.hardware.hardware.estimate_required_model_memory_gb",
return_value = (
14.0,
{"required_gb": 14.0, "model_size_source": "config"},
),
),
):
backend.start_training(
job_id = "test-job-uuid-auto",
model_name = "unsloth/test",
training_type = "LoRA/QLoRA",
gpu_ids = None,
)
config = mock_process.call_args.kwargs["kwargs"]["config"]
self.assertIsNone(config["resolved_gpu_ids"])
self.assertEqual(config["gpu_selection"]["selection_mode"], "inherit_parent_visible")
def test_inference_orchestrator_resolves_explicit_gpu_ids_before_spawn(self):
class DummyThread:
def __init__(self, *args, **kwargs):
pass
def start(self):
return None
with patch("core.inference.orchestrator.threading.Thread", DummyThread):
from core.inference.orchestrator import InferenceOrchestrator
orchestrator = InferenceOrchestrator()
config = SimpleNamespace(identifier = "unsloth/test", gguf_variant = None)
with (
patch(
"core.inference.orchestrator.prepare_gpu_selection",
return_value = ([1], {"selection_mode": "explicit"}),
),
patch.object(orchestrator, "_ensure_subprocess_alive", return_value = False),
patch.object(orchestrator, "_spawn_subprocess") as mock_spawn,
patch.object(
orchestrator,
"_wait_response",
return_value = {"success": True, "model_info": {}},
),
patch("utils.transformers_version.needs_transformers_5", return_value = False),
):
self.assertTrue(orchestrator.load_model(config = config, gpu_ids = [1]))
sub_config = mock_spawn.call_args.args[0]
self.assertEqual(sub_config["gpu_ids"], [1])
self.assertEqual(sub_config["resolved_gpu_ids"], [1])
self.assertEqual(sub_config["gpu_selection"]["selection_mode"], "explicit")
def test_inference_orchestrator_auto_selects_gpu_ids_when_omitted(self):
class DummyThread:
def __init__(self, *args, **kwargs):
pass
def start(self):
return None
with patch("core.inference.orchestrator.threading.Thread", DummyThread):
from core.inference.orchestrator import InferenceOrchestrator
orchestrator = InferenceOrchestrator()
config = SimpleNamespace(identifier = "unsloth/test", gguf_variant = None)
with (
patch(
"core.inference.orchestrator.prepare_gpu_selection",
return_value = ([0], {"selection_mode": "auto"}),
),
patch.object(orchestrator, "_ensure_subprocess_alive", return_value = False),
patch.object(orchestrator, "_spawn_subprocess") as mock_spawn,
patch.object(
orchestrator,
"_wait_response",
return_value = {"success": True, "model_info": {}},
),
patch("utils.transformers_version.needs_transformers_5", return_value = False),
):
self.assertTrue(orchestrator.load_model(config = config, gpu_ids = None))
sub_config = mock_spawn.call_args.args[0]
self.assertIsNone(sub_config["gpu_ids"])
self.assertEqual(sub_config["resolved_gpu_ids"], [0])
self.assertEqual(sub_config["gpu_selection"]["selection_mode"], "auto")
class TestRouteErrors(unittest.TestCase):
def test_prepare_gpu_selection_rejects_gpu_ids_on_non_accelerator_backend(self):
with patch("utils.hardware.hardware.get_device", return_value = DeviceType.CPU):
with self.assertRaises(ValueError) as exc_info:
prepare_gpu_selection([0], model_name = "unsloth/test")
self.assertIn("only supported on CUDA and Intel XPU", str(exc_info.exception))
def test_inference_route_resolves_gguf_gpu_ids(self):
# GGUF IDs use the normal resolver instead of a blanket rejection.
import utils.hardware.hardware as hardware_mod
inference_route = _load_route_module(
"inference_route_module_for_gguf_gpu_ids_test",
"routes/inference.py",
)
request = LoadRequest(model_path = "unsloth/test.gguf", gpu_ids = [0, 1])
model_config = SimpleNamespace(
is_gguf = True,
is_lora = False,
gguf_hf_repo = None,
gguf_file = "/tmp/test.gguf",
gguf_mmproj_file = None,
gguf_variant = None,
identifier = "unsloth/test.gguf",
display_name = "unsloth/test.gguf",
is_vision = False,
is_audio = False,
audio_type = None,
has_audio_input = False,
)
def _fake_resolve(ids, is_vulkan = False):
raise ValueError("SENTINEL requested GPUs are outside the parent-visible set")
with (
patch.object(
inference_route,
"ModelConfig",
SimpleNamespace(from_identifier = lambda **_kwargs: model_config),
),
patch("utils.hardware.resolve_requested_gpu_ids", _fake_resolve),
patch.object(hardware_mod, "resolve_requested_gpu_ids", _fake_resolve),
patch.object(
inference_route,
"_guard_chat_load_against_training",
return_value = None,
),
patch.object(inference_route.asyncio, "to_thread", new = _inline_to_thread),
patch.object(inference_route, "_hf_offline_if_unreachable", nullcontext),
):
with self.assertRaises(HTTPException) as exc_info:
asyncio.run(
inference_route._load_model_impl(
request,
SimpleNamespace(
app = SimpleNamespace(
state = SimpleNamespace(llama_parallel_slots = 1),
),
),
current_subject = "test-user",
)
)
self.assertEqual(exc_info.exception.status_code, 400)
self.assertIn("SENTINEL", exc_info.exception.detail)
self.assertNotIn("not supported for GGUF", exc_info.exception.detail)
def test_load_rejects_unavailable_vulkan_ordinal_before_training_guard(self):
inference_route = _load_route_module(
"inference_route_module_for_vulkan_preflight_test",
"routes/inference.py",
)
request = LoadRequest(model_path = "unsloth/test.gguf", gpu_ids = [99])
model_config = SimpleNamespace(
is_gguf = True,
is_lora = False,
gguf_hf_repo = None,
gguf_file = "/tmp/test.gguf",
gguf_mmproj_file = None,
gguf_variant = None,
identifier = "unsloth/test.gguf",
display_name = "unsloth/test.gguf",
is_vision = False,
is_audio = False,
audio_type = None,
has_audio_input = False,
)
with (
patch.object(
inference_route,
"ModelConfig",
SimpleNamespace(from_identifier = lambda **_kwargs: model_config),
),
patch("utils.hardware.get_device", return_value = DeviceType.CUDA),
patch.object(inference_route, "_classify_diffusion_gguf", return_value = None),
patch.object(
inference_route.LlamaCppBackend,
"_is_vulkan_backend",
return_value = True,
),
patch.object(
inference_route.LlamaCppBackend,
"_find_llama_server_binary",
return_value = "/tmp/llama-server",
),
patch.object(
inference_route.LlamaCppBackend,
"_get_gpu_memory",
return_value = [(0, 8 * 1024**3, 16 * 1024**3)],
),
patch.object(
inference_route,
"_guard_chat_load_against_training",
return_value = None,
) as training_guard,
patch.object(inference_route.asyncio, "to_thread", new = _inline_to_thread),
patch.object(inference_route, "_hf_offline_if_unreachable", nullcontext),
):
with self.assertRaises(HTTPException) as exc_info:
asyncio.run(
inference_route._load_model_impl(
request,
SimpleNamespace(
app = SimpleNamespace(
state = SimpleNamespace(llama_parallel_slots = 1),
),
),
current_subject = "test-user",
)
)
self.assertEqual(exc_info.exception.status_code, 400)
self.assertIn("Vulkan GPU ordinal(s) [99]", exc_info.exception.detail)
training_guard.assert_not_called()
def test_vulkan_ordinals_are_allowed_on_xpu_hosts(self):
import utils.hardware.hardware as hardware_mod
inference_route = _load_route_module(
"inference_route_module_for_xpu_vulkan_test",
"routes/inference.py",
)
config = SimpleNamespace(is_gguf = True)
with (
patch("utils.hardware.get_device", return_value = DeviceType.XPU),
patch.object(
inference_route.LlamaCppBackend,
"_is_vulkan_backend",
return_value = True,
),
patch.object(inference_route, "_classify_diffusion_gguf", return_value = False),
patch.object(hardware_mod, "resolve_requested_gpu_ids", return_value = [0, 1]),
patch.object(
inference_route.LlamaCppBackend,
"_find_llama_server_binary",
return_value = None,
),
):
resolved, uses_vulkan_ordinals = asyncio.run(
inference_route._resolve_gguf_gpu_ids_for_request(config, [1, 0])
)
self.assertEqual(resolved, [0, 1])
self.assertTrue(uses_vulkan_ordinals)
def test_diffusion_gpu_ids_accept_rocm_physical_path(self):
import utils.hardware as hardware_pkg
import utils.hardware.hardware as hardware_mod
inference_route = _load_route_module(
"inference_route_module_for_diffusion_rocm_path_test",
"routes/inference.py",
)
config = SimpleNamespace(is_gguf = True)
fake_backend = SimpleNamespace(
is_vulkan_build = lambda: False,
_backend_lacks_gpu_lib = lambda *a, **k: False,
)
with (
patch.object(hardware_mod, "IS_ROCM", True),
patch.object(hardware_pkg, "get_device", return_value = DeviceType.CUDA),
patch.object(
hardware_mod,
"resolve_requested_gpu_ids",
return_value = [0, 1],
),
patch.object(
inference_route,
"get_llama_cpp_backend",
return_value = fake_backend,
),
patch.object(inference_route.asyncio, "to_thread", new = _inline_to_thread),
):
self.assertEqual(hardware_mod._backend_label(DeviceType.CUDA), "rocm")
resolved, uses_vulkan_ordinals = asyncio.run(
inference_route._resolve_gguf_gpu_ids_for_request(
config,
[1, 0],
diffusion_kind = True,
)
)
self.assertEqual(resolved, [0, 1])
self.assertFalse(uses_vulkan_ordinals)
def test_inference_route_validates_gpu_ids_for_gguf(self):
import utils.hardware.hardware as hardware_mod
import utils.hardware as hardware_pkg
inference_route = _load_route_module(
"inference_route_module_for_gguf_gpu_ids_test2",
"routes/inference.py",
)
request = LoadRequest(model_path = "unsloth/test.gguf", gpu_ids = [0, 1])
model_config = SimpleNamespace(
is_gguf = True,
is_lora = False,
gguf_hf_repo = None,
gguf_file = "/tmp/test.gguf",
gguf_mmproj_file = None,
gguf_variant = None,
identifier = "unsloth/test.gguf",
display_name = "unsloth/test.gguf",
is_vision = False,
is_audio = False,
audio_type = None,
has_audio_input = False,
)
with (
patch.object(
inference_route,
"ModelConfig",
SimpleNamespace(from_identifier = lambda **_kwargs: model_config),
),
patch(
"utils.hardware.resolve_requested_gpu_ids",
side_effect = ValueError("Invalid gpu_ids [0, 1]: rejected by test"),
),
patch.object(
hardware_mod,
"resolve_requested_gpu_ids",
side_effect = ValueError("Invalid gpu_ids [0, 1]: rejected by test"),
),
patch.object(
inference_route,
"_guard_chat_load_against_training",
return_value = None,
),
patch.object(inference_route.asyncio, "to_thread", new = _inline_to_thread),
patch.object(inference_route, "_hf_offline_if_unreachable", nullcontext),
):
with self.assertRaises(HTTPException) as exc_info:
asyncio.run(
inference_route._load_model_impl(
request,
SimpleNamespace(
app = SimpleNamespace(
state = SimpleNamespace(llama_parallel_slots = 1),
),
),
current_subject = "test-user",
)
)
self.assertEqual(exc_info.exception.status_code, 400)
self.assertIn("gpu_ids", exc_info.exception.detail.lower())
def test_inference_route_rejects_gpu_ids_on_cpu_only_llama_build(self):
# A CPU-only llama.cpp build cannot honor a CUDA visibility pin.
import utils.hardware as hardware_pkg
inference_route = _load_route_module(
"inference_route_module_for_cpu_only_gpu_ids_test",
"routes/inference.py",
)
request = LoadRequest(model_path = "unsloth/test.gguf", gpu_ids = [0, 1])
model_config = SimpleNamespace(
is_gguf = True,
is_lora = False,
gguf_hf_repo = None,
gguf_file = "/tmp/test.gguf",
gguf_mmproj_file = None,
gguf_variant = None,
identifier = "unsloth/test.gguf",
display_name = "unsloth/test.gguf",
is_vision = False,
is_audio = False,
audio_type = None,
has_audio_input = False,
)
fake_backend = SimpleNamespace(
is_loaded = False,
model_identifier = None,
is_vulkan_build = lambda: False,
_backend_lacks_gpu_lib = lambda *a, **k: True,
)
with (
patch.object(
inference_route,
"ModelConfig",
SimpleNamespace(from_identifier = lambda **_kwargs: model_config),
),
patch.object(inference_route, "_classify_diffusion_gguf", return_value = False),
patch.object(inference_route, "_guard_chat_load_against_training", return_value = None),
patch.object(inference_route.asyncio, "to_thread", new = _inline_to_thread),
patch.object(inference_route, "_hf_offline_if_unreachable", nullcontext),
patch.object(inference_route, "get_llama_cpp_backend", return_value = fake_backend),
patch.object(hardware_pkg, "get_device", return_value = hardware_pkg.DeviceType.CUDA),
):
with self.assertRaises(HTTPException) as exc_info:
asyncio.run(
inference_route._load_model_impl(
request,
SimpleNamespace(
app = SimpleNamespace(state = SimpleNamespace(llama_parallel_slots = 1)),
),
current_subject = "test-user",
)
)
self.assertEqual(exc_info.exception.status_code, 400)
self.assertIn("cpu-only build", exc_info.exception.detail.lower())
def test_diffusion_gguf_on_vulkan_build_rejects_ordinal_pin(self):
# The GGUF picker supplies Vulkan ordinals, not the CUDA physical IDs the diffusion runner uses.
import utils.hardware as hardware_pkg
inference_route = _load_route_module(
"inference_route_module_for_diffusion_cuda_path_test",
"routes/inference.py",
)
request = LoadRequest(model_path = "unsloth/diffusion.gguf", gpu_ids = [0, 1])
model_config = SimpleNamespace(
is_gguf = True,
is_lora = False,
gguf_hf_repo = None,
gguf_file = "/tmp/diffusion.gguf",
gguf_mmproj_file = None,
gguf_variant = None,
identifier = "unsloth/diffusion.gguf",
display_name = "unsloth/diffusion.gguf",
is_vision = False,
is_audio = False,
audio_type = None,
has_audio_input = False,
)
fake_backend = SimpleNamespace(
is_loaded = False,
model_identifier = None,
is_vulkan_build = lambda: True,
_backend_lacks_gpu_lib = lambda *a, **k: False,
)
with (
patch.object(
inference_route,
"ModelConfig",
SimpleNamespace(from_identifier = lambda **_kwargs: model_config),
),
patch.object(inference_route, "_classify_diffusion_gguf", return_value = True),
patch.object(
_hw_module,
"resolve_requested_gpu_ids",
side_effect = AssertionError("Vulkan ordinal reached the CUDA resolver"),
),
patch.object(inference_route, "_guard_chat_load_against_training", return_value = None),
patch.object(inference_route.asyncio, "to_thread", new = _inline_to_thread),
patch.object(inference_route, "_hf_offline_if_unreachable", nullcontext),
patch.object(inference_route, "get_llama_cpp_backend", return_value = fake_backend),
patch.object(hardware_pkg, "get_device", return_value = hardware_pkg.DeviceType.CUDA),
):
with self.assertRaises(HTTPException) as exc_info:
asyncio.run(
inference_route._load_model_impl(
request,
SimpleNamespace(
app = SimpleNamespace(state = SimpleNamespace(llama_parallel_slots = 1)),
),
current_subject = "test-user",
)
)
self.assertEqual(exc_info.exception.status_code, 400)
self.assertIn("no defined mapping", exc_info.exception.detail)
def test_inference_route_defers_gpu_handoff_until_after_validation(self):
# A doomed chat load (GGUF + gpu_ids -> 400) must NOT reclaim the CHAT arbiter owner first: the handoff is deferred past
# validation, so a resident Images/Video pipeline is never evicted for a load that then errors.
import core.inference.gpu_arbiter as arb
inference_route = _load_route_module(
"inference_route_module_for_handoff_test",
"routes/inference.py",
)
request = LoadRequest(model_path = "unsloth/test.gguf", gpu_ids = [0, 1])
model_config = SimpleNamespace(
is_gguf = True,
is_lora = False,
gguf_hf_repo = None,
gguf_file = "/tmp/test.gguf",
gguf_mmproj_file = None,
gguf_mtp_file = None,
gguf_variant = None,
identifier = "unsloth/test.gguf",
display_name = "unsloth/test.gguf",
is_vision = False,
is_audio = False,
audio_type = None,
has_audio_input = False,
)
acquired = []
# Make [0, 1] invalid on any host (a duplicate id is rejected everywhere): the point is the ORDER, validation before the handoff.
request.gpu_ids = [0, 0]
with (
patch.object(
inference_route,
"ModelConfig",
SimpleNamespace(from_identifier = lambda **_kwargs: model_config),
),
patch.object(inference_route, "_guard_chat_load_against_training", return_value = None),
patch.object(inference_route.asyncio, "to_thread", new = _inline_to_thread),
patch.object(inference_route, "_hf_offline_if_unreachable_for", nullcontext),
# The chat handoff passes a `register` hook (the in-flight marker), so accept it.
patch.object(arb, "acquire_for", lambda owner, register = None: acquired.append(owner)),
):
with self.assertRaises(HTTPException) as exc_info:
asyncio.run(
inference_route._load_model_impl(
request,
SimpleNamespace(
app = SimpleNamespace(
state = SimpleNamespace(llama_parallel_slots = 1),
),
),
current_subject = "test-user",
)
)
self.assertEqual(exc_info.exception.status_code, 400)
self.assertEqual(acquired, []) # no CHAT handoff before the doomed load errored
def test_inference_route_checks_hub_download_conflict_before_the_handoff(self):
# A GGUF the download manager is fetching 409s and loads nothing, so that check must run BEFORE the CHAT handoff:
# afterwards it destroyed the resident Images/Video pipeline for a load that could never start.
import core.inference.gpu_arbiter as arb
import core.inference.llama_cpp as llama_cpp
inference_route = _load_route_module(
"inference_route_module_for_hub_conflict_test",
"routes/inference.py",
)
request = LoadRequest(model_path = "unsloth/Qwen3-4B-GGUF", gguf_variant = "Q4_K_M")
model_config = SimpleNamespace(
is_gguf = True,
is_lora = False,
gguf_hf_repo = "unsloth/Qwen3-4B-GGUF",
gguf_file = None,
gguf_mmproj_file = None,
gguf_variant = "Q4_K_M",
identifier = "unsloth/Qwen3-4B-GGUF",
display_name = "Qwen3-4B",
is_vision = False,
is_audio = False,
audio_type = None,
has_audio_input = False,
)
acquired = []
with (
patch.object(
inference_route,
"ModelConfig",
SimpleNamespace(from_identifier = lambda **_kwargs: model_config),
),
patch.object(inference_route, "_guard_chat_load_against_training", return_value = None),
patch.object(inference_route, "_resolve_inherited_extra_args", lambda *a, **k: None),
patch.object(inference_route.asyncio, "to_thread", new = _inline_to_thread),
patch.object(inference_route, "_hf_offline_if_unreachable_for", nullcontext),
patch.object(llama_cpp, "_hub_download_blocks_gguf_load", lambda *a, **k: True),
patch.object(arb, "acquire_for", lambda *a, **k: acquired.append(a[0])),
):
with self.assertRaises(HTTPException) as exc_info:
asyncio.run(
inference_route._load_model_impl(
request,
SimpleNamespace(
app = SimpleNamespace(
state = SimpleNamespace(llama_parallel_slots = 1),
),
),
current_subject = "test-user",
)
)
self.assertEqual(exc_info.exception.status_code, 409)
self.assertIn("download", exc_info.exception.detail.lower())
self.assertEqual(acquired, []) # nothing evicted for a load that cannot start
def test_inference_route_marks_the_chat_load_under_the_arbiter_lock(self):
# A chat load holds no llama-server process until its GGUF downloaded, so the arbiter is told through acquire_for's
# `register` hook (which runs under the arbiter lock). Passing no register left a competing acquire with nothing to cancel.
import core.inference.gpu_arbiter as arb
import core.inference.llama_cpp as llama_cpp
inference_route = _load_route_module(
"inference_route_module_for_chat_marker_test",
"routes/inference.py",
)
request = LoadRequest(model_path = "unsloth/Qwen3-4B-GGUF", gguf_variant = "Q4_K_M")
model_config = SimpleNamespace(
is_gguf = True,
is_lora = False,
gguf_hf_repo = "unsloth/Qwen3-4B-GGUF",
gguf_file = None,
gguf_mmproj_file = None,
gguf_variant = "Q4_K_M",
identifier = "unsloth/Qwen3-4B-GGUF",
display_name = "Qwen3-4B",
is_vision = False,
is_audio = False,
audio_type = None,
has_audio_input = False,
)
marked = []
def _acquire(owner, register = None):
# Under the arbiter lock the evictor must already be able to see this load.
if register is not None:
register()
marked.append(llama_cpp.chat_load_active())
raise RuntimeError("stop the load here")
with (
patch.object(
inference_route,
"ModelConfig",
SimpleNamespace(from_identifier = lambda **_kwargs: model_config),
),
patch.object(inference_route, "_guard_chat_load_against_training", return_value = None),
patch.object(inference_route, "_resolve_inherited_extra_args", lambda *a, **k: None),
patch.object(inference_route.asyncio, "to_thread", new = _inline_to_thread),
patch.object(inference_route, "_hf_offline_if_unreachable_for", nullcontext),
patch.object(llama_cpp, "_hub_download_blocks_gguf_load", lambda *a, **k: False),
patch.object(arb, "acquire_for", _acquire),
):
with self.assertRaises(HTTPException):
asyncio.run(
inference_route._load_model_impl(
request,
SimpleNamespace(
app = SimpleNamespace(
state = SimpleNamespace(llama_parallel_slots = 1),
),
),
current_subject = "test-user",
)
)
self.assertEqual(marked, [True])
# The marker is scoped to the request: it must not outlive the failed load.
self.assertFalse(llama_cpp.chat_load_active())
def test_training_route_returns_400_for_invalid_gpu_ids(self):
training_route = _load_route_module(
"training_route_module_for_test",
"routes/training.py",
)
request = TrainingStartRequest(
model_name = "unsloth/test",
training_type = "LoRA/QLoRA",
format_type = "alpaca",
gpu_ids = [99],
)
class DummyBackend:
current_job_id = None
def is_training_active(self):
return False
def start_training(self, **kwargs):
raise ValueError("Invalid gpu_ids [99]")
with (
patch.object(training_route, "get_training_backend", return_value = DummyBackend()),
patch.object(
training_route,
"_remote_untrainable_model_format",
return_value = None,
),
patch.object(training_route.asyncio, "to_thread", new = _inline_to_thread),
patch(
"routes.training_vram.summarize_resident_chat",
return_value = {"any": False, "hf": None, "gguf": None},
),
patch(
"core.export.get_export_backend",
return_value = SimpleNamespace(current_checkpoint = None),
),
):
with self.assertRaises(HTTPException) as exc_info:
asyncio.run(training_route.start_training(request, current_subject = "test-user"))
self.assertEqual(exc_info.exception.status_code, 400)
self.assertIn("gpu_ids [99]", exc_info.exception.detail)
def test_training_route_returns_400_for_uuid_parent_visibility_gpu_ids(self):
training_route = _load_route_module(
"training_route_module_for_uuid_parent_visibility_test",
"routes/training.py",
)
request = TrainingStartRequest(
model_name = "unsloth/test",
training_type = "LoRA/QLoRA",
format_type = "alpaca",
gpu_ids = [1],
)
class DummyBackend:
current_job_id = None
def is_training_active(self):
return False
def start_training(self, **kwargs):
raise ValueError(
"Invalid gpu_ids [1]: explicit physical GPU IDs are unsupported when CUDA_VISIBLE_DEVICES uses UUID/MIG entries"
)
with (
patch.object(training_route, "get_training_backend", return_value = DummyBackend()),
patch.object(
training_route,
"_remote_untrainable_model_format",
return_value = None,
),
patch.object(training_route.asyncio, "to_thread", new = _inline_to_thread),
patch(
"routes.training_vram.summarize_resident_chat",
return_value = {"any": False, "hf": None, "gguf": None},
),
patch(
"core.export.get_export_backend",
return_value = SimpleNamespace(current_checkpoint = None),
),
):
with self.assertRaises(HTTPException) as exc_info:
asyncio.run(training_route.start_training(request, current_subject = "test-user"))
self.assertEqual(exc_info.exception.status_code, 400)
self.assertIn("UUID/MIG", exc_info.exception.detail)
def test_inference_route_returns_400_for_invalid_gpu_ids(self):
inference_route = _load_route_module(
"inference_route_module_for_test",
"routes/inference.py",
)
request = LoadRequest(model_path = "unsloth/test", gpu_ids = [99])
model_config = SimpleNamespace(
is_gguf = False,
is_lora = False,
path = None,
identifier = "unsloth/test",
display_name = "unsloth/test",
is_vision = False,
is_audio = False,
audio_type = None,
has_audio_input = False,
)
class DummyInferenceBackend:
active_model_name = None
models = {}
def load_model(self, **kwargs):
raise ValueError("Invalid gpu_ids [99]")
with (
patch.object(
inference_route,
"ModelConfig",
SimpleNamespace(from_identifier = lambda **_kwargs: model_config),
),
patch.object(
inference_route,
"get_inference_backend",
return_value = DummyInferenceBackend(),
),
patch.object(
inference_route,
"get_llama_cpp_backend",
return_value = SimpleNamespace(is_loaded = False),
),
patch.object(
inference_route,
"_guard_chat_load_against_training",
return_value = None,
),
patch.object(inference_route.asyncio, "to_thread", new = _inline_to_thread),
patch.object(inference_route, "_hf_offline_if_unreachable", nullcontext),
patch(
"core.export.get_export_backend",
return_value = SimpleNamespace(current_checkpoint = None),
),
):
with self.assertRaises(HTTPException) as exc_info:
asyncio.run(
inference_route._load_model_impl(
request,
SimpleNamespace(
app = SimpleNamespace(
state = SimpleNamespace(llama_parallel_slots = 1),
),
),
current_subject = "test-user",
)
)
self.assertEqual(exc_info.exception.status_code, 400)
self.assertIn("gpu_ids [99]", exc_info.exception.detail)
def test_inference_route_returns_400_for_uuid_parent_visibility_gpu_ids(self):
inference_route = _load_route_module(
"inference_route_module_for_uuid_parent_visibility_test",
"routes/inference.py",
)
request = LoadRequest(model_path = "unsloth/test", gpu_ids = [1])
model_config = SimpleNamespace(
is_gguf = False,
is_lora = False,
path = None,
identifier = "unsloth/test",
display_name = "unsloth/test",
is_vision = False,
is_audio = False,
audio_type = None,
has_audio_input = False,
)
class DummyInferenceBackend:
active_model_name = None
models = {}
def load_model(self, **kwargs):
raise ValueError(
"Invalid gpu_ids [1]: explicit physical GPU IDs are unsupported when CUDA_VISIBLE_DEVICES uses UUID/MIG entries"
)
with (
patch.object(
inference_route,
"ModelConfig",
SimpleNamespace(from_identifier = lambda **_kwargs: model_config),
),
patch.object(
inference_route,
"get_inference_backend",
return_value = DummyInferenceBackend(),
),
patch.object(
inference_route,
"get_llama_cpp_backend",
return_value = SimpleNamespace(is_loaded = False),
),
patch.object(
inference_route,
"_guard_chat_load_against_training",
return_value = None,
),
patch.object(inference_route.asyncio, "to_thread", new = _inline_to_thread),
patch.object(inference_route, "_hf_offline_if_unreachable", nullcontext),
patch(
"core.export.get_export_backend",
return_value = SimpleNamespace(current_checkpoint = None),
),
):
with self.assertRaises(HTTPException) as exc_info:
asyncio.run(
inference_route._load_model_impl(
request,
SimpleNamespace(
app = SimpleNamespace(
state = SimpleNamespace(llama_parallel_slots = 1),
),
),
current_subject = "test-user",
)
)
self.assertEqual(exc_info.exception.status_code, 400)
self.assertIn("UUID/MIG", exc_info.exception.detail)
class TestRaiseIfOffloaded(unittest.TestCase):
def test_no_offload_is_noop(self):
from utils.hardware import raise_if_offloaded
model = SimpleNamespace(hf_device_map = {"model.embed_tokens": 0, "lm_head": 1})
raise_if_offloaded(model, "balanced", "Test")
def test_cpu_offload_raises(self):
from utils.hardware import raise_if_offloaded
model = SimpleNamespace(hf_device_map = {"model.layers.0": 0, "model.layers.1": "cpu"})
with self.assertRaisesRegex(ValueError, "offloaded"):
raise_if_offloaded(model, "balanced", "Test")
def test_no_device_map_attr_is_noop(self):
from utils.hardware import raise_if_offloaded
raise_if_offloaded(SimpleNamespace(), "sequential", "Test")
class TestMinGpuVram(unittest.TestCase):
def test_min_gpu_vram_decreases_with_more_gpus(self):
from utils.hardware.vram_estimation import (
ModelArchConfig,
TrainingVramConfig,
estimate_training_vram,
)
arch = ModelArchConfig(
hidden_size = 4096,
num_hidden_layers = 32,
num_attention_heads = 32,
num_key_value_heads = 8,
intermediate_size = 14336,
vocab_size = 128256,
tie_word_embeddings = False,
)
config = TrainingVramConfig(
training_method = "qlora",
load_in_4bit = True,
)
breakdown = estimate_training_vram(arch, config)
v1 = breakdown.min_gpu_vram(1)
v2 = breakdown.min_gpu_vram(2)
v4 = breakdown.min_gpu_vram(4)
self.assertGreater(v1, v2)
self.assertGreater(v2, v4)
self.assertGreater(v4, 0)
def test_total_equals_min_gpu_vram_1(self):
from utils.hardware.vram_estimation import (
ModelArchConfig,
TrainingVramConfig,
estimate_training_vram,
)
arch = ModelArchConfig(
hidden_size = 4096,
num_hidden_layers = 32,
num_attention_heads = 32,
num_key_value_heads = 8,
intermediate_size = 14336,
vocab_size = 128256,
tie_word_embeddings = False,
)
config = TrainingVramConfig(
training_method = "qlora",
load_in_4bit = True,
)
breakdown = estimate_training_vram(arch, config)
self.assertEqual(breakdown.total, breakdown.min_gpu_vram(1))
class TestPerGpuFitGuardAllCounts(unittest.TestCase):
def test_training_estimate_resolves_attention_without_raising(self):
with (
patch("utils.hardware.hardware.get_device", return_value = DeviceType.CUDA),
patch(
"utils.hardware.hardware.estimate_fp16_model_size_bytes",
return_value = (8 * (1024**3), "config"),
),
patch(
"utils.hardware.hardware._resolve_model_identifier_for_gpu_estimate",
return_value = "unsloth/test",
),
patch(
"utils.hardware.hardware._load_config_for_gpu_estimate",
return_value = SimpleNamespace(
hidden_size = 4096,
num_hidden_layers = 32,
num_attention_heads = 32,
num_key_value_heads = 8,
intermediate_size = 14336,
vocab_size = 128256,
tie_word_embeddings = False,
),
),
patch(
"utils.hardware.hardware._determine_attention_impl_for_gpu_estimate",
return_value = "eager",
),
patch("utils.hardware.hardware.get_visible_gpu_count", return_value = 1),
):
_, metadata = estimate_required_model_memory_gb(
"unsloth/test",
training_type = "LoRA/QLoRA",
load_in_4bit = True,
)
self.assertEqual(metadata.get("estimation_mode"), "detailed")
self.assertEqual(metadata.get("attention_implementation"), "eager")
def test_training_estimate_falls_back_when_attention_resolution_fails(self):
with (
patch("utils.hardware.hardware.get_device", return_value = DeviceType.CUDA),
patch(
"utils.hardware.hardware.estimate_fp16_model_size_bytes",
return_value = (8 * (1024**3), "config"),
),
patch(
"utils.hardware.hardware._resolve_model_identifier_for_gpu_estimate",
return_value = "unsloth/test",
),
patch(
"utils.hardware.hardware._load_config_for_gpu_estimate",
return_value = SimpleNamespace(
hidden_size = 4096,
num_hidden_layers = 32,
num_attention_heads = 32,
num_key_value_heads = 8,
intermediate_size = 14336,
vocab_size = 128256,
tie_word_embeddings = False,
),
),
patch(
"utils.hardware.hardware._determine_attention_impl_for_gpu_estimate",
side_effect = RuntimeError("attention unavailable"),
),
patch("utils.hardware.hardware.get_visible_gpu_count", return_value = 1),
):
_, metadata = estimate_required_model_memory_gb(
"unsloth/test",
training_type = "LoRA/QLoRA",
load_in_4bit = True,
)
self.assertEqual(metadata.get("estimation_mode"), "detailed")
self.assertEqual(
metadata.get("attention_implementation"),
"eager",
)
def test_attention_resolver_does_not_mutate_loaded_config(self):
from utils.hardware import hardware as hardware_module
config = SimpleNamespace(
hidden_size = 1024,
num_hidden_layers = 2,
num_attention_heads = 8,
num_key_value_heads = 8,
intermediate_size = 2048,
vocab_size = 1024,
tie_word_embeddings = True,
)
def _stub_resolver(model_class, cfg):
cfg._attn_implementation = "eager"
return "eager"
with patch.dict(sys.modules, _fake_unsloth_attention_modules(_stub_resolver)):
hardware_module._determine_attention_impl_for_gpu_estimate(config)
self.assertFalse(hasattr(config, "_attn_implementation"))
def test_attention_resolver_handles_missing_model_mapping(self):
from utils.hardware import hardware as hardware_module
config = SimpleNamespace(
hidden_size = 1024,
num_hidden_layers = 2,
num_attention_heads = 8,
num_key_value_heads = 8,
intermediate_size = 2048,
vocab_size = 1024,
tie_word_embeddings = True,
)
captured = {}
def _stub_resolver(model_class, cfg):
captured["model_class"] = model_class
return "eager"
from transformers import AutoModel, AutoModelForCausalLM
with (
patch.object(AutoModelForCausalLM, "_model_mapping", new = None),
patch.object(AutoModel, "_model_mapping", new = None),
patch.dict(sys.modules, _fake_unsloth_attention_modules(_stub_resolver)),
):
result = hardware_module._determine_attention_impl_for_gpu_estimate(config)
self.assertEqual(result, "eager")
self.assertIsNone(captured["model_class"])
def test_attention_resolver_does_not_mutate_nested_text_config(self):
from utils.hardware import hardware as hardware_module
text_config = SimpleNamespace(
hidden_size = 1024,
num_hidden_layers = 2,
num_attention_heads = 8,
num_key_value_heads = 8,
intermediate_size = 2048,
vocab_size = 1024,
tie_word_embeddings = True,
)
config = SimpleNamespace(
hidden_size = 1024,
num_hidden_layers = 2,
num_attention_heads = 8,
num_key_value_heads = 8,
intermediate_size = 2048,
vocab_size = 1024,
tie_word_embeddings = True,
text_config = text_config,
)
def _stub_resolver(model_class, cfg):
cfg._attn_implementation = "eager"
inner = getattr(cfg, "text_config", None)
if inner is not None:
inner._attn_implementation = "eager"
return "eager"
with patch.dict(sys.modules, _fake_unsloth_attention_modules(_stub_resolver)):
hardware_module._determine_attention_impl_for_gpu_estimate(config)
self.assertFalse(hasattr(config, "_attn_implementation"))
self.assertFalse(hasattr(text_config, "_attn_implementation"))
def test_min_per_gpu_generated_for_all_visible_counts(self):
with (
patch("utils.hardware.hardware.get_device", return_value = DeviceType.CUDA),
patch(
"utils.hardware.hardware.estimate_fp16_model_size_bytes",
return_value = (8 * (1024**3), "config"),
),
patch(
"utils.hardware.hardware._resolve_model_identifier_for_gpu_estimate",
return_value = "unsloth/test",
),
patch(
"utils.hardware.hardware._load_config_for_gpu_estimate",
return_value = SimpleNamespace(
hidden_size = 4096,
num_hidden_layers = 32,
num_attention_heads = 32,
num_key_value_heads = 8,
intermediate_size = 14336,
vocab_size = 128256,
tie_word_embeddings = False,
),
),
patch("utils.hardware.hardware.get_visible_gpu_count", return_value = 6),
):
_, metadata = estimate_required_model_memory_gb(
"unsloth/test",
training_type = "LoRA/QLoRA",
load_in_4bit = True,
)
self.assertEqual(metadata.get("estimation_mode"), "detailed")
breakdown = metadata["vram_breakdown"]
for n in range(1, 7):
self.assertIn(f"min_per_gpu_{n}", breakdown)
class TestAutoSelectWithNoneRequired(_GpuCacheResetMixin, unittest.TestCase):
def test_auto_select_falls_back_when_estimate_unavailable(self):
with (
patch("utils.hardware.hardware.get_device", return_value = DeviceType.CUDA),
patch(
"utils.hardware.hardware.estimate_required_model_memory_gb",
return_value = (None, {"model_size_source": "unavailable"}),
),
patch(
"utils.hardware.hardware._get_parent_visible_gpu_spec",
return_value = {
"raw": "0,1",
"numeric_ids": [0, 1],
"supports_explicit_gpu_ids": True,
},
),
patch(
"utils.hardware.hardware.get_parent_visible_gpu_ids",
return_value = [0, 1],
),
):
selected, metadata = auto_select_gpu_ids("unsloth/test")
self.assertEqual(selected, [0, 1])
self.assertEqual(metadata["selection_mode"], "fallback_all")
class TestXpuSelection(_GpuCacheResetMixin, unittest.TestCase):
def test_auto_select_supports_xpu(self):
with (
patch("utils.hardware.hardware.get_device", return_value = DeviceType.XPU),
patch(
"utils.hardware.hardware.estimate_required_model_memory_gb",
return_value = (1.0, {}),
),
patch(
"utils.hardware.hardware.get_visible_gpu_utilization",
return_value = {
"devices": [
{"index": 0, "vram_total_gb": 8, "vram_used_gb": 1},
]
},
),
patch(
"utils.hardware.hardware._get_parent_visible_gpu_spec",
return_value = {
"raw": None,
"numeric_ids": [0],
"supports_explicit_gpu_ids": True,
},
),
patch(
"utils.hardware.hardware.get_parent_visible_gpu_ids",
return_value = [0],
),
):
selected, metadata = auto_select_gpu_ids("unsloth/test")
self.assertEqual(selected, [0])
self.assertEqual(metadata["selection_mode"], "auto")
def test_prepare_gpu_selection_accepts_explicit_ids_on_xpu(self):
with (
patch("utils.hardware.hardware.get_device", return_value = DeviceType.XPU),
patch(
"utils.hardware.hardware._get_parent_visible_gpu_spec",
return_value = {
"raw": "0",
"numeric_ids": [0],
"supports_explicit_gpu_ids": True,
},
),
patch(
"utils.hardware.hardware.get_parent_visible_gpu_ids",
return_value = [0],
),
patch("utils.hardware.hardware.get_physical_gpu_count", return_value = 1),
):
selected, metadata = prepare_gpu_selection([0], model_name = "unsloth/test")
self.assertEqual(selected, [0])
self.assertEqual(metadata["selection_mode"], "explicit")
class TestEstimateFp16ModelSizeBytesPrefersLocalWeights(unittest.TestCase):
def _run(
self,
model_path,
*,
config_bytes,
local_bytes,
safetensors_params = None,
config = object(),
):
from utils.hardware import hardware as hardware_module
with (
patch.object(
hardware_module,
"_resolve_model_identifier_for_gpu_estimate",
return_value = model_path,
),
patch.object(
hardware_module,
"_get_hf_safetensors_total_params",
return_value = safetensors_params,
),
patch.object(
hardware_module,
"_load_config_for_gpu_estimate",
return_value = config,
),
patch.object(
hardware_module,
"_estimate_fp16_model_size_bytes_from_config",
return_value = config_bytes,
),
patch.object(
hardware_module,
"_get_local_weight_size_bytes",
return_value = local_bytes,
),
):
return hardware_module.estimate_fp16_model_size_bytes(model_path)
def test_local_weight_bytes_preferred_when_larger_than_config(self):
bytes_, src = self._run(
"/local/vlm",
config_bytes = 2 * (1 << 30),
local_bytes = 20 * (1 << 30),
)
self.assertEqual(bytes_, 20 * (1 << 30))
self.assertEqual(src, "weight_bytes")
def test_config_bytes_preferred_when_larger_than_local(self):
bytes_, src = self._run(
"/local/text-only",
config_bytes = 20 * (1 << 30),
local_bytes = 2 * (1 << 30),
)
self.assertEqual(bytes_, 20 * (1 << 30))
self.assertEqual(src, "config")
def test_config_bytes_returned_when_no_local_weights(self):
bytes_, src = self._run(
"/local/no-weights",
config_bytes = 5 * (1 << 30),
local_bytes = None,
)
self.assertEqual(bytes_, 5 * (1 << 30))
self.assertEqual(src, "config")
def test_local_bytes_returned_when_config_resolution_fails(self):
bytes_, src = self._run(
"/local/no-config",
config_bytes = None,
local_bytes = 7 * (1 << 30),
config = None,
)
self.assertEqual(bytes_, 7 * (1 << 30))
self.assertEqual(src, "weight_bytes")
def test_equal_local_and_config_keeps_config_label(self):
# Tie-breaker is "local must be strictly larger", so an exact match keeps the config-derived path.
same = 8 * (1 << 30)
bytes_, src = self._run(
"/local/equal",
config_bytes = same,
local_bytes = same,
)
self.assertEqual(bytes_, same)
self.assertEqual(src, "config")
def test_remote_safetensors_path_unaffected_by_local_weights(self):
from utils.hardware import hardware as hardware_module
with (
patch.object(
hardware_module,
"_resolve_model_identifier_for_gpu_estimate",
return_value = "owner/repo",
),
patch.object(
hardware_module,
"_get_hf_safetensors_total_params",
return_value = 1_000_000_000,
),
patch.object(
hardware_module,
"_load_config_for_gpu_estimate",
) as mock_load,
patch.object(
hardware_module,
"_get_local_weight_size_bytes",
) as mock_local,
):
bytes_, src = hardware_module.estimate_fp16_model_size_bytes("owner/repo")
self.assertEqual(bytes_, 2 * 1_000_000_000)
self.assertEqual(src, "safetensors")
mock_load.assert_not_called()
mock_local.assert_not_called()