unsloth/studio/backend/tests/test_linux_external_media_paths.py
Michael Han 0eb6b6c931
Studio: say when a scan folder cannot be read instead of showing no models (#9053)
* Studio: say when a scan folder cannot be read

A folder Unsloth is denied looks exactly like an empty one: the scan
catches the OSError, logs it, and moves on, so the model list is empty
with no reason given.

Add-time validation now opens the directory instead of trusting
os.access, which reads mode bits only and passes on folders macOS TCC or
a Windows ACL still refuses. The check runs after the denylist rules so a
denied path is never opened.

The scan keeps the error it already caught, and both scan-folders
endpoints return it as a per-folder status the dialog shows with the
setting that fixes it.

No new work on the healthy path: 0.06us per folder per scan, and the
folder list is a dict lookup with no syscalls.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: record scan folder status from the Hub inventory scan too

The folders dialog reads /api/hub/scan-folders, but the scan behind it is
the Hub inventory, not collect_local_models, so nothing was ever recorded
for it and every row stayed "ok".

Its custom-folder loop now records the same way. That alone is not enough:
the Hub's _scan_models_dir catches the OSError itself and returns an empty
list, so no exception reaches the loop. So an empty result is now the
trigger. One opendir says whether the folder is empty, gone, or refused,
and it runs only for a folder that returned no models.

A folder that found models still costs nothing, with a test that fails if
it ever touches the filesystem.

* Studio: catch denied model subdirs, and recheck when the dialog reopens

Two gaps in the folder status.

A root can list fine while every model under it is denied, on a NAS mount
or a drive owned by another user. The scanners skip an unreadable child
silently, so that arrives as the same empty list as an empty folder and
was reported as ok. The probe now also opens subdirectories, stopping at
the first refusal and capped at 64, so the denied-everything case costs
one extra open.

The row tells the user to fix permissions and reopen the dialog, but
nothing rechecked between inventory scans, so the warning stayed up after
access was restored. Listing the folders now rechecks the folders marked
bad, and only those. A healthy folder is not in the registry, so the list
still opens nothing.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: probe two levels down, and keep the tests collectable on Windows

The child probe descended one level, so <root>/<publisher>/<model> with the
model denied still reported ok: both levels above it list fine and the
scanners return nothing. It now walks two levels, depth first, on one
shared budget of 64 opens. Depth first means a denied mount is found in
three opens instead of after every publisher, and the budget bounds the
cost whatever the shape of the tree.

os.geteuid does not exist on Windows and a skipif condition is evaluated
at import, so collecting the test file there raised AttributeError before
the os.name check could skip anything. Resolved once into a shared marker,
with a test that runs the module body with geteuid removed.

* Studio: flag a folder with one denied model, and show status in the picker

A folder holding one readable model and one denied model returned the
readable one, so the scan looked successful and the denied model was
silently absent. The probe now runs whether or not models were found, and
reports "partial" in that case so the copy does not contradict the rows on
screen by claiming the folder cannot be read.

That is a real cost change on the healthy path, so it is measured rather
than claimed: 104us for a folder with 8 model dirs, 0.77ms for one with
300, capped by the same 64-open budget. The end-to-end scan stays inside
run-to-run variation, and the folder list still opens nothing unless a
folder is already marked bad. The old zero-syscall test is replaced by the
bound, which is now the guarantee that matters.

The inline model selector manages the same folders and rendered only the
path, so it shows the status too.

* Studio: stop treating an exhausted probe budget as healthy

Three fixes.

A denied directory past the open budget was reported as ok. Running out of
budget means the tail was never looked at, which is not the same as finding
it healthy, so it now returns an internal "unknown" that is never recorded
and never sent to the UI.

That alone would strand a wide folder in a warning it could never clear, so
the registry now remembers which directory refused. A recheck opens that one
directory, which settles a fixed folder in a single open no matter how wide
the folder is or where the denial sat.

A folder recorded as partial kept that status after being deleted, because
the recheck preserved partial for every non-ok probe. It now only holds
partial against a permission result, so missing and unreadable replace it.

One permission test was missing the marker that skips it as root and on
Windows, where chmod 000 does not deny.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Probe scan folders off the event loop, and stop a vanished model or a Windows device error condemning the folder

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Probe the commit directory inside an HF snapshots folder, and stop calling a shut root partial

Two ways the dialog told the user the wrong thing. The cache layout is
models--org--name/snapshots/<commit>/, three levels under a registered
root, so a probe that stops at two never opens the one directory the
weights live in: a denied commit dir made the model vanish from the list
while the folder still reported ok, which is the silent empty case this
PR exists to remove. The extra level is bought only for a directory
named snapshots, since a blanket third level would spend the open budget
descending into diffusers component directories.

And when the root itself becomes denied, the partial branch restored
partial regardless, so a folder none of which can be read kept saying
some models in it could not be read, sending the user hunting for one
bad model. The cause the probe already returns is the discriminator: it
is the root path itself for the root's own refusal and a nested entry
path otherwise.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
Co-authored-by: Daniel Han <danielhanchen@gmail.com>
2026-08-19 06:10:58 -07:00

299 lines
10 KiB
Python

# SPDX-License-Identifier: AGPL-3.0-only
# Copyright 2026-present the Unsloth AI Inc. team. All rights reserved. See /studio/LICENSE.AGPL-3.0
from __future__ import annotations
import ast
import os
import sys
from pathlib import Path
from types import SimpleNamespace
from typing import Optional
import pytest
from hub.storage import scan_folders
from storage import studio_db
from utils.paths import external_media
_BACKEND_ROOT = Path(__file__).resolve().parent.parent
class _ExistingScanFolderConn:
def __init__(self):
self.params = ()
def execute(
self,
_sql,
params = (),
):
self.params = params
return self
def fetchone(self):
return {"id": 1, "path": self.params[0], "created_at": "fake"}
def commit(self):
pass
def close(self):
pass
class _HTTPException(Exception):
def __init__(self, status_code: int, detail: str):
super().__init__(detail)
self.status_code = status_code
self.detail = detail
def _stub_linux_path_checks(monkeypatch, module):
monkeypatch.setattr(module.platform, "system", lambda: "Linux")
monkeypatch.setattr(module.os.path, "realpath", os.path.normpath)
monkeypatch.setattr(module.os.path, "expanduser", lambda p: p)
monkeypatch.setattr(module.os.path, "exists", lambda _p: True)
monkeypatch.setattr(module.os.path, "isdir", lambda _p: True)
monkeypatch.setattr(module.os, "access", lambda _p, _mode: True)
# The mount does not exist on the test host, so the readability probe that
# opens the directory has to be part of the same pretence.
monkeypatch.setattr(module, "is_readable_dir", lambda _p: True)
def _stub_hub_scan_folder_db(monkeypatch):
monkeypatch.setattr(scan_folders, "_ensure_schema", lambda _conn: None)
monkeypatch.setattr(scan_folders, "get_connection", _ExistingScanFolderConn)
def _stub_legacy_scan_folder_db(monkeypatch):
monkeypatch.setattr(studio_db, "get_connection", _ExistingScanFolderConn)
def test_linux_run_media_policy_accepts_mounted_volume_descendants(monkeypatch):
monkeypatch.setattr(external_media.platform, "system", lambda: "Linux")
assert external_media.is_linux_run_media_path("/run/media/dspofu/nvmeB")
assert external_media.is_linux_run_media_path("/run/media/dspofu/nvmeB/modelsAI/gguf/qwen3.6")
@pytest.mark.parametrize(
"path",
[
"/run",
"/run/media",
"/run/media/dspofu",
"/run/user/1000/models",
"/run/systemd/private",
"/run/not-media/dspofu/nvmeB",
],
)
def test_linux_run_media_policy_rejects_unrelated_run_paths(monkeypatch, path):
monkeypatch.setattr(external_media.platform, "system", lambda: "Linux")
assert not external_media.is_linux_run_media_path(path)
def test_linux_run_media_mount_roots_lists_readable_volume_roots(monkeypatch, tmp_path):
base = tmp_path / "run" / "media"
mount = base / "dspofu" / "nvmeB"
sensitive_mount = base / "dspofu" / ".ssh"
sensitive_aws_mount = base / "dspofu" / ".aws"
other_user_mount = base / "other" / "backup"
incomplete = base / "dspofu-only"
mount.mkdir(parents = True)
sensitive_mount.mkdir()
sensitive_aws_mount.mkdir()
other_user_mount.mkdir(parents = True)
incomplete.mkdir()
monkeypatch.setattr(external_media.platform, "system", lambda: "Linux")
roots = external_media.linux_run_media_mount_roots(base, user = "dspofu")
assert roots == [mount.resolve()]
def test_linux_run_media_mount_roots_skips_sensitive_resolved_volume_name(monkeypatch, tmp_path):
base = tmp_path / "run" / "media"
normal_mount = base / "dspofu" / "nvmeB"
sensitive_target = base / "dspofu" / ".config"
normal_mount.mkdir(parents = True)
sensitive_target.mkdir()
alias = base / "dspofu" / "config-alias"
alias.symlink_to(sensitive_target, target_is_directory = True)
monkeypatch.setattr(external_media.platform, "system", lambda: "Linux")
roots = external_media.linux_run_media_mount_roots(base, user = "dspofu")
assert roots == [normal_mount.resolve()]
def test_linux_run_media_mount_roots_skips_sensitive_resolved_descendant(monkeypatch, tmp_path):
base = tmp_path / "run" / "media"
normal_mount = base / "dspofu" / "nvmeB"
sensitive_descendant = normal_mount / ".ssh" / "models"
sensitive_descendant.mkdir(parents = True)
alias = base / "dspofu" / "models-alias"
alias.symlink_to(sensitive_descendant, target_is_directory = True)
monkeypatch.setattr(external_media.platform, "system", lambda: "Linux")
roots = external_media.linux_run_media_mount_roots(base, user = "dspofu")
assert roots == [normal_mount.resolve()]
def test_hub_scan_folder_accepts_linux_run_media_mount(monkeypatch):
_stub_linux_path_checks(monkeypatch, scan_folders)
monkeypatch.setattr(external_media.platform, "system", lambda: "Linux")
_stub_hub_scan_folder_db(monkeypatch)
target = "/run/media/dspofu/nvmeB/modelsAI/gguf/qwen3.6"
row = scan_folders.add_scan_folder(target)
assert row["path"] == target
@pytest.mark.parametrize(
"target",
[
"/run",
"/run/media",
"/run/media/dspofu",
"/run/user/1000/models",
"/run/systemd/private",
"/run/not-media/dspofu/nvmeB",
],
)
def test_hub_scan_folder_keeps_unrelated_run_paths_blocked(monkeypatch, target):
_stub_linux_path_checks(monkeypatch, scan_folders)
monkeypatch.setattr(external_media.platform, "system", lambda: "Linux")
_stub_hub_scan_folder_db(monkeypatch)
with pytest.raises(ValueError, match = "Path under /run is not allowed"):
scan_folders.add_scan_folder(target)
def test_hub_scan_folder_keeps_sensitive_dirs_blocked_under_run_media(monkeypatch):
_stub_linux_path_checks(monkeypatch, scan_folders)
monkeypatch.setattr(external_media.platform, "system", lambda: "Linux")
_stub_hub_scan_folder_db(monkeypatch)
with pytest.raises(ValueError, match = "Credential or configuration"):
scan_folders.add_scan_folder("/run/media/dspofu/nvmeB/.ssh/models")
def test_legacy_scan_folder_accepts_linux_run_media_mount(monkeypatch):
_stub_linux_path_checks(monkeypatch, studio_db)
monkeypatch.setattr(external_media.platform, "system", lambda: "Linux")
_stub_legacy_scan_folder_db(monkeypatch)
target = "/run/media/dspofu/nvmeB/modelsAI/gguf/qwen3.6"
row = studio_db.add_scan_folder(target)
assert row["path"] == target
@pytest.mark.parametrize(
"target",
[
"/run",
"/run/media",
"/run/media/dspofu",
"/run/user/1000/models",
"/run/systemd/private",
"/run/not-media/dspofu/nvmeB",
],
)
def test_legacy_scan_folder_keeps_unrelated_run_paths_blocked(monkeypatch, target):
_stub_linux_path_checks(monkeypatch, studio_db)
monkeypatch.setattr(external_media.platform, "system", lambda: "Linux")
_stub_legacy_scan_folder_db(monkeypatch)
with pytest.raises(ValueError, match = "Path under /run is not allowed"):
studio_db.add_scan_folder(target)
def test_legacy_scan_folder_keeps_sensitive_dirs_blocked_under_run_media(monkeypatch):
_stub_linux_path_checks(monkeypatch, studio_db)
monkeypatch.setattr(external_media.platform, "system", lambda: "Linux")
_stub_legacy_scan_folder_db(monkeypatch)
with pytest.raises(ValueError, match = "Credential or configuration"):
studio_db.add_scan_folder("/run/media/dspofu/nvmeB/.aws/models")
def test_legacy_browse_allowlist_includes_linux_run_media_mounts(monkeypatch, tmp_path):
tree = ast.parse((_BACKEND_ROOT / "routes" / "models.py").read_text(encoding = "utf-8"))
function_names = {
"_build_browse_allowlist",
"_browse_relative_parts",
"_is_path_inside_allowlist",
"_match_browse_child",
"_normalize_browse_request_path",
"_resolve_browse_target",
}
functions = [
node
for node in tree.body
if isinstance(node, ast.FunctionDef) and node.name in function_names
]
module = ast.Module(body = functions, type_ignores = [])
ast.fix_missing_locations(module)
home = tmp_path / "home"
media_root = tmp_path / "run" / "media" / "dspofu" / "nvmeB"
model_dir = media_root / "modelsAI" / "gguf" / "qwen3.6"
home.mkdir()
model_dir.mkdir(parents = True)
(media_root / ".ssh").mkdir()
fake_paths = SimpleNamespace(
hf_default_cache_dir = lambda: tmp_path / "missing-default-hf",
legacy_hf_cache_dir = lambda: tmp_path / "missing-legacy-hf",
well_known_model_dirs = lambda: [],
studio_root = lambda: tmp_path / "missing-studio",
outputs_root = lambda: tmp_path / "missing-outputs",
exports_root = lambda: tmp_path / "missing-exports",
)
fake_external_media = SimpleNamespace(
linux_run_media_mount_roots = lambda: [media_root],
macos_volume_roots = lambda: [],
windows_drive_roots = lambda: [],
)
fake_paths.external_media = fake_external_media
fake_studio_db = SimpleNamespace(
list_scan_folders = lambda: [],
contains_sensitive_path_component = studio_db.contains_sensitive_path_component,
# The media root is a legitimate mount, not denied; the .ssh 403 below
# comes from the credential check. A False stub keeps this OS-independent
# (on macOS tmp_path lives under the denied /private/var).
is_denied_system_path = lambda _p: False,
)
monkeypatch.setitem(sys.modules, "utils.paths", fake_paths)
monkeypatch.setitem(sys.modules, "utils.paths.external_media", fake_external_media)
monkeypatch.setitem(sys.modules, "storage.studio_db", fake_studio_db)
ns = {
"HTTPException": _HTTPException,
"os": os,
"Path": Path,
"Optional": Optional,
"_safe_is_dir": lambda p: Path(p).is_dir(),
"_resolve_hf_cache_dir": lambda: tmp_path / "missing-hf",
"logger": SimpleNamespace(debug = lambda *_args, **_kwargs: None),
}
exec(compile(module, "<extracted routes/models.py>", "exec"), ns)
allowlist = ns["_build_browse_allowlist"]()
assert media_root.resolve() in allowlist
assert ns["_resolve_browse_target"](str(model_dir), allowlist) == model_dir.resolve()
with pytest.raises(_HTTPException) as exc:
ns["_resolve_browse_target"](str(media_root / ".ssh"), allowlist)
assert exc.value.status_code == 403
ssh_root = media_root / ".ssh"
with pytest.raises(_HTTPException) as exc_root:
ns["_resolve_browse_target"](str(ssh_root), [ssh_root])
assert exc_root.value.status_code == 403