mirror of
https://github.com/unslothai/unsloth.git
synced 2026-08-25 08:42:25 +00:00
* Studio: say when a scan folder cannot be read A folder Unsloth is denied looks exactly like an empty one: the scan catches the OSError, logs it, and moves on, so the model list is empty with no reason given. Add-time validation now opens the directory instead of trusting os.access, which reads mode bits only and passes on folders macOS TCC or a Windows ACL still refuses. The check runs after the denylist rules so a denied path is never opened. The scan keeps the error it already caught, and both scan-folders endpoints return it as a per-folder status the dialog shows with the setting that fixes it. No new work on the healthy path: 0.06us per folder per scan, and the folder list is a dict lookup with no syscalls. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Studio: record scan folder status from the Hub inventory scan too The folders dialog reads /api/hub/scan-folders, but the scan behind it is the Hub inventory, not collect_local_models, so nothing was ever recorded for it and every row stayed "ok". Its custom-folder loop now records the same way. That alone is not enough: the Hub's _scan_models_dir catches the OSError itself and returns an empty list, so no exception reaches the loop. So an empty result is now the trigger. One opendir says whether the folder is empty, gone, or refused, and it runs only for a folder that returned no models. A folder that found models still costs nothing, with a test that fails if it ever touches the filesystem. * Studio: catch denied model subdirs, and recheck when the dialog reopens Two gaps in the folder status. A root can list fine while every model under it is denied, on a NAS mount or a drive owned by another user. The scanners skip an unreadable child silently, so that arrives as the same empty list as an empty folder and was reported as ok. The probe now also opens subdirectories, stopping at the first refusal and capped at 64, so the denied-everything case costs one extra open. The row tells the user to fix permissions and reopen the dialog, but nothing rechecked between inventory scans, so the warning stayed up after access was restored. Listing the folders now rechecks the folders marked bad, and only those. A healthy folder is not in the registry, so the list still opens nothing. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Studio: probe two levels down, and keep the tests collectable on Windows The child probe descended one level, so <root>/<publisher>/<model> with the model denied still reported ok: both levels above it list fine and the scanners return nothing. It now walks two levels, depth first, on one shared budget of 64 opens. Depth first means a denied mount is found in three opens instead of after every publisher, and the budget bounds the cost whatever the shape of the tree. os.geteuid does not exist on Windows and a skipif condition is evaluated at import, so collecting the test file there raised AttributeError before the os.name check could skip anything. Resolved once into a shared marker, with a test that runs the module body with geteuid removed. * Studio: flag a folder with one denied model, and show status in the picker A folder holding one readable model and one denied model returned the readable one, so the scan looked successful and the denied model was silently absent. The probe now runs whether or not models were found, and reports "partial" in that case so the copy does not contradict the rows on screen by claiming the folder cannot be read. That is a real cost change on the healthy path, so it is measured rather than claimed: 104us for a folder with 8 model dirs, 0.77ms for one with 300, capped by the same 64-open budget. The end-to-end scan stays inside run-to-run variation, and the folder list still opens nothing unless a folder is already marked bad. The old zero-syscall test is replaced by the bound, which is now the guarantee that matters. The inline model selector manages the same folders and rendered only the path, so it shows the status too. * Studio: stop treating an exhausted probe budget as healthy Three fixes. A denied directory past the open budget was reported as ok. Running out of budget means the tail was never looked at, which is not the same as finding it healthy, so it now returns an internal "unknown" that is never recorded and never sent to the UI. That alone would strand a wide folder in a warning it could never clear, so the registry now remembers which directory refused. A recheck opens that one directory, which settles a fixed folder in a single open no matter how wide the folder is or where the denial sat. A folder recorded as partial kept that status after being deleted, because the recheck preserved partial for every non-ok probe. It now only holds partial against a permission result, so missing and unreadable replace it. One permission test was missing the marker that skips it as root and on Windows, where chmod 000 does not deny. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Probe scan folders off the event loop, and stop a vanished model or a Windows device error condemning the folder * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Probe the commit directory inside an HF snapshots folder, and stop calling a shut root partial Two ways the dialog told the user the wrong thing. The cache layout is models--org--name/snapshots/<commit>/, three levels under a registered root, so a probe that stops at two never opens the one directory the weights live in: a denied commit dir made the model vanish from the list while the folder still reported ok, which is the silent empty case this PR exists to remove. The extra level is bought only for a directory named snapshots, since a blanket third level would spend the open budget descending into diffusers component directories. And when the root itself becomes denied, the partial branch restored partial regardless, so a folder none of which can be read kept saying some models in it could not be read, sending the user hunting for one bad model. The cause the probe already returns is the discriminator: it is the root path itself for the root's own refusal and a nested entry path otherwise. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> Co-authored-by: Daniel Han <danielhanchen@gmail.com>
299 lines
10 KiB
Python
299 lines
10 KiB
Python
# SPDX-License-Identifier: AGPL-3.0-only
|
|
# Copyright 2026-present the Unsloth AI Inc. team. All rights reserved. See /studio/LICENSE.AGPL-3.0
|
|
|
|
from __future__ import annotations
|
|
|
|
import ast
|
|
import os
|
|
import sys
|
|
from pathlib import Path
|
|
from types import SimpleNamespace
|
|
from typing import Optional
|
|
|
|
import pytest
|
|
|
|
from hub.storage import scan_folders
|
|
from storage import studio_db
|
|
from utils.paths import external_media
|
|
|
|
|
|
_BACKEND_ROOT = Path(__file__).resolve().parent.parent
|
|
|
|
|
|
class _ExistingScanFolderConn:
|
|
def __init__(self):
|
|
self.params = ()
|
|
|
|
def execute(
|
|
self,
|
|
_sql,
|
|
params = (),
|
|
):
|
|
self.params = params
|
|
return self
|
|
|
|
def fetchone(self):
|
|
return {"id": 1, "path": self.params[0], "created_at": "fake"}
|
|
|
|
def commit(self):
|
|
pass
|
|
|
|
def close(self):
|
|
pass
|
|
|
|
|
|
class _HTTPException(Exception):
|
|
def __init__(self, status_code: int, detail: str):
|
|
super().__init__(detail)
|
|
self.status_code = status_code
|
|
self.detail = detail
|
|
|
|
|
|
def _stub_linux_path_checks(monkeypatch, module):
|
|
monkeypatch.setattr(module.platform, "system", lambda: "Linux")
|
|
monkeypatch.setattr(module.os.path, "realpath", os.path.normpath)
|
|
monkeypatch.setattr(module.os.path, "expanduser", lambda p: p)
|
|
monkeypatch.setattr(module.os.path, "exists", lambda _p: True)
|
|
monkeypatch.setattr(module.os.path, "isdir", lambda _p: True)
|
|
monkeypatch.setattr(module.os, "access", lambda _p, _mode: True)
|
|
# The mount does not exist on the test host, so the readability probe that
|
|
# opens the directory has to be part of the same pretence.
|
|
monkeypatch.setattr(module, "is_readable_dir", lambda _p: True)
|
|
|
|
|
|
def _stub_hub_scan_folder_db(monkeypatch):
|
|
monkeypatch.setattr(scan_folders, "_ensure_schema", lambda _conn: None)
|
|
monkeypatch.setattr(scan_folders, "get_connection", _ExistingScanFolderConn)
|
|
|
|
|
|
def _stub_legacy_scan_folder_db(monkeypatch):
|
|
monkeypatch.setattr(studio_db, "get_connection", _ExistingScanFolderConn)
|
|
|
|
|
|
def test_linux_run_media_policy_accepts_mounted_volume_descendants(monkeypatch):
|
|
monkeypatch.setattr(external_media.platform, "system", lambda: "Linux")
|
|
|
|
assert external_media.is_linux_run_media_path("/run/media/dspofu/nvmeB")
|
|
assert external_media.is_linux_run_media_path("/run/media/dspofu/nvmeB/modelsAI/gguf/qwen3.6")
|
|
|
|
|
|
@pytest.mark.parametrize(
|
|
"path",
|
|
[
|
|
"/run",
|
|
"/run/media",
|
|
"/run/media/dspofu",
|
|
"/run/user/1000/models",
|
|
"/run/systemd/private",
|
|
"/run/not-media/dspofu/nvmeB",
|
|
],
|
|
)
|
|
def test_linux_run_media_policy_rejects_unrelated_run_paths(monkeypatch, path):
|
|
monkeypatch.setattr(external_media.platform, "system", lambda: "Linux")
|
|
|
|
assert not external_media.is_linux_run_media_path(path)
|
|
|
|
|
|
def test_linux_run_media_mount_roots_lists_readable_volume_roots(monkeypatch, tmp_path):
|
|
base = tmp_path / "run" / "media"
|
|
mount = base / "dspofu" / "nvmeB"
|
|
sensitive_mount = base / "dspofu" / ".ssh"
|
|
sensitive_aws_mount = base / "dspofu" / ".aws"
|
|
other_user_mount = base / "other" / "backup"
|
|
incomplete = base / "dspofu-only"
|
|
mount.mkdir(parents = True)
|
|
sensitive_mount.mkdir()
|
|
sensitive_aws_mount.mkdir()
|
|
other_user_mount.mkdir(parents = True)
|
|
incomplete.mkdir()
|
|
monkeypatch.setattr(external_media.platform, "system", lambda: "Linux")
|
|
|
|
roots = external_media.linux_run_media_mount_roots(base, user = "dspofu")
|
|
|
|
assert roots == [mount.resolve()]
|
|
|
|
|
|
def test_linux_run_media_mount_roots_skips_sensitive_resolved_volume_name(monkeypatch, tmp_path):
|
|
base = tmp_path / "run" / "media"
|
|
normal_mount = base / "dspofu" / "nvmeB"
|
|
sensitive_target = base / "dspofu" / ".config"
|
|
normal_mount.mkdir(parents = True)
|
|
sensitive_target.mkdir()
|
|
alias = base / "dspofu" / "config-alias"
|
|
alias.symlink_to(sensitive_target, target_is_directory = True)
|
|
monkeypatch.setattr(external_media.platform, "system", lambda: "Linux")
|
|
|
|
roots = external_media.linux_run_media_mount_roots(base, user = "dspofu")
|
|
|
|
assert roots == [normal_mount.resolve()]
|
|
|
|
|
|
def test_linux_run_media_mount_roots_skips_sensitive_resolved_descendant(monkeypatch, tmp_path):
|
|
base = tmp_path / "run" / "media"
|
|
normal_mount = base / "dspofu" / "nvmeB"
|
|
sensitive_descendant = normal_mount / ".ssh" / "models"
|
|
sensitive_descendant.mkdir(parents = True)
|
|
alias = base / "dspofu" / "models-alias"
|
|
alias.symlink_to(sensitive_descendant, target_is_directory = True)
|
|
monkeypatch.setattr(external_media.platform, "system", lambda: "Linux")
|
|
|
|
roots = external_media.linux_run_media_mount_roots(base, user = "dspofu")
|
|
|
|
assert roots == [normal_mount.resolve()]
|
|
|
|
|
|
def test_hub_scan_folder_accepts_linux_run_media_mount(monkeypatch):
|
|
_stub_linux_path_checks(monkeypatch, scan_folders)
|
|
monkeypatch.setattr(external_media.platform, "system", lambda: "Linux")
|
|
_stub_hub_scan_folder_db(monkeypatch)
|
|
target = "/run/media/dspofu/nvmeB/modelsAI/gguf/qwen3.6"
|
|
|
|
row = scan_folders.add_scan_folder(target)
|
|
|
|
assert row["path"] == target
|
|
|
|
|
|
@pytest.mark.parametrize(
|
|
"target",
|
|
[
|
|
"/run",
|
|
"/run/media",
|
|
"/run/media/dspofu",
|
|
"/run/user/1000/models",
|
|
"/run/systemd/private",
|
|
"/run/not-media/dspofu/nvmeB",
|
|
],
|
|
)
|
|
def test_hub_scan_folder_keeps_unrelated_run_paths_blocked(monkeypatch, target):
|
|
_stub_linux_path_checks(monkeypatch, scan_folders)
|
|
monkeypatch.setattr(external_media.platform, "system", lambda: "Linux")
|
|
_stub_hub_scan_folder_db(monkeypatch)
|
|
|
|
with pytest.raises(ValueError, match = "Path under /run is not allowed"):
|
|
scan_folders.add_scan_folder(target)
|
|
|
|
|
|
def test_hub_scan_folder_keeps_sensitive_dirs_blocked_under_run_media(monkeypatch):
|
|
_stub_linux_path_checks(monkeypatch, scan_folders)
|
|
monkeypatch.setattr(external_media.platform, "system", lambda: "Linux")
|
|
_stub_hub_scan_folder_db(monkeypatch)
|
|
|
|
with pytest.raises(ValueError, match = "Credential or configuration"):
|
|
scan_folders.add_scan_folder("/run/media/dspofu/nvmeB/.ssh/models")
|
|
|
|
|
|
def test_legacy_scan_folder_accepts_linux_run_media_mount(monkeypatch):
|
|
_stub_linux_path_checks(monkeypatch, studio_db)
|
|
monkeypatch.setattr(external_media.platform, "system", lambda: "Linux")
|
|
_stub_legacy_scan_folder_db(monkeypatch)
|
|
target = "/run/media/dspofu/nvmeB/modelsAI/gguf/qwen3.6"
|
|
|
|
row = studio_db.add_scan_folder(target)
|
|
|
|
assert row["path"] == target
|
|
|
|
|
|
@pytest.mark.parametrize(
|
|
"target",
|
|
[
|
|
"/run",
|
|
"/run/media",
|
|
"/run/media/dspofu",
|
|
"/run/user/1000/models",
|
|
"/run/systemd/private",
|
|
"/run/not-media/dspofu/nvmeB",
|
|
],
|
|
)
|
|
def test_legacy_scan_folder_keeps_unrelated_run_paths_blocked(monkeypatch, target):
|
|
_stub_linux_path_checks(monkeypatch, studio_db)
|
|
monkeypatch.setattr(external_media.platform, "system", lambda: "Linux")
|
|
_stub_legacy_scan_folder_db(monkeypatch)
|
|
|
|
with pytest.raises(ValueError, match = "Path under /run is not allowed"):
|
|
studio_db.add_scan_folder(target)
|
|
|
|
|
|
def test_legacy_scan_folder_keeps_sensitive_dirs_blocked_under_run_media(monkeypatch):
|
|
_stub_linux_path_checks(monkeypatch, studio_db)
|
|
monkeypatch.setattr(external_media.platform, "system", lambda: "Linux")
|
|
_stub_legacy_scan_folder_db(monkeypatch)
|
|
|
|
with pytest.raises(ValueError, match = "Credential or configuration"):
|
|
studio_db.add_scan_folder("/run/media/dspofu/nvmeB/.aws/models")
|
|
|
|
|
|
def test_legacy_browse_allowlist_includes_linux_run_media_mounts(monkeypatch, tmp_path):
|
|
tree = ast.parse((_BACKEND_ROOT / "routes" / "models.py").read_text(encoding = "utf-8"))
|
|
function_names = {
|
|
"_build_browse_allowlist",
|
|
"_browse_relative_parts",
|
|
"_is_path_inside_allowlist",
|
|
"_match_browse_child",
|
|
"_normalize_browse_request_path",
|
|
"_resolve_browse_target",
|
|
}
|
|
functions = [
|
|
node
|
|
for node in tree.body
|
|
if isinstance(node, ast.FunctionDef) and node.name in function_names
|
|
]
|
|
module = ast.Module(body = functions, type_ignores = [])
|
|
ast.fix_missing_locations(module)
|
|
|
|
home = tmp_path / "home"
|
|
media_root = tmp_path / "run" / "media" / "dspofu" / "nvmeB"
|
|
model_dir = media_root / "modelsAI" / "gguf" / "qwen3.6"
|
|
home.mkdir()
|
|
model_dir.mkdir(parents = True)
|
|
(media_root / ".ssh").mkdir()
|
|
|
|
fake_paths = SimpleNamespace(
|
|
hf_default_cache_dir = lambda: tmp_path / "missing-default-hf",
|
|
legacy_hf_cache_dir = lambda: tmp_path / "missing-legacy-hf",
|
|
well_known_model_dirs = lambda: [],
|
|
studio_root = lambda: tmp_path / "missing-studio",
|
|
outputs_root = lambda: tmp_path / "missing-outputs",
|
|
exports_root = lambda: tmp_path / "missing-exports",
|
|
)
|
|
fake_external_media = SimpleNamespace(
|
|
linux_run_media_mount_roots = lambda: [media_root],
|
|
macos_volume_roots = lambda: [],
|
|
windows_drive_roots = lambda: [],
|
|
)
|
|
fake_paths.external_media = fake_external_media
|
|
fake_studio_db = SimpleNamespace(
|
|
list_scan_folders = lambda: [],
|
|
contains_sensitive_path_component = studio_db.contains_sensitive_path_component,
|
|
# The media root is a legitimate mount, not denied; the .ssh 403 below
|
|
# comes from the credential check. A False stub keeps this OS-independent
|
|
# (on macOS tmp_path lives under the denied /private/var).
|
|
is_denied_system_path = lambda _p: False,
|
|
)
|
|
monkeypatch.setitem(sys.modules, "utils.paths", fake_paths)
|
|
monkeypatch.setitem(sys.modules, "utils.paths.external_media", fake_external_media)
|
|
monkeypatch.setitem(sys.modules, "storage.studio_db", fake_studio_db)
|
|
|
|
ns = {
|
|
"HTTPException": _HTTPException,
|
|
"os": os,
|
|
"Path": Path,
|
|
"Optional": Optional,
|
|
"_safe_is_dir": lambda p: Path(p).is_dir(),
|
|
"_resolve_hf_cache_dir": lambda: tmp_path / "missing-hf",
|
|
"logger": SimpleNamespace(debug = lambda *_args, **_kwargs: None),
|
|
}
|
|
exec(compile(module, "<extracted routes/models.py>", "exec"), ns)
|
|
|
|
allowlist = ns["_build_browse_allowlist"]()
|
|
|
|
assert media_root.resolve() in allowlist
|
|
assert ns["_resolve_browse_target"](str(model_dir), allowlist) == model_dir.resolve()
|
|
|
|
with pytest.raises(_HTTPException) as exc:
|
|
ns["_resolve_browse_target"](str(media_root / ".ssh"), allowlist)
|
|
assert exc.value.status_code == 403
|
|
|
|
ssh_root = media_root / ".ssh"
|
|
with pytest.raises(_HTTPException) as exc_root:
|
|
ns["_resolve_browse_target"](str(ssh_root), [ssh_root])
|
|
assert exc_root.value.status_code == 403
|