mirror of
https://github.com/unslothai/unsloth.git
synced 2026-08-07 07:54:10 +00:00
Some checks are pending
Core / Core (HF=default + TRL=default) (push) Waiting to run
Core / Core (HF=4.57.6 + TRL<1) (push) Waiting to run
Core / Core (HF=latest + TRL=latest) (push) Waiting to run
Core / llama.cpp build + smoke (push) Waiting to run
Cross-platform parity / parity (macos-latest) (push) Waiting to run
Cross-platform parity / parity (ubuntu-latest) (push) Waiting to run
Cross-platform parity / parity (windows-latest) (push) Waiting to run
Lint CI / Source lint (Python + shell + YAML + JSON + safety nets) (push) Waiting to run
MLX CI on Mac M1 / dispatch (push) Waiting to run
Scorecard supply-chain security / Scorecard analysis (push) Waiting to run
Security audit / advisory audit (pip + npm + cargo) (push) Waiting to run
Security audit / pip scan-packages :: extras (push) Waiting to run
Security audit / pip scan-packages :: studio (push) Waiting to run
Security audit / pip scan-packages :: hf-stack (push) Waiting to run
Security audit / npm scan-packages (Unsloth frontend tarballs) (push) Waiting to run
Security audit / workflow-trigger lint (pull_request_target / cache-poisoning) (push) Waiting to run
Security audit / pytest tests/security (push) Waiting to run
Security audit / npm provenance + new install-script diff (push) Waiting to run
Unsloth API CI / Unsloth API & Auth Tests (push) Waiting to run
Backend CI / (Python 3.10) (push) Waiting to run
Backend CI / (Python 3.11) (push) Waiting to run
Backend CI / (Python 3.12) (push) Waiting to run
Backend CI / (Python 3.13) (push) Waiting to run
Backend CI / Repo tests (CPU) (push) Waiting to run
Unsloth export capability / capability (macos-latest) (push) Waiting to run
Unsloth export capability / capability (ubuntu-latest) (push) Waiting to run
Unsloth export capability / capability (windows-latest) (push) Waiting to run
Frontend CI / Frontend build + bundle sanity (push) Waiting to run
Unsloth GGUF CI / OpenAI, Anthropic API tests (push) Waiting to run
Unsloth GGUF CI / Tool calling Tests (push) Waiting to run
Unsloth GGUF CI / JSON, images (push) Waiting to run
Unsloth load-orchestrator CI / test (push) Waiting to run
Mac Studio API CI / Unsloth API & Auth Tests (push) Waiting to run
Mac Studio GGUF CI / OpenAI, Anthropic API tests (push) Waiting to run
Mac Studio GGUF CI / Tool calling Tests (push) Waiting to run
Mac Studio GGUF CI / JSON, images (push) Waiting to run
Mac Studio Install Matrix CI / Install + load (macos-14) (push) Waiting to run
Mac Studio Install Matrix CI / Install + load (macos-15) (push) Waiting to run
Mac Studio Install Matrix CI / Install + load (macos-26) (push) Waiting to run
Mac Studio Install Matrix CI / Install + load (macos-15-intel) (push) Waiting to run
Mac Studio Install Matrix CI / Install + load (macos-26-intel) (push) Waiting to run
Mac Studio UI CI / Chat UI Tests (push) Waiting to run
Mac Studio Update CI / Unsloth Updating Tests (push) Waiting to run
Unsloth Tauri CI / Tauri Linux debug build (no codesign) (push) Waiting to run
Unsloth UI CI / Chat UI Tests (push) Waiting to run
Unsloth Update CI / Unsloth Updating Tests (push) Waiting to run
Windows Unsloth API CI / Unsloth API & Auth Tests (push) Waiting to run
Windows Unsloth GGUF CI / OpenAI, Anthropic API tests (push) Waiting to run
Windows Unsloth GGUF CI / Tool calling Tests (push) Waiting to run
Windows Unsloth GGUF CI / JSON, images (push) Waiting to run
Windows Unsloth GGUF CI / Unsloth install + inference without Visual Studio (push) Waiting to run
Windows Unsloth GGUF CI / GPU prebuilt resolves without Visual Studio (push) Waiting to run
Windows Unsloth GGUF CI / setup.ps1 unit tests (VS 2026 / CMake guard) (push) Waiting to run
Windows Unsloth GGUF CI / real-VS detection (VS 2022) (push) Waiting to run
Windows Unsloth GGUF CI / real-VS detection (VS 2026) (push) Waiting to run
Windows Unsloth GGUF CI / VC++ runtime detect + install round-trip (windows-2025-vs2026) (push) Waiting to run
Windows Unsloth GGUF CI / VC++ runtime detect + install round-trip (windows-latest) (push) Waiting to run
Windows Unsloth UI CI / Chat UI Tests (push) Waiting to run
Windows Unsloth Update CI / Unsloth Updating Tests (push) Waiting to run
Wheel CI / Wheel build + content sanity + import smoke (push) Waiting to run
* feat(install): opt-in Vulkan llama.cpp backend and HIP gfx fallback (#7357) Add UNSLOTH_LLAMA_BACKEND=vulkan and --llama-backend vulkan to force the upstream Vulkan prebuilt on any host, persist llama_backend in the install marker, and re-assert it during Studio updates. On Windows AMD, auto-fallback to Vulkan when no detected gfx arch is in the upstream win-hip-radeon GPU_TARGETS set (e.g. gfx803 / RX 480). Mixed setups where at least one card is HIP-supported still default to HIP unless opted in. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * fix(install): address Codex P2s on Vulkan gfx routing (#7357) Honor ROCm family tokens (gfx110X), include fork-supported gfx1103, require a known active gfx before auto-Vulkan, and base the HIP floor check on the visible-device target instead of every physical GPU in hipinfo. * Address Codex review: env namespace, physical-NVIDIA guard, test kwarg - llama_backend_from_env: stop reading UNSLOTH_LLAMA_CPP_BACKEND. That is a separate pre-existing setup variable meaning auto/cpu; setup.sh/setup.ps1 warn and ignore other values, so reading it here forced Vulkan behind that warning. Vulkan opt-in stays on UNSLOTH_LLAMA_BACKEND / UNSLOTH_FORCE_VULKAN. - _should_auto_vulkan_for_amd_windows: gate on not has_physical_nvidia (not merely has_usable_nvidia). A CUDA-masked NVIDIA card keeps has_physical_nvidia while has_usable_nvidia goes False; Vulkan ignores CUDA_VISIBLE_DEVICES and could enumerate the reserved card. Mirrors the Intel auto path. Explicit opt-in still overrides. - test fakes: validate_prebuilt_attempts/validate_prebuilt_choice gained a llama_backend kwarg; the four fake signatures in the fallback tests now accept it, clearing the TypeError that reddened Backend CI / Repo tests (CPU). Tests: UNSLOTH_LLAMA_CPP_BACKEND=vulkan no longer triggers Vulkan; hidden physical NVIDIA suppresses AMD auto-Vulkan while explicit opt-in overrides. * Keep gfx1034 on the ROCm path (fork gfx103X bundle covers it) The WINDOWS_HIP_PREBUILT_GFX_TARGETS allow-list omitted gfx1034, so _route_to_vulkan_prebuilt downgraded RX 6500/6400-class hosts to the upstream Vulkan prebuilt before published_rocm_choice_for_host could match the fork windows-rocm gfx103X bundle (whose members include gfx1034). Add gfx1034 to the allow-list and a regression test asserting it stays on the fork ROCm asset. * Fix auto-Vulkan stealing fork windows-rocm gfx908/gfx90a hosts for PR #7373 * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Fix Vulkan marker claiming a backend that was never installed for PR #7373 * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Tighten the Vulkan backend routing comments for PR #7373 * Keep the visible-device-aware gfx when setup forwards --rocm-gfx setup.ps1 resolves the gfx arch from its own probe, and that pick is not fully visible-device aware: neither the hipinfo nor the amd-smi branch reads CUDA_VISIBLE_DEVICES, and the amd-smi branch matches a bare integer only, so a comma-separated HIP/ROCR mask such as 1,0 also falls back to GPU 0. The resulting arch was then forwarded through --rocm-gfx and replaced the arch detect_host() had already resolved for the runtime-visible GPU. On a mixed-AMD Windows host that flipped the auto-Vulkan decision: with GPU 0 gfx1100 and a masked-in gfx1010, the forward reinstated gfx1100, _should_auto_vulkan_for_amd_windows() saw a HIP-supported arch and the HIP bundle was installed for a GPU that cannot run it. Fold the forward in as a fill rather than a replacement: it still supplies the arch on amd-smi-only, driver-only and name-inferred hosts where the probe reports none, which is what --rocm-gfx exists for, but no longer overwrites a successfully detected active arch. An explicit UNSLOTH_ROCM_GFX_ARCH stays authoritative, since it is the documented manual override for hosts whose arch the probes get wrong. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Scope the Windows AMD Vulkan fallback per device and per repo Three follow-ups on the auto-Vulkan routing for #7357. Keep an explicit --rocm-gfx authoritative. The previous round stopped a forwarded gfx from replacing an arch detect_host() had already resolved, but --rocm-gfx is also the documented operator override for hosts whose probe is wrong or stale, and both arrive as the same argv. Narrow the advisory case to the two shapes setup can actually be describing: an arch the probe saw on this host (setup picked a different physical GPU of the same box), or a family label such as gfx110X, which is a bundle name the update path derives from the marker asset rather than a real GPU arch. Any other value is an override for an arch no probe reported and stays authoritative. Keeping family labels advisory also preserves the rule that an in-generation-but-unbuilt arch (gfx1033) is never upgraded into the gfx103X bundle. Do not auto-route to Vulkan from a HIP-only device mask. HIP_VISIBLE_DEVICES, ROCR_VISIBLE_DEVICES and CUDA_VISIBLE_DEVICES select the active arch, but the Vulkan runtime honours none of them: it enumerates through GGML_VK_VISIBLE_DEVICES and Vulkan ordinals in LlamaCppBackend._get_gpu_free_memory_vulkan. Masking down to a below-floor card therefore used to install a backend that could still enumerate the HIP-capable card the user deliberately hid, possibly one reserved for another workload. Require every physical AMD gfx to be below the floor, matching the has_physical_nvidia gate right above it. So the per-GPU list survives to that check, a forward that agrees with the probe no longer collapses rocm_gfx_targets to a single entry. Make the HIP support predicate repository-specific. The floor constant is a union of ggml-org's windows-hip gpu_targets and the fork's windows-rocm bundles, so it only answers "is this arch served" for the fork. With --published-repo ggml-org/llama.cpp, direct_upstream_release_plan() offers win-hip-radeon then CPU and never Vulkan, so the four fork-only archs (gfx908, gfx90a, gfx1034, gfx1103) were declared supported and fell through to CPU instead of the Vulkan bundle that would actually run. Add UPSTREAM_WINDOWS_HIP_GFX_TARGETS and select the set from the planned repo. * Keep probe-confirmed AMD GPUs in the physical list when a gfx is forwarded rocm_gfx_targets is the physical inventory _should_auto_vulkan_for_amd_windows() reads, so a forwarded --rocm-gfx that the probe never reported was deleting cards the probe had confirmed. On a mixed Windows AMD box whose active device is masked down to a below-floor card, a stale UNSLOTH_ROCM_GFX_ARCH or a name-inferred arch for the other GPU collapsed the list to that one arch, the floor check concluded no AMD GPU on the host reaches the Windows HIP prebuilt, and the install auto-fell back to Vulkan, which honours no HIP mask and would enumerate the reserved HIP-capable card. Add the forwarded arch to the list instead of replacing it: it selects the HIP target, it does not redefine what hardware is present. An empty probe still yields a single-entry list, so the driver-only Windows AMD host the forward exists for keeps its automatic Vulkan fallback, and an explicit --llama-backend vulkan is unaffected. * Do not auto-fall back to Vulkan when a HIP device mask filtered the probe hipinfo is itself a HIP application, and AMD documents HIP_VISIBLE_DEVICES as "only devices whose index is present in the sequence are visible to HIP", with that spelling recommended on Windows. Under a mask the Windows probe therefore enumerates the visible devices, so rocm_gfx_targets is what survived the mask rather than the physical inventory the auto-Vulkan floor check assumes. A masked-out gfx1100 next to a visible gfx803 made the check conclude that no AMD GPU on the box reaches the Windows HIP prebuilt and route the install to Vulkan, which honours none of these masks and would enumerate the reserved card. Decline to guess when a mask is set: the physical inventory is unknowable from a masked probe, so keep the HIP / fork / source path. This only ever turns the automatic fallback off, never on. The driver-only single-GPU host the fallback exists for sets no mask, an all-hiding "" / -1 mask is still handled as no active target rather than a partial view, and an explicit --llama-backend vulkan or UNSLOTH_LLAMA_BACKEND=vulkan is unaffected. Reading the physical inventory through an unmasked re-probe would also correct _pick_rocm_gfx_target, which indexes the token list by the mask value and so already assumes an unmasked probe. That is pre-existing behaviour on main and is left alone here. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Treat an all-hiding HIP device mask as suppressing the Vulkan fallback too The mask guard exempted an empty or -1 value on the grounds that the probe reports no active target under it, but that only holds for the probe: a forwarded --rocm-gfx still reconstructs an active arch, and setup infers that arch from the display-adapter name, which no HIP mask touches. A user who hid every AMD GPU from HIP could therefore still be auto-routed to Vulkan, which honours none of these masks and would then use all of them. That is the strongest form of the hazard the guard exists for, not an exemption from it. Presence of any of the three variables is now the whole test, which also removes the value parsing. An explicit --llama-backend vulkan or UNSLOTH_LLAMA_BACKEND is still unaffected. * Grant the fork-only Windows HIP coverage to the fork, not to every mirror The floor set is a union of the fork's windows-rocm bundles and only the fork is planned from its manifest: resolve_simple_install_release_plans() compares == DEFAULT_PUBLISHED_REPO and sends every other --published-repo through direct_upstream_release_plan(), whose AMD branch offers win-hip-radeon then CPU and never Vulkan. Exempting only the exact ggml-org spelling therefore told a mirror carrying upstream-standard assets that fork-only archs such as gfx1034, gfx1103 and gfx908 were HIP-served, landing them on HIP or CPU instead of the Vulkan bundle that would actually run. Gate on the fork instead. Matching the dispatch exactly, spelling included, also fixes a differently cased repo: that really does take the upstream path, so it must be answered with upstream coverage rather than the fork superset. An empty repo still defaults to the fork, as the resolver does. * Derive the Windows HIP gfx floor guard from the published manifest The guard compared WINDOWS_HIP_PREBUILT_GFX_TARGETS against a second hardcoded tuple in the same test file, so a windows-rocm arch newly published by the fork passed both. Affected hosts would then be routed off the hash-approved fork ROCm bundle onto an unhashed upstream Vulkan build with nothing failing. Read the fork's llama-prebuilt-manifest.json through the installer's own resolver instead, and assert the floor, the family labels, and the routing tuple all still cover what it publishes. The manifest ships only as a release asset, so an unreachable release skips with an explicit reason rather than flaking. Both literals match the manifest as published today. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Compress the Vulkan backend routing comments and docstrings for PR #7373 * Correct the family-label rationale in the Windows HIP coverage check The comment justified serving gfx103X / gfx110X against any repository by claiming upstream's windows-hip targets build every member of those families. The fork manifest maps gfx103X to gfx1030..1032 plus gfx1034 and gfx110X to gfx1100..1102 plus gfx1103, and UPSTREAM_WINDOWS_HIP_GFX_TARGETS carries neither gfx1034 nor gfx1103, so the stated reason is wrong even though the answer is right. State the real reason instead. A family label is a bundle name, not an arch, so the concrete GPU is unknown at this point; answering unsupported to cover the two uncovered members would move gfx1030..1032 and gfx1100..1102 off a working HIP build onto Vulkan for a card the label cannot identify. Those two archs still reach Vulkan through the concrete-arch branch below, which does answer per repository. Comment only. No behaviour change: the 5850-combination override sweep still reports 0 rocm_gfx_target changes, 0 auto_vulkan False to True flips and 680 True to False flips all backed by a probe-confirmed HIP GPU, and both the feature and override profile matrices are byte-identical. * Pin that a deliberate CPU install outranks Vulkan for PR #7373 UNSLOTH_LLAMA_CPP_BACKEND (setup.sh / setup.ps1, "auto" or "cpu") and UNSLOTH_LLAMA_BACKEND (this module, a backend name) are separate variables at separate layers, and both accept "cpu". setup translates its own =cpu into --force-cpu, which is what pins the CPU-only bundle on a GPU host and keeps Intel iGPU Vulkan crashes away (#7213), so no trigger this PR adds may outrank it. _route_to_vulkan_prebuilt already gets this right, since force_cpu short-circuits ahead of the forced, auto-Intel and auto-no-HIP triggers. Cover it so it stays that way: the matrix runs [Linux, Windows, macOS] x [NVIDIA, AMD, Intel, CPU only] x [unset, vulkan, hip, rocm, cpu] with the legacy UNSLOTH_FORCE_VULKAN set as well, and asserts the published bundle survives every one. WSL presents as Linux to this resolver, so it rides the Linux row. Also assert the guard is not vacuous: the same host still takes Vulkan once the CPU pin is gone, so the matrix cannot pass on a resolver that had simply stopped routing to Vulkan. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> Co-authored-by: LeoBorcherding <borchborchmail@gmail.com> Co-authored-by: Daniel Han <danielhanchen@gmail.com>
757 lines
31 KiB
Python
757 lines
31 KiB
Python
# SPDX-License-Identifier: AGPL-3.0-only
|
|
# Copyright 2026-present the Unsloth AI Inc. team. All rights reserved. See /studio/LICENSE.AGPL-3.0
|
|
|
|
"""In-app llama.cpp prebuilt update.
|
|
|
|
Builds on utils.llama_cpp_freshness (which detects whether a newer prebuilt
|
|
release exists) and adds the *apply* half: run install_llama_prebuilt.py to
|
|
download the newest bundle for this host and atomically swap it in place, so
|
|
the next model load uses it.
|
|
|
|
Design notes:
|
|
- Detection is delegated to check_prebuilt_freshness(). We surface an
|
|
``update_available`` flag (installed_tag != latest_tag) which is laxer than
|
|
freshness' ``stale`` (which additionally requires the install to be >= 3 days
|
|
old). The UI shows the "Update llama.cpp" affordance on update_available.
|
|
- The install is slow (download + extract + validate), so it runs on a daemon
|
|
thread; callers poll get_update_status() for the job state.
|
|
- Everything fails open: a missing marker / offline GitHub / source build just
|
|
reports update_available=False and never blocks the app.
|
|
- The mechanics (managed-root resolution, local-link detection, the resolve
|
|
probe, the streamed installer run) live in utils.prebuilt.update_flow; this
|
|
module keeps the llama policy and the job dict its callers poll.
|
|
- This is the single main update item: whisper.cpp piggybacks on it. Status
|
|
folds in a whisper sub-status (update_available becomes the union) and apply
|
|
chains a whisper phase after the llama phase when whisper is behind (see
|
|
update_flow.run_chained_update and whisper_cpp_update.chained_phase_plan).
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
import os
|
|
import re
|
|
import subprocess
|
|
import sys
|
|
import threading
|
|
from pathlib import Path
|
|
from typing import Optional
|
|
|
|
import structlog
|
|
|
|
from utils.llama_cpp_freshness import (
|
|
_INSTALL_MARKER_NAME,
|
|
check_prebuilt_freshness,
|
|
latest_published_release,
|
|
latest_release_assets,
|
|
parse_base_build,
|
|
read_install_marker,
|
|
reset_caches,
|
|
update_download_size_bytes,
|
|
)
|
|
from utils.prebuilt import update_flow as _flow
|
|
|
|
logger = structlog.get_logger(__name__)
|
|
|
|
DEFAULT_PUBLISHED_REPO = "unslothai/llama.cpp"
|
|
_INSTALL_TIMEOUT_SECONDS = 1800 # 30 min ceiling for download + build/validate
|
|
# install_llama_prebuilt.py EXIT_NO_SPACE: out of disk, retrying will not help.
|
|
_EXIT_NO_SPACE = 4
|
|
|
|
# Background job state. Single in-flight update at a time, guarded by _job_lock.
|
|
_JOB_IDLE = _flow.JOB_IDLE
|
|
_JOB_RUNNING = _flow.JOB_RUNNING
|
|
_JOB_SUCCESS = _flow.JOB_SUCCESS
|
|
_JOB_ERROR = _flow.JOB_ERROR
|
|
|
|
_job_lock = threading.Lock()
|
|
_job: dict = _flow.new_job()
|
|
|
|
_utcnow = _flow.utcnow
|
|
_is_under = _flow.is_under
|
|
_is_external_link = _flow.is_external_link
|
|
_rocm_install_args = _flow.rocm_install_args
|
|
|
|
|
|
def _find_binary() -> Optional[str]:
|
|
"""Locate the active llama-server binary via the inference backend's own
|
|
resolver, so update targets exactly what Unsloth runs. Lazy import keeps the
|
|
heavy inference module off this module's import path."""
|
|
try:
|
|
from core.inference.llama_cpp import LlamaCppBackend
|
|
return LlamaCppBackend._find_llama_server_binary()
|
|
except Exception as exc: # pragma: no cover - defensive
|
|
logger.debug("llama update: binary discovery failed", error = str(exc))
|
|
return None
|
|
|
|
|
|
def _install_dir_for(binary_path: Optional[str]) -> Optional[Path]:
|
|
"""The directory holding UNSLOTH_PREBUILT_INFO.json -- i.e. the install root
|
|
install_llama_prebuilt.py wrote and the one we re-install into."""
|
|
return _flow.install_dir_for(binary_path, marker_name = _INSTALL_MARKER_NAME)
|
|
|
|
|
|
def _installer_script() -> Optional[Path]:
|
|
"""Locate install_llama_prebuilt.py (UNSLOTH_LLAMA_INSTALLER wins)."""
|
|
return _flow.find_installer_script(
|
|
env_var = "UNSLOTH_LLAMA_INSTALLER", script_name = "install_llama_prebuilt.py"
|
|
)
|
|
|
|
|
|
# Markerless (source-build) installs have no UNSLOTH_PREBUILT_INFO.json, so we
|
|
# ask the installer whether an official prebuilt now exists for this host.
|
|
_resolve_memo: dict = {}
|
|
|
|
|
|
def _resolve_prebuilt_for_host(*, force_refresh: bool = False) -> Optional[dict]:
|
|
"""Run install_llama_prebuilt.py --resolve-prebuilt (no download) and return
|
|
{prebuilt_available, repo, release_tag, llama_tag, asset, install_kind} or
|
|
None. Fail-open: any error -> None so a source build never blocks the app."""
|
|
return _flow.resolve_prebuilt_for_host(
|
|
force_refresh = force_refresh,
|
|
memo = _resolve_memo,
|
|
installer_script = lambda: _installer_script(),
|
|
log_message = "llama update: resolve-prebuilt failed",
|
|
)
|
|
|
|
|
|
def _installed_build_number(binary: Optional[str]) -> Optional[int]:
|
|
"""Best-effort build number from ``llama-server --version`` (e.g.
|
|
'version: 9585 (abc)'). None when unparseable or <= 1: a source build with
|
|
no git tags reports 'version: 1', which we treat as unknown (offer update)."""
|
|
if not binary:
|
|
return None
|
|
try:
|
|
proc = subprocess.run([binary, "--version"], capture_output = True, text = True, timeout = 20)
|
|
except Exception: # pragma: no cover - defensive
|
|
return None
|
|
m = re.search(r"version:\s*(\d+)", (proc.stderr or "") + (proc.stdout or ""))
|
|
if not m:
|
|
return None
|
|
n = int(m.group(1))
|
|
return n if n > 1 else None
|
|
|
|
|
|
def get_installed_llama_version() -> Optional[str]:
|
|
"""Display string for the active llama.cpp install (e.g. 'b9585' or
|
|
'b9601-mix-a0e2906'), or None.
|
|
|
|
Prefers the install marker's release_tag -- the full unsloth release
|
|
identity, the same field the update banner compares as installed (see
|
|
#6219) -- so a 'b9601-mix-a0e2906' build reads back in full rather than
|
|
collapsing to its base 'b9601'. The marker's bare ``tag`` is only the
|
|
upstream llama.cpp build (no '-mix-<commit>' suffix), so it's the fallback.
|
|
Last resort is ``b<build>`` parsed from ``llama-server --version`` for
|
|
source/custom builds that have no marker.
|
|
|
|
Lightweight: reads the local marker and at most runs ``--version``. Does no
|
|
network or release-freshness work (unlike get_update_status), so it is safe
|
|
to call from latency-sensitive paths like the About panel.
|
|
"""
|
|
binary = _find_binary()
|
|
marker = read_install_marker(binary)
|
|
if marker:
|
|
tag = marker.get("release_tag") or marker.get("tag")
|
|
if tag:
|
|
return tag
|
|
# Markerless/source build: the fallback execs ``llama-server --version``.
|
|
# Skip it while an update is swapping the tree -- on Windows that exec can
|
|
# make the installer's os.replace fail (the same race get_update_status's
|
|
# source-build probe guards against). The panel just omits the row.
|
|
with _job_lock:
|
|
job_running = _job["state"] == _JOB_RUNNING
|
|
if job_running:
|
|
return None
|
|
n = _installed_build_number(binary)
|
|
return f"b{n}" if n is not None else None
|
|
|
|
|
|
def _llama_install_root(binary: Optional[str]) -> Optional[Path]:
|
|
"""The Unsloth-managed llama.cpp root the active binary lives under, or None
|
|
when the binary is unmanaged (see update_flow.managed_install_root)."""
|
|
return _flow.managed_install_root(
|
|
binary,
|
|
marker_root = _install_dir_for(binary),
|
|
server_path_var = "LLAMA_SERVER_PATH",
|
|
cpp_path_var = "UNSLOTH_LLAMA_CPP_PATH",
|
|
dir_name = "llama.cpp",
|
|
)
|
|
|
|
|
|
def _source_build_status(binary: str, *, force_refresh: bool) -> Optional[dict]:
|
|
"""Update status for a markerless (source-build) install: offer the official
|
|
prebuilt when one exists for this host and is newer than the installed
|
|
binary. None -> caller falls through to the no-marker default (unsupported)."""
|
|
res = _resolve_prebuilt_for_host(force_refresh = force_refresh)
|
|
if not res or not res.get("prebuilt_available"):
|
|
return None
|
|
# llama_tag is the upstream base (bNNNN, what --version reports); release_tag
|
|
# is the full tag, either a same-base mix (bNNNN-mix-<sha>) or a fork wrapper
|
|
# (e.g. v1.0). Compare the numeric base against llama_tag.
|
|
base_tag = res.get("llama_tag") or res.get("release_tag")
|
|
release_tag = res.get("release_tag")
|
|
if not base_tag:
|
|
return None
|
|
# No resolvable install root (e.g. a pinned LLAMA_SERVER_PATH we cannot
|
|
# manage) means an apply would not take effect, so do not offer.
|
|
if _llama_install_root(binary) is None:
|
|
return None
|
|
installed_build = _installed_build_number(binary)
|
|
latest_build = parse_base_build(base_tag)
|
|
# A same-base mix adds patches the bare base lacks, so it is newer even at an
|
|
# unchanged build number (the marker path's is_behind already does this). The
|
|
# bNNNN anchor keeps a fork wrapper tag from being read as a mix.
|
|
latest_is_mix = (
|
|
isinstance(release_tag, str)
|
|
and latest_build is not None
|
|
and parse_base_build(release_tag) == latest_build
|
|
and release_tag.strip() != f"b{latest_build}"
|
|
)
|
|
if installed_build is None or latest_build is None:
|
|
# Unknown installed/latest version (the involuntary source-build case):
|
|
# treat as behind so we still offer the prebuilt.
|
|
update_available = True
|
|
elif installed_build < latest_build:
|
|
update_available = True
|
|
elif installed_build == latest_build:
|
|
# Same upstream base: offer the extra-patch mix, never a bare rebuild.
|
|
update_available = latest_is_mix
|
|
else:
|
|
# Source build newer than the latest prebuilt: downgrade guard.
|
|
update_available = False
|
|
# Display the mix tag when that's what makes it newer; otherwise the base.
|
|
latest = release_tag if latest_is_mix else base_tag
|
|
# Size of the resolved prebuilt, so source builds show it like the marker
|
|
# path. Fails open to None (offline / asset absent from the release).
|
|
update_size_bytes = None
|
|
if update_available:
|
|
asset_name = res.get("asset")
|
|
if isinstance(asset_name, str) and asset_name:
|
|
try:
|
|
assets = latest_release_assets(res.get("repo"), force_refresh = force_refresh)
|
|
if assets:
|
|
update_size_bytes = assets.get(asset_name)
|
|
except Exception as exc: # pragma: no cover - network defensive
|
|
logger.debug("llama update: source-build size lookup failed", error = str(exc))
|
|
with _job_lock:
|
|
job = dict(_job)
|
|
return {
|
|
"supported": True,
|
|
"update_available": update_available,
|
|
"stale": False,
|
|
"installed_tag": (f"b{installed_build}" if installed_build else None),
|
|
"latest_tag": latest,
|
|
"published_repo": res.get("repo"),
|
|
"installed_at_utc": None,
|
|
"age_days": None,
|
|
"source_build": True,
|
|
"update_size_bytes": update_size_bytes,
|
|
"job": job,
|
|
}
|
|
|
|
|
|
def _active_install_is_local_link(binary: Optional[str]) -> bool:
|
|
"""True when the active llama-server resolves through a --with-llama-cpp-dir
|
|
local link at the canonical llama.cpp directory (see
|
|
update_flow.active_install_is_local_link)."""
|
|
return _flow.active_install_is_local_link(binary, dir_name = "llama.cpp")
|
|
|
|
|
|
def _local_link_status() -> dict:
|
|
"""Status payload for a local-link install: unmanaged, no update offered."""
|
|
return _flow.local_link_status(_job, _job_lock)
|
|
|
|
|
|
def _whisper_chain_status(
|
|
*, force_refresh: bool = False, paired_llama_will_update: bool = False
|
|
) -> Optional[dict]:
|
|
"""Whisper's piggyback plan for the combined update item (see
|
|
whisper_cpp_update.chained_phase_plan). None disables the piggyback --
|
|
fail-open so whisper can never break the llama status or apply."""
|
|
try:
|
|
from utils import whisper_cpp_update
|
|
return whisper_cpp_update.chained_phase_plan(
|
|
force_refresh = force_refresh,
|
|
paired_llama_will_update = paired_llama_will_update,
|
|
)
|
|
except Exception as exc: # pragma: no cover - defensive
|
|
logger.debug("llama update: whisper piggyback probe failed", error = str(exc))
|
|
return None
|
|
|
|
|
|
def _merge_whisper_status(status: dict, *, force_refresh: bool = False) -> dict:
|
|
"""Fold the whisper sub-status into the llama status payload: the llama
|
|
update item is the single UI surface, so update_available becomes the union
|
|
(llama behind OR whisper behind) while llama_update_available keeps the
|
|
llama-only flag. All pre-existing top-level fields are preserved."""
|
|
status["llama_update_available"] = bool(status.get("update_available"))
|
|
plan = _whisper_chain_status(
|
|
force_refresh = force_refresh,
|
|
paired_llama_will_update = status["llama_update_available"],
|
|
)
|
|
if plan is None:
|
|
status["whisper"] = None
|
|
status["update_component"] = "llama" if status["llama_update_available"] else None
|
|
return status
|
|
sub = plan.get("status") or {}
|
|
status["whisper"] = {
|
|
"update_available": bool(plan.get("update_available")),
|
|
"installed_tag": sub.get("installed_tag"),
|
|
"latest_tag": sub.get("latest_tag"),
|
|
"update_size_bytes": sub.get("update_size_bytes"),
|
|
"skip_reason": plan.get("skip_reason"),
|
|
}
|
|
whisper_update_available = bool(plan.get("update_available"))
|
|
if whisper_update_available:
|
|
status["update_available"] = True
|
|
status["update_component"] = (
|
|
"llama"
|
|
if status["llama_update_available"]
|
|
else "whisper"
|
|
if whisper_update_available
|
|
else None
|
|
)
|
|
return status
|
|
|
|
|
|
def get_update_status(*, force_refresh: bool = False) -> dict:
|
|
"""Report whether an update is available plus the current job state.
|
|
|
|
This is the single main update item: llama.cpp drives it and the whisper
|
|
piggyback is folded in (see _merge_whisper_status). force_refresh bypasses
|
|
the 24h release cache for an explicit "check now".
|
|
"""
|
|
status = _llama_only_status(force_refresh = force_refresh)
|
|
return _merge_whisper_status(status, force_refresh = force_refresh)
|
|
|
|
|
|
def _llama_only_status(*, force_refresh: bool = False) -> dict:
|
|
"""The llama.cpp half of get_update_status (no whisper sub-status)."""
|
|
binary = _find_binary()
|
|
# A --with-llama-cpp-dir local link is the user's own tree; never offer to
|
|
# replace it. Bail before any network/freshness work.
|
|
if _active_install_is_local_link(binary):
|
|
return _local_link_status()
|
|
marker = read_install_marker(binary)
|
|
|
|
with _job_lock:
|
|
job_running = _job["state"] == _JOB_RUNNING
|
|
|
|
# No marker = source build / custom path. Offer the official prebuilt if one
|
|
# now exists for this host (this is why macOS source builds showed no button).
|
|
# Skipped while the updater swaps the tree: each 3s poll would exec the
|
|
# half-replaced binary (on Windows that exec can make the installer's
|
|
# os.replace fail) and the poller only consumes job progress.
|
|
if marker is None and binary is not None and not job_running:
|
|
src = _source_build_status(binary, force_refresh = force_refresh)
|
|
if src is not None:
|
|
return src
|
|
|
|
repo = (marker or {}).get("published_repo") or DEFAULT_PUBLISHED_REPO
|
|
|
|
if force_refresh and repo:
|
|
# Prime the cache so the freshness read below sees the newest tag.
|
|
try:
|
|
latest_published_release(repo, force_refresh = True)
|
|
except Exception as exc: # pragma: no cover - network defensive
|
|
logger.debug("llama update: force refresh failed", error = str(exc))
|
|
|
|
freshness = check_prebuilt_freshness(binary)
|
|
installed = freshness.get("installed_tag")
|
|
latest = freshness.get("latest_tag")
|
|
# `behind` compares the full release identity with a base-build guard, so a
|
|
# lagging /releases/latest or a mix-tagged latest can't show a false update
|
|
# (see llama_cpp_freshness.is_behind).
|
|
update_available = bool(freshness.get("has_marker") and freshness.get("behind"))
|
|
|
|
# Size of the prebuilt that Update would download, for the banner. Only when
|
|
# an update is offered; fails open to None (offline / no matching asset).
|
|
update_size_bytes = None
|
|
if update_available:
|
|
try:
|
|
update_size_bytes = update_download_size_bytes(
|
|
marker,
|
|
latest,
|
|
freshness.get("published_repo") or repo,
|
|
force_refresh = force_refresh,
|
|
)
|
|
except Exception as exc: # pragma: no cover - network defensive
|
|
logger.debug("llama update: size lookup failed", error = str(exc))
|
|
|
|
with _job_lock:
|
|
job = dict(_job)
|
|
|
|
return {
|
|
"supported": bool(freshness.get("has_marker")),
|
|
"update_available": update_available,
|
|
"stale": bool(freshness.get("stale")),
|
|
"installed_tag": installed,
|
|
"latest_tag": latest,
|
|
"published_repo": freshness.get("published_repo") or repo,
|
|
"installed_at_utc": freshness.get("installed_at_utc"),
|
|
"age_days": freshness.get("age_days"),
|
|
"source_build": False,
|
|
"update_size_bytes": update_size_bytes,
|
|
"job": job,
|
|
}
|
|
|
|
|
|
def _run_llama_phase(
|
|
install_dir: Path,
|
|
repo: str,
|
|
asset: Optional[str],
|
|
script: Path,
|
|
pin_release_tag: Optional[str],
|
|
set_progress,
|
|
force_cpu: bool = False,
|
|
llama_backend: Optional[str] = None,
|
|
) -> dict:
|
|
"""The llama phase of a chained update: put the backend into a maintenance
|
|
state, run the installer for the latest prebuilt, then refresh caches so the
|
|
next load uses the new build. Returns {to_tag, reload_required, message};
|
|
raises on failure.
|
|
|
|
pin_release_tag pins the installer to that exact published release instead
|
|
of letting it re-resolve "latest" itself (see start_update for why)."""
|
|
backend = None
|
|
model_was_active = False
|
|
try:
|
|
# Block loads and free the binary while the installer swaps it.
|
|
try:
|
|
from routes.inference import get_llama_cpp_backend
|
|
backend = get_llama_cpp_backend()
|
|
except Exception as exc:
|
|
logger.debug(
|
|
"llama update: backend unavailable, skipping load coordination", error = str(exc)
|
|
)
|
|
backend = None
|
|
|
|
if backend is not None:
|
|
try:
|
|
with backend._serial_load_lock:
|
|
backend._llama_update_in_progress = True
|
|
# Active processes can lock the exe on Windows.
|
|
if getattr(backend, "is_active", False):
|
|
model_was_active = True
|
|
backend.unload_model()
|
|
except Exception as exc:
|
|
logger.debug("llama update: load coordination failed", error = str(exc))
|
|
|
|
cmd = [
|
|
sys.executable,
|
|
str(script),
|
|
"--install-dir",
|
|
str(install_dir),
|
|
"--llama-tag",
|
|
"latest",
|
|
"--published-repo",
|
|
repo,
|
|
]
|
|
if pin_release_tag:
|
|
cmd.extend(["--published-release-tag", pin_release_tag])
|
|
cmd.extend(_rocm_install_args(asset))
|
|
# Re-assert a deliberate CPU install (--force-cpu) so detect_host on a GPU host
|
|
# does not re-route to a GPU/Vulkan bundle and revive the crash (#7213). --force-cpu
|
|
# (not --cpu-fallback) also re-persists force_cpu, keeping the choice across future
|
|
# updates. A natural fallback (or a legacy marker without the flag) heals to GPU (#6097).
|
|
if force_cpu:
|
|
cmd.append("--force-cpu")
|
|
if llama_backend == "vulkan":
|
|
cmd.extend(["--llama-backend", "vulkan"])
|
|
logger.info("llama update: installing", cmd = " ".join(cmd))
|
|
env = dict(os.environ, UNSLOTH_PROGRESS_PERCENT_STEP = "5")
|
|
# Preserve a Vulkan install across updates: detect_host on a CUDA/ROCm box would
|
|
# otherwise re-route and silently replace it. Re-assert via setup's env/CLI flags.
|
|
if llama_backend == "vulkan" or (asset and "vulkan" in asset.lower()):
|
|
env["UNSLOTH_FORCE_VULKAN"] = "1"
|
|
env["UNSLOTH_LLAMA_BACKEND"] = "vulkan"
|
|
_flow.stream_installer(
|
|
cmd,
|
|
env,
|
|
set_progress = set_progress,
|
|
timeout_seconds = _INSTALL_TIMEOUT_SECONDS,
|
|
)
|
|
|
|
# Drop stale caches so the banner re-checks the swapped marker.
|
|
# If GitHub is offline, latest stays unknown and the banner fails open.
|
|
reset_caches(drop_disk = True)
|
|
try:
|
|
latest_published_release(repo, force_refresh = True)
|
|
except Exception as exc: # pragma: no cover - network defensive
|
|
logger.debug("llama update: post-install freshness refresh failed", error = str(exc))
|
|
new_marker = read_install_marker(_find_binary())
|
|
new_tag = (new_marker or {}).get("release_tag") or (new_marker or {}).get("tag")
|
|
|
|
# Pinned install must land on that exact release; a same-repo mismatch
|
|
# means the pin was ignored (Vulkan/Intel reroute to another repo is fine).
|
|
if (
|
|
pin_release_tag
|
|
and new_tag
|
|
and (new_marker or {}).get("published_repo") == repo
|
|
and new_tag != pin_release_tag
|
|
):
|
|
raise RuntimeError(f"pinned release {pin_release_tag} but installer produced {new_tag}")
|
|
|
|
logger.info("llama update: success", to_tag = new_tag)
|
|
return {
|
|
"to_tag": new_tag,
|
|
"reload_required": model_was_active,
|
|
"message": (
|
|
f"Updated llama.cpp to {new_tag}."
|
|
+ (" Reload your model to use it." if model_was_active else "")
|
|
),
|
|
}
|
|
except _flow.InstallerExit as exc:
|
|
# Raw "installer exited 4: <log tail>" says nothing actionable in the UI.
|
|
if exc.returncode == _EXIT_NO_SPACE:
|
|
logger.warning("llama update: out of disk space")
|
|
raise RuntimeError(
|
|
"Not enough disk space to install llama.cpp. Free up space or point "
|
|
"UNSLOTH_STUDIO_HOME/TMPDIR at a larger volume, then retry."
|
|
) from exc
|
|
logger.warning("llama update: failed", error = str(exc))
|
|
raise
|
|
except Exception as exc:
|
|
logger.warning("llama update: failed", error = str(exc))
|
|
raise
|
|
finally:
|
|
# Always clear maintenance state.
|
|
if backend is not None:
|
|
try:
|
|
backend._llama_update_in_progress = False
|
|
except Exception: # pragma: no cover - defensive
|
|
pass
|
|
|
|
|
|
# Combined-job progress split when both phases run (download sizes: the llama
|
|
# bundle dwarfs the whisper one); normalized to 0..1 when a phase is skipped.
|
|
_LLAMA_PHASE_WEIGHT = 0.7
|
|
_WHISPER_PHASE_WEIGHT = 0.3
|
|
|
|
|
|
def _plan_llama_phase() -> dict:
|
|
"""Decide how the llama phase of a combined update runs. Returns {"spec"}
|
|
when llama should install, else {"skip_reason", "refusal"}: skip_reason
|
|
marks the phase skipped inside a chained job, refusal is the started=False
|
|
response when the whisper phase has nothing to run either."""
|
|
binary = _find_binary()
|
|
# Refuse to update a --with-llama-cpp-dir local link: installing a prebuilt
|
|
# here would write through the link into the user's own checkout (or fail)
|
|
# and silently drop the link the flag created.
|
|
if _active_install_is_local_link(binary):
|
|
return {
|
|
"skip_reason": "local_link",
|
|
"refusal": {
|
|
"started": False,
|
|
"reason": "local_link",
|
|
"message": (
|
|
"llama.cpp is a local directory linked with --with-llama-cpp-dir; "
|
|
"Unsloth won't replace it. Update your own llama.cpp checkout instead."
|
|
),
|
|
},
|
|
}
|
|
marker = read_install_marker(binary)
|
|
script = _installer_script()
|
|
if script is None:
|
|
return {
|
|
"skip_reason": "installer_missing",
|
|
"refusal": {
|
|
"started": False,
|
|
"reason": "installer_missing",
|
|
"message": "install_llama_prebuilt.py could not be located.",
|
|
},
|
|
}
|
|
|
|
if marker:
|
|
# Mirror the detection guard: a direct POST or a stale banner must not
|
|
# start an install when the latest is not actually newer (force a fresh
|
|
# check so a stale 24h cache can't wrongly block a real update either).
|
|
status = _llama_only_status(force_refresh = True)
|
|
if not status.get("update_available"):
|
|
return {
|
|
"skip_reason": "up_to_date",
|
|
"refusal": {
|
|
"started": False,
|
|
"reason": "up_to_date",
|
|
"message": "The installed llama.cpp build is already at the latest prebuilt.",
|
|
},
|
|
}
|
|
install_dir = _install_dir_for(binary)
|
|
repo = marker.get("published_repo") or DEFAULT_PUBLISHED_REPO
|
|
from_tag = marker.get("tag") or marker.get("release_tag")
|
|
asset = marker.get("asset")
|
|
force_cpu = bool(marker.get("force_cpu"))
|
|
llama_backend = marker.get("llama_backend")
|
|
if llama_backend == "vulkan" or (asset and "vulkan" in str(asset).lower()):
|
|
llama_backend = "vulkan"
|
|
# Install exactly the release the banner offered: the installer's own
|
|
# "latest" is commit-date ordered and can lag the published_at pick
|
|
# above, reinstalling the current build in a loop (the #6219 class).
|
|
# Not on macOS, which needs the older-release walk-back a pin disables
|
|
# (skipping too-new prebuilts); elsewhere an unusable latest now fails
|
|
# the job loudly (retryable) instead of walking back.
|
|
pin_release_tag = None if sys.platform == "darwin" else status.get("latest_tag")
|
|
else:
|
|
# Source build / custom path: only proceed when the same detection logic
|
|
# would offer the update (prebuilt exists, install is behind, root is
|
|
# manageable), so a direct POST cannot downgrade a newer source build.
|
|
src = _source_build_status(binary, force_refresh = True) if binary else None
|
|
if src is None:
|
|
return {
|
|
"skip_reason": "no_prebuilt_available",
|
|
"refusal": {
|
|
"started": False,
|
|
"reason": "no_prebuilt_available",
|
|
"message": (
|
|
"No official llama.cpp prebuilt is available for this host, "
|
|
"so the source build cannot be swapped automatically."
|
|
),
|
|
},
|
|
}
|
|
if not src.get("update_available"):
|
|
return {
|
|
"skip_reason": "up_to_date",
|
|
"refusal": {
|
|
"started": False,
|
|
"reason": "up_to_date",
|
|
"message": (
|
|
"The installed llama.cpp build is already at or newer than the "
|
|
"latest prebuilt."
|
|
),
|
|
},
|
|
}
|
|
res = _resolve_prebuilt_for_host()
|
|
install_dir = _llama_install_root(binary)
|
|
repo = (res or {}).get("repo") or DEFAULT_PUBLISHED_REPO
|
|
from_tag = None
|
|
asset = (res or {}).get("asset")
|
|
# Source builds carry no forced-CPU marker, so nothing to preserve here.
|
|
force_cpu = False
|
|
llama_backend = None
|
|
# No pin: source-build detection resolves via --resolve-prebuilt latest,
|
|
# the same resolver the unpinned apply uses, so the two already agree.
|
|
pin_release_tag = None
|
|
|
|
if install_dir is None:
|
|
return {
|
|
"skip_reason": "no_install_dir",
|
|
"refusal": {
|
|
"started": False,
|
|
"reason": "no_install_dir",
|
|
"message": "Could not determine the llama.cpp install directory.",
|
|
},
|
|
}
|
|
return {
|
|
"spec": {
|
|
"install_dir": install_dir,
|
|
"repo": repo,
|
|
"asset": asset,
|
|
"script": script,
|
|
"pin_release_tag": pin_release_tag,
|
|
"from_tag": from_tag,
|
|
"force_cpu": force_cpu,
|
|
"llama_backend": llama_backend,
|
|
}
|
|
}
|
|
|
|
|
|
def start_update() -> dict:
|
|
"""Kick off a background update job. The job chains the llama phase (the
|
|
existing flow) with a whisper phase that runs only when whisper is actually
|
|
behind; either phase no-ops cleanly when its component is current or
|
|
unmanaged. Idempotent: a second call while one is running returns the
|
|
in-flight job rather than starting another."""
|
|
# A job already in flight wins over any freshness re-check below (and skips
|
|
# its network calls). The final lock block re-checks to close the TOCTOU.
|
|
with _job_lock:
|
|
if _job["state"] == _JOB_RUNNING:
|
|
return {"started": False, "reason": "already_running", "job": dict(_job)}
|
|
|
|
llama_plan = _plan_llama_phase()
|
|
llama_spec = llama_plan.get("spec")
|
|
whisper_plan = _whisper_chain_status(
|
|
force_refresh = True,
|
|
paired_llama_will_update = llama_spec is not None,
|
|
)
|
|
whisper_spec = (whisper_plan or {}).get("phase")
|
|
if llama_spec is None and whisper_spec is None:
|
|
# Nothing to run in either phase: answer with the llama refusal so the
|
|
# existing reasons (local_link / up_to_date / ...) keep their meaning.
|
|
refusal = dict(llama_plan["refusal"])
|
|
with _job_lock:
|
|
refusal["job"] = dict(_job)
|
|
return refusal
|
|
|
|
whisper_run = None
|
|
if whisper_spec is not None:
|
|
from utils import whisper_cpp_update as _whisper
|
|
whisper_run = lambda set_progress: _whisper.run_chained_phase(whisper_spec, set_progress)
|
|
|
|
phases = [
|
|
{
|
|
"name": "llama",
|
|
"weight": _LLAMA_PHASE_WEIGHT,
|
|
"failure_message": "llama.cpp update failed.",
|
|
"skip_reason": llama_plan.get("skip_reason"),
|
|
"run": (
|
|
(
|
|
lambda set_progress: _run_llama_phase(
|
|
llama_spec["install_dir"],
|
|
llama_spec["repo"],
|
|
llama_spec["asset"],
|
|
llama_spec["script"],
|
|
llama_spec["pin_release_tag"],
|
|
set_progress,
|
|
force_cpu = llama_spec.get("force_cpu", False),
|
|
llama_backend = llama_spec.get("llama_backend"),
|
|
)
|
|
)
|
|
if llama_spec
|
|
else None
|
|
),
|
|
},
|
|
{
|
|
"name": "whisper",
|
|
"weight": _WHISPER_PHASE_WEIGHT,
|
|
"failure_message": "whisper.cpp update failed.",
|
|
# The sidecar reload is whisper-internal; it must not trip the
|
|
# job-level reload flag the chat frontend resyncs on.
|
|
"affects_job_reload": False,
|
|
"skip_reason": (whisper_plan or {}).get("skip_reason") or "unavailable",
|
|
"run": whisper_run,
|
|
},
|
|
]
|
|
running = " + ".join(
|
|
name for name, spec in (("llama.cpp", llama_spec), ("whisper.cpp", whisper_spec)) if spec
|
|
)
|
|
|
|
with _job_lock:
|
|
if _job["state"] == _JOB_RUNNING:
|
|
return {"started": False, "reason": "already_running", "job": dict(_job)}
|
|
_job.update(
|
|
state = _JOB_RUNNING,
|
|
message = f"Downloading and installing the latest {running} prebuilt...",
|
|
from_tag = (llama_spec or {}).get("from_tag"),
|
|
to_tag = None,
|
|
reload_required = None,
|
|
error = None,
|
|
progress = 0.0,
|
|
started_at = _utcnow(),
|
|
finished_at = None,
|
|
phases = None,
|
|
)
|
|
job_snapshot = dict(_job)
|
|
|
|
thread = threading.Thread(
|
|
target = _flow.run_chained_update,
|
|
args = (phases,),
|
|
kwargs = {"job": _job, "job_lock": _job_lock},
|
|
name = "llama-cpp-update",
|
|
daemon = True,
|
|
)
|
|
thread.start()
|
|
return {"started": True, "reason": None, "job": job_snapshot}
|
|
|
|
|
|
def _reset_job_for_tests() -> None:
|
|
"""Test-only: return the job tracker to idle."""
|
|
_flow.reset_job(_job, _job_lock)
|