unsloth/studio/backend/utils/hidden_models.py
Maheswar Kumar d20db3f1f0
studio: add audio page with tts/stt create tab, train tab, and openai audio endpoints (#7984)
* add studio audio page: tts and stt create tab, audio train panel, openai audio endpoints

new /audio page mirroring images: create tab with speak (tts via the main
inference slot) and transcribe (stt via the dictation sidecars) modes, an
always-visible capability line so the loaded model's task is never ambiguous,
and a train tab driving the generic /api/train/* audio branches.

backend: audio_gallery.py persists tts clips as wav + json sidecar pairs;
/v1/audio/speech (openai createSpeech shape, raw wav out) and
/v1/audio/transcriptions (multipart, json/text) on the dual-mounted router;
gallery list/file/delete/clear on the studio router; the tts core of
/audio/generate extracted into _generate_tts_wav so both routes share it and
persist clips; keep-warm suffix and transcriptions body cap wired.

frontend: AUDIO_CATALOG (orpheus, csm, spark, oute, llasa as tts; whisper and
qwen3-asr as stt) painted into the model selector with a per-group task tag,
chat-picker speech picks rerouting to /audio, persistent mount in __root,
sidebar row under more below video, nav registry + personalization defaults,
and the audio nav label in all 12 locales.

tests: audio gallery unit tests, speech/transcriptions route round-trips with
a faked tts/stt core, middleware body-cap and /v1 surface additions, and the
sidebar parity fixtures extended for the audio id.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* studio/audio: fix TTS detection, surface backend capability errors, searchable train pickers

Detection: _AUDIO_TOKEN_PATTERNS is first-match-wins, and audio_vlm's generic
<|audio|> was tested before the codec fingerprints. Orpheus carries both that
token and 28k <custom_token_N> SNAC codes, so it typed as audio_vlm, leaving
is_audio False and the Audio page refusing a model that had loaded fine. Codec
patterns now go first; Orpheus reports snac.

Errors: safe_error_detail flattened "Text-to-speech is not supported on the MLX
backend yet" into "An internal error occurred", so a safetensors TTS load on
Apple Silicon failed with no reason. Capability answers are now a typed
AudioBackendUnsupportedError tagged by the worker and returned as 501 with the
message and the GGUF workaround.

Tests: /audio/generate persists every clip, so suites driving it with a fake TTS
core wrote silent wavs into the real gallery, where the page listed them. An
autouse conftest fixture redirects studio_root.

Picker: Recommended seeds curated rows in the order given, so a fixed order left
every STT row below the fold on Transcribe. The active mode's task now leads.
Adds Whisper Tiny/Base, which both sidecars already carry. Audio opts into
community models via includeCommunity: non-unsloth TTS/ASR appear in search and
trending ones below the unsloth rows, and Search Hub is restored for that case.

Train: base model and dataset are search-as-you-type over the Hub with curated
entries pinned; default dataset is Etherll/kaira (audio + text columns, no
overrides). Panel is sectioned model/data/parameters with a run preview and its
left edge tracks the header selector.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* studio/audio: drop the audio train panel, send Train to the Train page

Audio fine-tuning was a second, thinner copy of a flow the Train page already
owns. Removes the panel and its charts and Hub comboboxes; the Train pill now
toasts that Unsloth trains TTS and STT there, given the right base and a dataset
with audio plus transcript columns, and navigates to /studio.

Also stops TTS picks dead-ending on Mac. Safetensors loads through MLX, which has
no TTS branch, so the model loaded and every generation 501'd. A TTS pick with a
GGUF build published now loads that instead and says why, since llama.cpp is the
only backend carrying the snac/bicodec/dac decoders. Picks with no GGUF still get
the 501, which now explains itself.

* studio/audio: trim duplicated comments on the GGUF fallback

* studio/audio: reword the Train page redirect toast

* studio/audio: match the media pane heading treatment, drop the cross-page links

Generate audio and Transcribe now use the same heading block as the Images and
Video Create panes from #7986: text-xl with leading-none, an 18px icon on the
heading line, and a text-xs line under it.

Also drops the Images and Video links from the header. They belong between the
two visual pages; audio is a different kind of output, so the row was noise here.

* studio/audio: per-clip actions menu in History, drop the Audio cross-page links

History rows were a single button with no per-row actions, so deleting one clip
meant selecting it first and using the player's buttons. Each row now carries a
dots menu (use text again, copy text, download WAV, delete), revealed on hover,
focus or while open. The row becomes a shell div since the trigger is a button
and cannot nest inside one.

Downloading from a row fetches the clip bytes on demand: only the selected clip
has them cached.

Removes the Audio link from the Images and Video headers, matching the Audio
page dropping its links to them.

* studio/audio: address the Codex review findings

Community models in the picker introduced most of these.

- Route uncurated ASR picks to the STT sidecar. audioTaskFor returns null for a
  repo outside AUDIO_CATALOG, so community Whisper repos loaded into the TTS
  slot. Picks now carry their Hub pipeline tag and fall back to it.
- Let community safetensors into Recommended. The curated-artifact clause in
  keep() can never pass for a community row, so third-party TTS and ASR
  checkpoints were browse-invisible. Community rows use the rest of the gate.
- Feed the community listings into resultGgufIds, so a tag-only GGUF repo opens
  the variant expander instead of loading as safetensors.
- Stop recording when the page goes inactive. The page stays mounted, so the
  unmount cleanup never ran and the mic stayed hot after navigating away.
- Keep generated audio when the gallery write fails. Persistence is best-effort
  server-side and still returns the WAV; the page now plays it.
- Guard gallery pagination with an in-flight flag. Repeated scrolls reused one
  offset and appended the same page, duplicating clips and React keys.
- Select the mtmd engine for Qwen3-ASR in /v1/audio/transcriptions. Whisper ids
  are shared with the Transformers sidecar, so those keep the default.
- Persist TTS clips via asyncio.to_thread, matching the image gallery routes.
- Reword the MLX hint: only Orpheus publishes a GGUF build, so it now names the
  host as the general fix and GGUF as the conditional one.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* studio/audio: fix lifecycle and device inventory

* Fix community audio routing and pagination

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Fix fresh audio review findings

* Fix audio lifecycle review findings

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Fix remaining audio review findings

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Update Whisper cache inventory contract

* Fix final audio convergence findings

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Cancel hidden audio model loads

* Cancel hidden transcriptions and bound audio work

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Fix audio model discovery and gallery paging

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Fix concurrent audio cancellation and streaming

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Scope STT startup cancellation to request owner

* Stabilize health auth test across hardware states

* Scope audio STT lifecycle ownership

* Close STT load cancellation race

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Hide unsupported filesystem ASR rows

* Fix audio runtime residency edge cases

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Bound stalled audio cancellation

* Scope audio load cancellation by request

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Fix audio model handoff routing

* Trim redundant audio comments

* Update the source and fixture contracts this branch changed

The audio page renamed diffusionPageForTask to mediaPageForTask, added
isAudioRoute to isChatLike, moved the recommendable-format gate into keep,
folded the community listing into the recommended pager, forwards a scoped
load_cancel_event through the GGUF loader, and emits cached Whisper repos as
ASR rows for the Audio page. Point the exact-source and fixture assertions at
the new shape; behaviour is unchanged.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* studio/audio: fix dictation residency, audio dataset decoding and the cancel paths

Leo's run on Windows 11 with an RX 9060 XT reported two blockers on PR 7984, and a
high-effort review of the branch found ten more defects. Both reports are addressed
here.

Dictation kept one model resident per engine, so a Transformers Whisper and a
llama.cpp Qwen3-ASR sat in VRAM together for the whole 5-minute keep-alive, and a
speech model loaded beside a dictation model the Audio page no longer needed.
`stt_registry.load` now releases every other engine before loading, under a lock so
two loads on different engines cannot interleave, and leaving Transcribe releases the
sidecar that tab loaded. Eject and the mode transition share one release path.

Audio datasets were unreadable whenever torchcodec cannot dlopen its FFmpeg
libraries, which is the Windows default: `disable_torchcodec_if_broken` clears
`datasets.config.TORCHCODEC_AVAILABLE`, and datasets 4.x then raises "To support
decoding audio data, please install 'torchcodec'" for the format check and all six
audio trainer paths. `utils/datasets/audio_decode` installs a soundfile decoder in
that case, restoring the pre-4.0 `{"path", "array", "sampling_rate"}` contract those
callers already read. `audio_array_and_rate` reads a cell from either backend, which
also fixes the `.get("array")` reads that raised AttributeError against the
torchcodec AudioDecoder on a working host.

Cancelling a GGUF dictation request used to SIGTERM the shared whisper-server, so the
next dictation paid a relaunch plus a model load. The sidecar now speaks http.client
and shuts the socket instead, sharing `_close_connection_on_cancel` with the mtmd
sidecar. `_transcribe_audio_result`'s CancelledError branch no longer calls the
lock-taking `cancel_transcription` inline on the event loop.

`refreshGallery` replaced the whole clip list with the newest page, collapsing a
paginated History and moving the player to a different clip on every delete and
generate. It merges the page into the list now and only reselects when the selected
clip is gone. A superseded refresh returns the clips its own fetch saw, so a
generation whose clip did persist is no longer told it was not saved.

The chat picker routed cached repos tagged text-to-speech to /audio with no
runtime-support gate, though the Audio page filters exactly those out, and forwarded
`meta.pipelineTag` rather than the task that chose the route. Both now go through
`audioPickIsRoutable` and `pickedTask`.

Also: a drain-cancelled generation raises `AudioGenerationCancelledError` so an idle
auto-unload reports 499 rather than a flattened 500; the load path's cancel handshake
is bounded, since only the cancelling unload sets it and that unload needs a pool
thread of its own; the cache inventory reuses the metadata it already probed; and the
Audio page adopts the container-query layout Images and Video use, so the 408px rail
stacks below 50rem instead of squeezing the preview.

The review also flagged the TTS cancel drain tearing down the worker. Left alone: the
stopping criteria is checked per token, not per decode step, and the pre-audio_started
teardown is a deliberate invariant that `test_audio_tts_cancellation.py` asserts.

Verified on CPU: 794 backend tests over the stt, audio, whisper, inventory and model
sweeps, 1591 frontend node tests, typecheck. No GPU here, so the dictation residency
and decode fixes still need a run on Leo's ROCm host.

* studio/audio: drop the duplicate seed spread the merge left in recommendedMeta

* Bound the scoped load cancel handshake

_run_tracked_load_model_impl waited on cancel_complete with no timeout. Only
/unload's finally sets it for a running attempt, so a disconnect or a shutdown
between the cancel and that finally left nobody to set it: /load then parked
forever while holding inference_lifecycle_gate, and since asyncio.to_thread runs
on non-daemon executor threads the process could not exit either. Reproduced by
dropping the handshake and watching both the request and the interpreter hang.

Wait 15s, log, and release.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* studio/audio: correct the decode probe, dictation release order and gallery merge

Follow-up to the previous commit, from a multi-agent review of its own diff. Seven
defects it introduced, each with a reproduction.

The decode fix did not fire in the API process. `datasets.config.TORCHCODEC_AVAILABLE`
is `find_spec("torchcodec") is not None`, which is true for an installed torchcodec
whose native libraries cannot dlopen, and only `unsloth.import_fixes` corrects that.
The API process never imports unsloth, so the dataset format check still reached the
broken decoder. `ensure_audio_decoding` now probes the import itself and clears the
flag, so `datasets`' own gates agree with it.

`np.mean(array, axis = tuple(range(array.ndim - 1)))` was copied from datasets'
torchcodec shim, which yields (channels, frames); soundfile yields (frames, channels).
Every stereo clip collapsed to one sample per channel and trained as near-silence with
no warning. Now `axis = -1`, with a test asserting the frame count survives.

`Audio.encode_example` needs torchcodec too, and the audio VLM path maps without
`remove_columns`, so reading the decoded array writes it back through `cast_storage`.
The gate passed and the run then died on the error it was meant to prevent, so encode
is patched alongside decode.

`MtmdSttSidecar.unload(wait = False)` called `RLock.locked()`, added in Python 3.14,
on a 3.10 to 3.13 matrix. It raised, `stt_registry.unload` logged and swallowed it, and
the llama-server kept its model: the exact doubling the release exists to prevent. The
lock was also the wrong probe, since `transcribe` runs `_post_transcribe` outside it and
counts `_active_requests` instead. That is what it checks now.

The registry released other engines before the target's preflight, so a 409 for a model
that is not downloaded cost the user the engine they were using. `_load_locked` orders
preflight ahead of release for that reason; the registry now does too.

On the frontend, `owned` tested residency rather than ownership. The activation resync
adopts whatever a sidecar holds, including a model chat dictation loaded, so leaving
Transcribe could unload it. An explicit `sttLoadedByThisPage` ref now gates the release,
and the selection is forgotten only once the unload lands, so a failed unload still has
an Eject to retry with.

`mergeGalleryPage` stitched unconditionally. "Clear all" merged an empty page into the
cache and left every deleted row on screen, and a cache with no ids in common with the
page rendered a gap as contiguous with a cursor that could never reach it. It reports
whether it stitched, and the cursor is only preserved when it did.

Also reverted: `audio_array_and_rate` and its five trainer call sites.
`unsloth_zoo.patch_torchcodec_audio_decoder` already gives the torchcodec AudioDecoder
a `.get`, and the soundfile decoder returns a plain dict, so the pre-existing reads
worked on both backends and the helper was an unrelated refactor.

The chat picker now refuses an unrunnable speech pick with a message instead of falling
through to a chat load that evicts the resident model. The Audio header pill takes the
Images page's `px-3` below 68rem so it stops covering the model name. Tests that passed
on deleted code are anchored, `audioPickIsRoutable` gets behavioural cases, and the
backend CI job installs soundfile and librosa so the decode module no longer skips.

Verified on CPU: 1149 backend tests over the stt, audio, whisper, inventory, monitor,
load and admission sweeps, 1600 frontend node tests, typecheck, catalog:check, i18n
strict. Two backend failures and 23 in tests/studio/install/test_rocm_support.py
reproduce on a clean checkout. No GPU here, so the dictation residency and audio decode
paths still need a run on Leo's ROCm host.

* studio/audio: drop the duplicate fallthrough in the encode shim

* Gate the Audio recorder on the browser capability check

Safari and other WebKit builds ship no MediaRecorder, and Studio reached over
plain http on a LAN address (-H 0.0.0.0) is not a secure context, so
navigator.mediaDevices is undefined in every engine there. The composer already
gates its microphone on StudioModelDictationAdapter.isSupported(); the Audio
page did not, so Record was enabled and could only ever fail with 'Could not
access the microphone', which blames the wrong thing.

Reuse the same check, say why in the field hint, and leave file upload
available so transcription still works on those hosts.

* Fix eight review findings for PR #7984

Backend:
- Propagate disconnect cancellation to the base64 JSON transcribe route, so a
  client that goes away no longer leaves the sidecar transcribing under its lock.
- Serialize installation of the soundfile audio decoder. Two first-time callers
  could both pass the _installed check, and the loser captured the shim as
  _ORIGINAL_ENCODE, recursing into itself until RecursionError.
- Match the GGUF audio read timeout to the exposed token limit instead of a fixed
  300s, keeping 300s as the floor.

Frontend:
- Keep TTS generation running across route changes. Only unmount aborts, matching
  Images and Video and the note in routes/audio.tsx.
- Distinguish a failed gallery refresh from a failed save, so a clip the server
  did persist is no longer reported as unsaved.
- Drop server-deleted clips when merging gallery pages, and do not stitch a page
  that no longer overlaps the cache.
- Preserve Hub evidence (base model, tags, library) when routing community audio
  picks, so a checkpoint whose family is only in its metadata still routes.
- Refresh Audio residency after a global model eject.

Tests: regression coverage for the decoder install race and the JSON transcribe
request forwarding; restore Audio.encode_example in the decode fixture; skip the
decode module without librosa; stub the torchcodec probe so the left-alone case
holds on hosts without it; read trainer.py rather than importing the whole
torch stack for a source-contract assertion.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Tighten comments in the audio code for PR #7984

Condense 51 multi-line docstrings and JSDoc blocks added by this PR across 24
files, dropping 43 lines while keeping what each one is there to say. Comments,
docstrings and whitespace only: no code, signature or behaviour changes.

* studio/audio: fetch Spark-TTS into the HF cache, and stop reporting an unreadable repo as non-audio

snapshot_download(repo, local_dir = repo.split('/')[-1]) resolves against the
process CWD. Under the desktop shell that is studio/src-tauri, so the model landed
inside the Tauri crate, the dev watcher rebuilt on every file and killed the backend
mid-load. It also bypasses hf_cache_settings, so the copy was invisible to the
inventory and re-downloaded per CWD, while the trainer's local_files_only branch was
already reading the cache. Dropped local_dir at all four Spark-TTS sites.

detect_audio_type folded 'not an audio model' and 'could not read the repo' into a
bare None, so a gated repo (401 on tokenizer_config.json) reached the Train page as
a definitively non-audio model and the run was refused with 'This model does not
support audio'. _detect_audio_from_tokenizer already tracked this correctly; the
value was just discarded. Exposed it as detect_audio_type_checked, carried it to
/api/models/config as audio_type_known, and the modality gate now blocks only on a
known negative. A gated audio repo behaves like a gated text one: the run starts and
fails on the real download error. The frontend flag is negative (audioCapabilityUnknown)
so an absent value keeps the old blocking behaviour.

Also: datasets 4.3 imports torchcodec.encoders at the top of Audio.encode_example
before it inspects the value, so the soundfile shim delegating str/Path/bytes to the
original re-raised the ImportError it exists to avoid. Those forms need no encoder
and are handled directly now.

(cherry picked from commit 1b6b6750eb1fad4e804f40091286432c7c0888a4)

* studio/audio: summarise decoded audio cells in the dataset preview

_serialize_preview_value compressed the undecoded {bytes, path} shape only. When
torchcodec cannot load its FFmpeg libraries the soundfile fallback decodes instead
and the dataset formatter returns {path, array, sampling_rate} with the waveform as
a plain list, so the preview serialised one float per sample. Ten rows of a few
seconds each is tens of MB of JSON; the client died with 'Maximum call stack size
exceeded' before it could POST /api/train/start, which is why Spark-TTS training
never reached the backend on a no-FFmpeg host. A decoded cell now collapses to
'<audio, N samples @ R Hz, Ds>' the way a binary cell collapses.

Also short-circuit detect_audio_type_checked on a falsy model name. Callers already
passed None on every poll, which interpolated into the Hub URL and fetched
/None/resolve/main/tokenizer_config.json every few seconds. Previously silent; the
new not-definitive log made it visible.

(cherry picked from commit 2ec97e54c917e946674fec3f89a3b038ce93b4da)

* studio/audio: tag trained checkpoints with their codec and offer them on the Audio page

A scan row carried no modality: scan_trained_models returns only
(display_name, path, lora|merged) and LoRAInfo had no audio field. So a TTS
checkpoint fine-tuned in Studio read as a text model everywhere -- the Audio
page's task gate filtered it out, and chat sent it to the GGUF auto-switch,
which cannot resolve a local adapter directory and answered 'is not downloaded
on this server' for a model sitting in outputs/.

/models/loras now reports audio_type, detected from the checkpoint's own
tokenizer first (a merged export has one) and falling back to the base repo an
adapter names. The Audio page feeds the TTS ones to the picker through the same
additionalOnDeviceModels path Transcribe already uses for downloaded STT
artifacts, so a checkpoint trained here is selectable where it was trained.

Verified against the real run output: the trained Orpheus adapter detects as
snac, a text model still reports None.

(cherry picked from commit 6259cded9e05abc9c8ed792a0ae114ef6c283976)

* studio/audio: label a trained checkpoint by name, not its directory

renderAdditionalOnDeviceModelRow always labelled with model.id and linked it to
the Hub. That reads fine for a repo id, but a checkpoint trained here is
identified by its output directory, so the Audio picker showed two rows of
truncated 'C:\Users\...' with a Hub link that goes nowhere. A local path now
shows the model name with its base model as the meta line.

(cherry picked from commit ba76272b409d5b6b351f182165712bd0c0941850)

* studio/audio: stop the TTS watchdog killing a Transformers generation, and name checkpoints

Three things from a field run of a trained Orpheus LoRA.

Spark-TTS datasets could not be trained at all. The audio text-column allowlist is
text/sentence/transcript/transcription/label; every svjack/SparkTTS_* set names the
line to speak 'prompt', so detection found no text column, requires_manual_mapping
came back True, and the mapping dialog left Continue disabled with no way to satisfy
it. Orpheus's dataset uses 'text' and sailed through, which is the whole difference
between the two. Added prompt and normalized_text (LJSpeech derivatives).

Generation timed out at 120s. That bound only governs the Transformers subprocess
path -- llama.cpp TTS never reaches it -- and was tuned against GGUF speeds, where
the same clip returns in seconds. A safetensors LoRA needs minutes for it. The worker
emits audio_started once and nothing until audio_done, so there is no progress signal
to build a stall timeout on; raised the bound instead. A dead worker is already caught
every second by _ensure_subprocess_alive, so this only has to bound a live wedged one.

The load toast read 'Loading C:\Users\...\outputs\unsloth_orpheus-3b-0.1-ft_1786351654'.
A trained checkpoint is identified by its directory, so it now shows the leaf with the
training epoch stripped.

(cherry picked from commit 00211b548c4a83087e0b889e4cbdd92d2b8ca1bc)

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Fix the chat runtime LoRA type for PR #7984

toLoraSummary reads lora.audio_type, which was added to BackendLoraInfo but not
to this function's own parameter type, so the frontend build failed with TS2339
and the Tauri Linux job stopped before the Rust steps.

* Fix six review findings for PR #7984

/v1/audio/speech asked for the chat default of 2048 new tokens, which the OpenAI
CreateSpeech shape gives a client no way to raise, so any input past roughly half
a minute of speech came back as a truncated WAV with HTTP 200. It now asks for
AUDIO_GENERATION_MAX_TOKENS, the same ceiling the Audio page's slider uses.

The gallery had no retention limit, so an automated client on that route could
grow the Studio data directory until the disk filled. Oldest owned pairs beyond
UNSLOTH_AUDIO_GALLERY_MAX_CLIPS (default 2000) are now pruned after a save.

A TTS cancel arriving before audio_started armed the 5s drain deadline even
though _cancel_generation is gated on the worker having started, so the window
expired with no cancel ever sent and the teardown unloaded the model the user had
just loaded. The pre-start wait now has its own 30s teardown budget, and the 5s
drain is armed where the cancel is actually delivered.

The worker's audio_error carried no cancelled flag, so a cancellation that sets
the worker's shared event without the route's own event (an unload, a training
admission, the GPU arbiter) surfaced as HTTP 500 rather than a cancellation.

Every repo whose config sniffs as Whisper was un-hidden, but the can_chat guard
was keyed on the seven curated ids, so a third-party or fine-tuned Whisper
checkpoint stayed eligible for chat auto-load.

The Audio page sent temperature on every request, which the backend reads as an
explicit client override, so per-model recommendations (Spark-TTS 0.8, OuteTTS
0.4) never applied. It is now sent only once the slider has been moved. The page
also loaded at exactly the max-token ceiling, leaving no context for the prompt
itself; it now reserves room for both.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Fix two more review findings for PR #7984

The mtmd sidecar's wait=False unload read _active_requests without the lock and
then acquired it, so a transcription claiming the slot in between had
llama-server killed underneath it and lost the recording. Rechecked under the
lock before releasing.

Audio detection interpolated a local filesystem path into a Hub URL once the
local read found no tokenizer_config.json. The /loras scan hits that for every
adapter directory without its own tokenizer, and a transient failure is never
cached, so each pass paid two 15s timeouts per checkpoint while blocking the
event loop that called it. A local path now stops after the local read.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* studio/audio: load the Spark-TTS tokenizer from LLM/, not the repo root

With the dataset gate fixed, a Spark run reaches pre_detect_and_load_tokenizer and
dies there: unsloth/Spark-TTS-0.5B keeps only BiCodec/, config.yaml, src/ and
wav2vec2-* at its repo root, so AutoTokenizer finds no vocab and raises 'Couldn't
instantiate the backend tokenizer ... You need to have sentencepiece or tiktoken
installed'. Both are installed; the message sends you after the wrong thing.

_load_model already reads weights from LLM/. The tokenizer pre-detect now agrees,
via subfolder, and only for a bicodec repo root -- a local checkpoint or an alias
that already names LLM/ is left alone.

Verified against the real repo: root raises, subfolder='LLM' returns Qwen2Tokenizer
with 165158 tokens.

* studio/audio: three Poseidon findings -- speech token budget, GGUF read timeout, checkpoint scan

P1, /v1/audio/speech was pinned to the 2048 chat default. The route builds a
ChatCompletionRequest with no max_tokens, so _tts_max_new_tokens fell through to
'or 2048' and CreateSpeech has no field a client could use to raise it. Anything
past roughly half a minute of speech came back as a truncated WAV with HTTP 200 and
no signal. Ask for AUDIO_GENERATION_MAX_TOKENS; the orchestrator clamps and scales
its watchdog off the same value.

My own regression: raising _AUDIO_GENERATION_TIMEOUT to 900s for the Transformers
path also moved llama.cpp, which imports _audio_generation_timeout for its GGUF read
timeout -- I had claimed llama.cpp never reaches it and was wrong. max(300.0, ...)
went dead and every GGUF read got 900s minimum, up to 3600s. Since /audio/speech is
in _INFERENCE_SUFFIXES a wedged server holds other_inference_request_count() up for
that whole window, blocking idle auto-unload and 409-ing a training start.
_audio_generation_timeout takes a base now: 900s subprocess, 300s GGUF.

Also mine: _audio_type_of_checkpoint called detect_audio_type with no
local_files_only, turning a filesystem scan into N Hub reads per poll, and a
non-definitive miss is deliberately uncached so a gated or offline base re-fetched
every time.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Resolve the audio timeout base at call time for PR #7984

base defaulted to _AUDIO_GENERATION_TIMEOUT in the signature, so it was bound once
at import and reassigning the module constant afterwards had no effect. Resolved
inside the function instead. Both backends keep their intended budgets: 900s
scaling to 3600s for the Transformers subprocess, 300s scaling to 1200s for GGUF.

* Fix four more review findings for PR #7984

/v1/audio/speech is reachable after any /api/inference/load, including the default
max_seq_length=0 that becomes 2048, so asking for the full 8192 ceiling overflowed
or truncated. Capped to what is left of the loaded context once the prompt is
accounted for.

A generation whose gallery refresh missed it selected an id that is not in clips,
so the player fell through to the empty state and the audio could not be played or
downloaded. My earlier fix for the mislabelling caused that. The response WAV is
now kept as a fallback until the real record is observed, labelled as saved rather
than unsaved.

download_status() clears model once the worker thread stops, so a cancellation the
user made while the Audio page was hidden matched nothing on return and the
deferred load restarted the whole multi-GB download. All three sidecars now report
cancelled_model alongside cancelled, which the page matches on. Kept separate from
model so the Downloads panel does not start tracking a cancelled download.

A locally trained Whisper checkpoint was tagged automatic-speech-recognition and
routed to the Audio page, which hands the filesystem path to /audio/stt/load,
where resolve_model_id takes only a curated key or an owner/model Hub id and 422s.
Local checkpoints no longer get the ASR tag; TTS still routes, since that loads
through the main slot, which accepts a local path.

* Fix four more review findings for PR #7984

The mtmd sidecar's active-request guard only covered wait=False, but the training
VRAM path unloads with wait=True, so llama-server was killed under a live
transcription. A blocking unload now drains active requests for up to 30s first,
then proceeds so training is not stalled by a long recording.

/v1/audio/speech floored an over-context prompt at one output token and forwarded
it anyway, failing deep in generation. It now returns 400 while the caller can
still shorten the input.

Clearing the gallery reset the module cache but not the React clips state, so a
failed follow-up refresh left every cleared row rendered against a revoked object
URL. Cleared synchronously on the DELETE.

The Spark-TTS tokenizer helper treated an LLM/ child as proof the path was already
the tokenizer directory, but a cache-pinned or offline snapshot root has one, which
is exactly the case needing the subfolder. Only a path ending in LLM, or one
carrying its own tokenizer_config.json, skips it now.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Fix three more review findings for PR #7984

selectClip nulls the fallback clip, which undid the setFallbackClip immediately
before it, so a clip the server persisted but the refresh missed still rendered
the empty state. My earlier fix for that case was a no-op. selectClip now takes
keepFallback for the one caller that needs it.

Deleting a clip left the row on screen against an already-revoked object URL when
the follow-up refresh failed, since refreshGallery returns the cache without
calling setClips. The row is dropped on the DELETE now, as clear-all does.

A merged Spark-TTS export reached the non-LoRA BiCodec branch, which called
snapshot_download on an absolute path and then looked for an LLM/ child that a
merged export does not have. It now loads the LLM from the export directory and
resolves BiCodec assets from the base model recorded in export_metadata.json,
mirroring the processor fallback already in this file.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Fix five more review findings for PR #7984

Switching STT engines loaded the new sidecar while the old one was still resident
and released it only afterwards, so a switch could OOM on a device with room for
either model alone. The other engines are now released before the allocation, but
only once the checkpoint is known to be on disk, so a 409 for a model that was
never downloaded still cannot cost the user the engine they were using.

_tts_max_new_tokens ignored the prompt, so a Max tokens slider near the ceiling
plus a long prompt overflowed the context the page loads with. It now subtracts
the prompt from the loaded context, covering both the Studio and OpenAI routes.

UNSLOTH_AUDIO_GALLERY_MAX_CLIPS documented that a non-numeric value disables
pruning, but the parser restored the 2000-clip default, which would then delete
the oldest recordings an operator had asked to keep.

A client disconnecting mid-decode was only noticed after PyAV reached EOF or the
30-minute cap. The cancel event is polled in the frame loop now.

A microphone recording had no duration or size bound and no timeslice, so an
over-long take was buffered whole and uploaded only to be refused. It stops at
the sidecar's own limits and reports why.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Fix five more review findings for PR #7984

The GGUF and MTMD sidecars still called the now-cancellable decoder without the
event, so only the Transformers path actually stopped on a disconnect.

The recorder's byte cap was 96 MiB against the raw route's 25 MiB
STT_AUDIO_RAW_MAX_BYTES, so a dense codec could still build a recording that was
refused with 413. It mirrors the raw limit now and stops before appending the
chunk that would cross it.

The TTS prompt reserve used len(prompt) // 3, which under-counts CJK and emoji
badly, which is exactly the input that then overflows the context. It asks the
loaded tokenizer where one is reachable, and otherwise estimates by character
class rather than a flat ratio.

Trained TTS checkpoints were offered on macOS even though MLX has no TTS decoder,
so selecting one always returned "not supported on the MLX backend yet". Only
GGUF exports are listed there now, matching how the catalog rows are filtered.

The transcript download revoked its blob URL immediately after the synthetic
click, which races browsers that resolve that navigation asynchronously. Deferred,
as the gallery download already does.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Fix the CI backend failure and five review findings for PR #7984

The (Python 3.11) Backend tests job was failing on two counts this branch owns.
scan_loras probed detect_audio_type without an hf_token, which
test_security_gate_consistency forbids because a token-less probe misclassifies a
gated model and poisons a token-keyed cache; the route takes the token as a
dependency now and threads it through. The health-gate test stubbed
_hardware_snapshot as a two-tuple, and main has since added chat_only_detail, so
health_check raised IndexError reading snapshot[2] after the merge.

Review findings:
- /v1/audio/speech preflighted with len(input) // 3 while the budget helper used
  the tokenizer-aware estimate, so dense text passed the check and was then
  floored to one output token. Both use the same estimate now, and no budget left
  is a 400 rather than a one-token clip.
- The audio routing evidence map held only remote search results, so a cached
  community Whisper row picked from the chat picker was judged on its id alone and
  refused routing to the page that does list it. Cached rows are included now.
- Leaving Transcribe fired the sidecar release and returned, so a following TTS
  load allocated while the sidecar still held its model. The load waits for that
  teardown.
- An audio dataset carrying both an instruction-like prompt column and a real
  transcript was mapped by schema order, which silently trains ASR against the
  instructions. Transcript names are matched first, prompt and normalized_text
  only as a fallback.

Also merged current main.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Fix three more review findings for PR #7984

The merged Spark export's BiCodec fetch ran snapshot_download without the request
token, so a private or gated base 401'd while the load that followed it would have
authenticated fine.

Only /v1/audio/speech rejected an over-context prompt; /api/inference/audio/generate
floored the budget at one token and generated a clip too short to hold codec tokens.
The guard moved into _generate_tts_wav, the core both routes share, so they cannot
diverge again. The two route tests that covered it were retargeted at the helper and
the shared core, since the route tests fake that core.

The response fallback was kept when a refresh missed a persisted clip, as intended,
but never cleared once the record arrived. Deleting the now-visible clip then made
the fallback reappear from a stale data URL, labelled as saved.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Skip the Transformers TTS cancellation test without the training stack

core.inference.inference imports peft transitively, which the backend-test CI job
does not install, so this test failed the whole job with ModuleNotFoundError
instead of reporting a skip for something that cannot run there. It is the only
backend failure in CI that is not also present on main.

* Fix /loras audio probe re-walking the cache on every poll for PR #7984

Measured before/after from two isolated installs at the merge base and the head.
With 50 trained checkpoints, GET /api/models/loras went 6.0ms -> 26.5ms steady
state, +340%, and it runs on the event loop, so it delayed unrelated requests too.

The per-checkpoint audio probe answers non-definitively for an adapter directory
without its own tokenizer and for a base repo that is not downloaded, and a
non-definitive answer is deliberately never cached, so both repeated on every poll.

- Remember an offline miss for 60s instead of re-probing. Bounded rather than
  permanent because both cases can become answerable without a restart: the base
  gets downloaded, or a training run finishes writing its tokenizer.
- Key that on the raw name, before the casing resolution, since resolving a repo
  id that is not cached walks every HF cache directory, which is the cost itself.
- Drop the per-row "could not determine" log to debug when offline. It was one
  line per checkpoint per poll, and offline it is the ordinary answer.
- Run the scan in a worker thread. It was already blocking before the probe was
  added; the probe made the block long enough to matter.

Steady state is now 6.0ms -> 7.3ms, +0.027ms per checkpoint. The latency tail is
unchanged: over 400 samples p99 is 103ms before and 121ms after, max 264ms and
260ms, which is this box, not the PR.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Reserve prompt overhead, bound the gallery by bytes, own STT by identity for PR #7984

Three review findings, all confirmed:

- The TTS budget was context minus the RAW text, but no backend generates from that:
  llama_cpp's _TTS_PROMPTS wraps it in codec delimiters and the Transformers path
  builds its own prompt, so zero headroom meant those tokens pushed prompt plus
  max_new_tokens back over the context. Reserve 32 tokens for the wrapper.
- The gallery cap counted clips, so 2000 maximum-length WAVs was still tens of
  gigabytes on a route an API client drives. Added a byte quota alongside it,
  whichever binds first, keeping the newest clip so a single oversized request
  does not read as a silent failure.
- STT ownership was a boolean, so when another surface replaced the sidecar's model
  while Audio was inactive the activation resync adopted it and Eject unloaded a
  model this page never loaded. Store the model and engine and require both to
  match current residency before releasing.

Backend 877 passed for the audio, gallery and STT suites; frontend 1716 passed;
typecheck clean.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Fix STT engine mismatch, routed pick loss, cached load id and Spark alias for PR #7984

Four review findings, all confirmed:

- Keying STT ownership on model AND engine, which I added last commit, broke the
  fallback case: a "gguf" pick on a host without whisper-server is served by the
  Transformers fallback and comes back resident under that engine, so the compare
  never matched and the sidecar was never freed. Key on the model alone. The
  registry keeps one model resident, and the unload resolves the serving engine
  server-side already (_resolve_serving_stt_engine).
- A pick arriving while a cancelled TTS load was still settling hit the in-flight
  guard and was dropped, and the route effect had already cleared ?model=, so
  nothing retried it. Queue the loser and replay it when the load settles.
- meta.loadId was discarded, so a row cached in a non-active HF cache was sent as
  its display repo id: it failed to load offline, or downloaded again into the
  active cache. Thread it through as the load target, as chat-page.tsx does.
- A merged BiCodec export records its base as the registry alias
  "Spark-TTS-0.5B/LLM", which names a load subdirectory rather than a repo, so
  snapshot_download rejected it. Resolve it the way the trainer does. Extracted to
  core/inference/spark_tts_paths.py, a dependency-light leaf, so the mapping is
  testable without the Unsloth stack.

Backend 1303 passed across the audio, gallery, STT, LoRA, capability and Spark
suites; frontend 1734 passed; typecheck clean.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Settle a text tokenizer without parsing it, for PR #7984

The cold /api/models/loras path: with 50 checkpoints the first call after a restart
was 172ms against 9ms before the PR, and profiling put 84 percent of the scan in the
per-checkpoint audio probe. Most of that was json.loads on tokenizer_config files
that were never going to match anything.

A pattern can only match if its marker text appears in the file at all, so the raw
text is scanned for the markers first and an ordinary text checkpoint is settled
without a parse. Two details that matter:

- The markers cannot be derived from _AUDIO_TOKEN_PATTERNS, which is lambdas, so a
  codec added there without a marker here would silently stop being detected. A test
  pins the pattern set and drives every pattern through the marker scan, and it fails
  when a codec is added.
- A marker miss counts as "read" only when the text ends in a closing brace. Without
  that, a training run part-way through writing its tokenizer would become a
  definitive "not audio" and be cached for the life of the process, where before it
  stayed unknown because json.loads raised on the truncated text.

Also stops the snac count summing all 28k of Orpheus's codes to answer a question
settled by the first 10,001.

First call 172ms -> 74ms, of which 27ms is now the scan itself, measured in process
on the same fixtures. Steady state is unchanged at +1.4ms. Backend 1152 passed for
the detection suites.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Fix Spark-TTS capability detection and a broken whisper runtime for PR #7984

From the Windows/ROCm report on this PR.

Spark-TTS training was blocked by the modality gate. Root cause is capability
detection, not the gate: the Train page passes the registry alias
"Spark-TTS-0.5B/LLM", which names a load subdirectory rather than a repo, so the
probe fetched a repo that does not exist, got a 404 on every candidate path, and
read that as a DEFINITIVE "not an audio model" rather than "not a repo id".

  unsloth/Spark-TTS-0.5B  ->  ('bicodec', True)   correct
  Spark-TTS-0.5B/LLM      ->  (None, True)        wrong, and definitive

Resolved through load_scan_target first, the same way routes/training.py already
resolves it for the trainer's own preflight.

A whisper.cpp build that starts, answers GET /, and then dies on the first
inference kept reporting as available: the binary and every linked library are
present, so slim_runtime_intact() is satisfied and _resolve_serving_stt_engine
never fell back, which left every recording 501-ing next to a loaded chip. Only
inference can prove this case, so a failure there now marks the engine unavailable
and the existing Transformers fallback takes over. A cancel closes that socket
deliberately and is excluded, and a later success clears the flag.

Also report the dictation device as rocm rather than cuda on ROCm. Torch keeps the
"cuda" device name for HIP, which is right for the API and reads as a bug on an
AMD card.

Backend: 1324 passed across the STT, audio, capability and model route suites.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Drop Llasa from the audio picker: Studio cannot decode XCodec2, for PR #7984

Swept all 13 curated audio models against a running Studio rather than reading the
catalog. Eleven are fine. Two are not:

  unsloth/Llasa-1B    is_audio=False  audio_type=None  known=True
  unsloth/Llasa-3B    is_audio=False  audio_type=None  known=True

Llasa speaks XCodec2: 65,536 <|s_N|> tokens, confirmed from its tokenizer_config.
That is in neither _AUDIO_TOKEN_PATTERNS nor AudioCodecManager, which decodes snac,
csm, bicodec and dac only. So the row loaded and then failed at generation with
"loaded but is not a supported TTS model", which is the same shape as the Orpheus
defect this PR was opened to fix.

The policy module already said so and contradicted itself: "the main-slot TTS
backend decodes only the four codec families below", above a list of five. Removed
Llasa from both the curated catalog and that community family list, so a searched
Llasa repo is not admitted either.

Studio can still TRAIN Llasa (unsloth_Llasa-3B.yaml is untouched); this catalog only
feeds the Generate picker. Re-add both together with an xcodec2 decoder.

An existing test asserted a community Llasa row WAS runnable, encoding the same
wrong assumption. Corrected it, with the live reading recorded next to it.

The other two entries that do not report as audio are the Qwen3-ASR GGUF pair, and
that one is expected: a GGUF repo has no tokenizer_config.json at its root, and
those models are served by the mtmd sidecar, which routes by catalog engine rather
than by this flag. Left alone.

Frontend 1736 passed, typecheck and catalog check clean.

* Consolidate: one alias resolver, and name the STT tests after their subject

Looked for duplication rather than assuming it. There is less than expected: the
three STT sidecars share only ~116 near-identical lines between ggml and mtmd and
~26 across all three, and their big methods (_run, load, transcribe) are 25 to 40
percent similar, so a shared base would churn delicate concurrency code to save
about 100 lines. Not worth it. Across 309 STT and audio tests exactly one pair is
a near-duplicate, and that pair is the same check for two different engines, so it
should stay. This PR is large because it does a lot, not because it repeats itself.

Two things were worth consolidating.

I had added core/inference/spark_tts_paths.py with its own copy of the
"Spark-TTS-0.5B/LLM" to "unsloth/Spark-TTS-0.5B" mapping, while the capability
probe and routes/training.py both resolve it with load_scan_target. Three copies
of one mapping is how they drift. The export path now uses load_scan_target too,
and the module and its test file are gone. The test that replaces them reads
inference.py as text rather than importing it, because that import pulls the whole
Unsloth stack, which was what made a second copy tempting.

Four test files were named after the review rounds that produced them, which told a
reader when they were written and nothing about what they cover. 105 tests, none
duplicated:

  test_stt_review_fixes_3 + _4  ->  test_stt_mtmd_sidecar  (43 tests; the mtmd
                                    sidecar had no file of its own)
  test_stt_review_fixes         ->  merged into test_stt_ggml_sidecar (its subject)
  test_stt_review_fixes_2       ->  test_stt_install_and_snapshot_validation

105 tests before, 105 after, five files down to three, and 1009 passed across the
STT, audio, Spark and capability suites.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Fix nine review findings across the audio page and its backend

Backend:
- audio_decode: pick the token belonging to the repository a URL points at,
  the way datasets.Audio.decode_example does. Concatenated or interleaved
  streaming splits carry one token per source repo, so taking an arbitrary
  value sent one repo's credential to another repo's host.
- stt_ggml_sidecar fallback: a GGUF pick downloads one .bin, so when the
  runtime turns out to be broken at inference time the Transformers engine it
  is redirected to has no snapshot and every retry raised
  SttModelNotDownloadedError. Fetch the equivalent snapshot in the background
  so the promised fallback is real; /audio/stt/status reports its progress.
- TTS budget: recheck the prompt against the context after an idle-evicted
  model is restored. With nothing loaded there is no context to measure, so
  the first request after an eviction reached generation over-context and came
  back as a one-token clip.
- mtmd unload: the under-lock recheck of _active_requests only guarded
  wait=False, so a blocking unload could reap llama-server underneath a
  transcription that started during the acquire. Drain outside the lock and
  retry, bounded by the same window (draining under the lock would block the
  request being waited on).
- TTS prompt estimate: count every non-ASCII character as a token. The cut at
  U+2E7F billed Arabic, Cyrillic, Hebrew and the Indic scripts at the Latin
  third-of-a-token rate, so a long prompt in any of them passed the guard and
  overflowed during generation.

Frontend:
- Cancel a deactivating load with the target it was actually started with, not
  the repo id, so a load keyed on a path or a gguf file is really cancelled.
- Replay a queued TTS pick only once the page is active again, and drain the
  queue on activation.
- Abort the TTS handoff when releasing the STT sidecar fails, instead of
  loading TTS on top of a resident sidecar.
- Gallery merge: a complete first page (has_more=false) is everything the
  server holds, so drop the cached tail rather than rendering clips another
  client deleted or the size cap pruned.

* Move the imports left mid-file by the test merge to the top of the module

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Finish the merge: drop the expander prop main replaced, and type loadTarget

#8383 moved GGUF partition identity to the backend and removed
filenamePrefix from GgufVariantExpander, so the four call sites this branch
touches keep pipelineTag and drop the prop that no longer exists.

loadTarget was added to the pendingTtsLoad value but not to the ref's type,
which only the project build (tsc -b) catches.

* Hold speech generation until the transcribe sidecar release settles

Switching straight from Transcribe to Speak with a speech model already
resident needs no load, so the gate in the load path never ran and Generate
could allocate beside the dictation model. Generate now waits on the same
release, and a release that failed puts the page back in Transcribe instead of
showing Speak while the sidecar is still resident.

* Skip the two torch-dependent STT tests on a runner without torch

Both drive a path that reaches `import torch` before the behaviour under test,
so on a bare cross-platform runner they failed for a missing wheel rather than
for anything they assert.

* Pin the language list the GGUF language check reads in its test

Without Transformers the helper returns None and the check is skipped, so the
request fell through to the download guard and the test asserted nothing on any
runner that lacks the dependency.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Give the whisper.cpp discovery tests the filenames Windows actually uses

The launcher looks for whisper-server.exe there, and the slim guard checks the
libraries the marker names verbatim, so fixtures that hardcoded whisper-server
and libggml.so.0 were invisible to both: six tests failed on a Windows runner
for the spelling rather than for anything they assert. The product code already
handled both platforms.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Keep dictation to one resident engine, and scope releases to the claimed model

- Implicit loads on the transcribe routes went straight to the sidecar, and only
  the registry releases the other engines, so an API client alternating between
  Qwen3-ASR and Whisper through /v1/audio/transcriptions held both until their
  independent idle timers fired. Routed through the shared lifecycle, which is a
  no-op once the model is resident.
- Unload now takes the model the caller claims and compares it under the
  sidecar's own lock. Ownership is decided by the caller, so another surface can
  switch the same engine before a queued Eject or mode transition lands, and an
  unscoped release tore down a model the caller never owned. Threaded through the
  registry, the orchestrator, the route and the Audio page.
- A whisper.cpp pick on a host without whisper-server is served and loaded
  through Transformers, but the residency resolver read only the gguf block, so
  the refresh that completes the load found nothing and the Transcribe controls
  stayed disabled until the page was revisited. Same fallback sttEngineStatusFor
  already applies.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Claim the generation slot before awaiting the transcribe release

The Generate button only disables on busy, so awaiting the release first let
every click during a slow unload through: each resumed into its own
generateAudio while generateAbort tracked only the last, and either finally
cleared the busy state out from under the other.

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
Co-authored-by: shimmyshimmer <danielhanchen@gmail.com>
Co-authored-by: Etherll <61019402+Etherll@users.noreply.github.com>
Co-authored-by: danielhanchen <elliegouldingstuff@gmail.com>
Co-authored-by: LeoBorcherding <borchborchmail@gmail.com>
2026-08-11 04:47:26 -07:00

216 lines
8.7 KiB
Python

# SPDX-License-Identifier: AGPL-3.0-only
# Copyright 2026-present the Unsloth AI Inc. team. All rights reserved. See /studio/LICENSE.AGPL-3.0
"""Infra-only model detection shared by the model routes and the hub
inventory. Lives directly under ``utils`` (not ``utils.models``) so the hub
cache scanner can import it without pulling in ``utils/models/__init__.py``,
which eagerly loads the model-config/checkpoint stack, and without importing
``routes.models`` (import-time side effects, would cycle)."""
from __future__ import annotations
import json
import re
from pathlib import Path
from typing import Optional
# Hub repo id shape ("owner/name", no leading separator); anything else is
# treated as a local filesystem path.
_HF_REPO_ID_RE = re.compile(r"^[A-Za-z0-9][\w.\-]*/[\w.\-]+$")
# The llama.cpp install-validation probe repo. Always hidden.
_PROBE_REPO_ID = "ggml-org/models"
# The probe's on-disk filename. Carries the ".gguf" so it stays specific and
# does not hide unrelated repos like ``user/stories260K-finetune-GGUF``.
_PROBE_FILENAME = "stories260k.gguf"
# Keep previously cached defaults hidden after settings changes.
_DEFAULT_EMBEDDING_REPO_IDS = {
"unsloth/bge-small-en-v1.5",
"unsloth/bge-small-en-v1.5-GGUF",
}
# Local copies do not always retain the repo id. Keep a narrow basename
# fallback for Studio's static default embedder only; configured custom repos
# remain exact-match-only.
_DEFAULT_EMBEDDING_PATH_BASENAMES = {"bge-small-en-v1.5"}
# Curated dictation checkpoints (STT, never chat), hidden from the chat
# inventory and pickers: Transformers safetensors repos (unsloth/whisper-*) and
# their GGUF companions (unslothai/whisper-*-GGUF). Custom checkpoints are caught
# by config below, but the GGUF companions carry a raw .bin (no config.json), so
# they must be listed here by id or they leak into chat pickers. The Qwen3-ASR
# GGUFs are listed for the same reason: llama.cpp will happily load one as a
# chat model, where it only answers with transcripts.
_HIDDEN_STT_REPO_IDS = frozenset(
{
"unsloth/whisper-tiny",
"unsloth/whisper-base",
"unsloth/whisper-small",
"unsloth/whisper-large-v3-turbo",
"unsloth/whisper-large-v3",
"unslothai/whisper-tiny-GGUF",
"unslothai/whisper-base-GGUF",
"unslothai/whisper-small-GGUF",
"unslothai/whisper-large-v3-turbo-GGUF",
"unslothai/whisper-large-v3-GGUF",
"unslothai/Qwen3-ASR-0.6B-GGUF",
"unslothai/Qwen3-ASR-1.7B-GGUF",
}
)
_HIDDEN_STT_REPO_IDS_LOWER = frozenset(repo_id.lower() for repo_id in _HIDDEN_STT_REPO_IDS)
def is_curated_stt_repo_id(value: str | None) -> bool:
"""True only for Studio's exact curated STT Hub repositories.
Still hidden from chat, but task-scoped inventory consumers need the real cache rows
so the Audio page need not reimplement size, format, variants and lifecycle.
"""
return bool(value and value.strip().lower() in _HIDDEN_STT_REPO_IDS_LOWER)
def _config_is_whisper(path: Path) -> bool:
"""True if a config.json declares a Whisper model."""
try:
with open(path, "r", encoding = "utf-8") as file:
config = json.load(file)
except Exception:
return False
if not isinstance(config, dict):
return False
model_type = config.get("model_type")
if isinstance(model_type, str) and model_type.strip().lower() == "whisper":
return True
architectures = config.get("architectures")
return isinstance(architectures, list) and any(
isinstance(name, str) and name == "WhisperForConditionalGeneration"
for name in architectures
)
def _path_is_whisper_model(value: str) -> bool:
"""Inspect an existing local model path's config; never hides name-only matches."""
if _HF_REPO_ID_RE.fullmatch(value.strip()):
return False
path = Path(value).expanduser()
try:
if path.is_file():
path = path.parent
candidates = [path / "config.json"]
snapshots = path / "snapshots"
if snapshots.is_dir():
candidates.extend(child / "config.json" for child in snapshots.iterdir())
except OSError:
return False
return any(_config_is_whisper(candidate) for candidate in candidates)
def _safe_resolve(path: Path) -> Optional[str]:
"""resolve() to a string, or None when the path is inaccessible."""
try:
return str(path.resolve())
except OSError:
return None
def _existing_resolved_path(value: str) -> Optional[str]:
"""Resolve an existing local path."""
path = Path(value).expanduser()
try:
if not path.exists():
return None
except OSError:
return None
return _safe_resolve(path)
def _path_contains_repo_id(value: str, repo_ids: set[str]) -> bool:
"""Match exact repo-derived path segments."""
parts = [part for part in value.lower().replace("\\", "/").split("/") if part]
for repo_id in repo_ids:
owner, name = repo_id.split("/", 1)
if f"models--{owner}--{name}" in parts:
return True
if any(
parts[index] == owner and parts[index + 1] == name for index in range(len(parts) - 1)
):
return True
return False
def _path_basename_is_default_embedder(value: str) -> bool:
"""Match a default embedder folder or a suffixed local weight filename."""
normalized = value.lower().replace("\\", "/").rstrip("/")
basename = normalized.rsplit("/", 1)[-1]
return any(
basename == needle
or any(basename.startswith(f"{needle}{separator}") for separator in ("-", "_", "."))
for needle in _DEFAULT_EMBEDDING_PATH_BASENAMES
)
def is_hidden_model(*values: str | None) -> bool:
"""True if any id/path is the RAG embedding model (the effective embedder
or its GGUF companion repo), the llama.cpp install validation probe
(ggml-org/models / stories260K), or a curated/custom Whisper dictation
model, so pickers hide them (GGUF and non-GGUF). None are usable chat
models; the probe can be cached as a side effect of installing the prebuilt
llama-server and otherwise sorts smallest, so it would be auto-selected.
Hub repo ids are matched EXACTLY (case-insensitive full "owner/name"), so a
custom embedder with a generic basename like "org/model" cannot substring
hide unrelated cached repos such as "user/model-chat" or "org/model-GGUF".
Existing paths take precedence over the identical ``owner/name`` repo
shape. Cache and LM Studio paths use exact repo-derived segments. Local
copies of the static default embedder also use a boundary-aware basename
fallback; configured custom repos never do."""
from core.rag import config as rag_config
hidden_repo_ids = {
_PROBE_REPO_ID.lower(),
*(repo_id.lower() for repo_id in _DEFAULT_EMBEDDING_REPO_IDS),
*_HIDDEN_STT_REPO_IDS_LOWER,
}
exact_paths: list[str] = []
for model in {
rag_config.EMBEDDING_MODEL,
rag_config.default_gguf_repo(),
rag_config.effective_embedding_model(),
rag_config.effective_gguf_repo(),
}:
existing_path = _existing_resolved_path(model)
if existing_path:
exact_paths.append(existing_path.lower())
elif _HF_REPO_ID_RE.match(model):
hidden_repo_ids.add(model.lower())
else:
resolved = _safe_resolve(Path(model).expanduser())
if resolved:
exact_paths.append(resolved.lower())
for v in values:
if not v:
continue
low = v.lower()
if _HF_REPO_ID_RE.match(v):
# A repo id ("owner/name"): match the hidden set exactly. It is
# never a filesystem path, so skip the path/filename checks.
if low in hidden_repo_ids:
return True
continue
# Anything else is treated as a filesystem path (the cached snapshot
# path, or a local model id). Match the probe by its exact filename and
# any configured local-path embedder by exact resolved path. Split on
# both separators so a Windows-style path ("...\\stories260K.gguf") is
# matched even when this runs on a POSIX interpreter (and vice versa).
if low.replace("\\", "/").rsplit("/", 1)[-1] == _PROBE_FILENAME:
return True
if _path_basename_is_default_embedder(v):
return True
if _path_contains_repo_id(v, hidden_repo_ids):
return True
# Custom Whisper checkpoints keep no curated repo id, so match by config.
if _path_is_whisper_model(v):
return True
if exact_paths:
resolved = _safe_resolve(Path(v).expanduser())
if resolved and resolved.lower() in exact_paths:
return True
return False