mirror of
https://github.com/unslothai/unsloth.git
synced 2026-08-24 08:13:59 +00:00
* add studio audio page: tts and stt create tab, audio train panel, openai audio endpoints new /audio page mirroring images: create tab with speak (tts via the main inference slot) and transcribe (stt via the dictation sidecars) modes, an always-visible capability line so the loaded model's task is never ambiguous, and a train tab driving the generic /api/train/* audio branches. backend: audio_gallery.py persists tts clips as wav + json sidecar pairs; /v1/audio/speech (openai createSpeech shape, raw wav out) and /v1/audio/transcriptions (multipart, json/text) on the dual-mounted router; gallery list/file/delete/clear on the studio router; the tts core of /audio/generate extracted into _generate_tts_wav so both routes share it and persist clips; keep-warm suffix and transcriptions body cap wired. frontend: AUDIO_CATALOG (orpheus, csm, spark, oute, llasa as tts; whisper and qwen3-asr as stt) painted into the model selector with a per-group task tag, chat-picker speech picks rerouting to /audio, persistent mount in __root, sidebar row under more below video, nav registry + personalization defaults, and the audio nav label in all 12 locales. tests: audio gallery unit tests, speech/transcriptions route round-trips with a faked tts/stt core, middleware body-cap and /v1 surface additions, and the sidebar parity fixtures extended for the audio id. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * studio/audio: fix TTS detection, surface backend capability errors, searchable train pickers Detection: _AUDIO_TOKEN_PATTERNS is first-match-wins, and audio_vlm's generic <|audio|> was tested before the codec fingerprints. Orpheus carries both that token and 28k <custom_token_N> SNAC codes, so it typed as audio_vlm, leaving is_audio False and the Audio page refusing a model that had loaded fine. Codec patterns now go first; Orpheus reports snac. Errors: safe_error_detail flattened "Text-to-speech is not supported on the MLX backend yet" into "An internal error occurred", so a safetensors TTS load on Apple Silicon failed with no reason. Capability answers are now a typed AudioBackendUnsupportedError tagged by the worker and returned as 501 with the message and the GGUF workaround. Tests: /audio/generate persists every clip, so suites driving it with a fake TTS core wrote silent wavs into the real gallery, where the page listed them. An autouse conftest fixture redirects studio_root. Picker: Recommended seeds curated rows in the order given, so a fixed order left every STT row below the fold on Transcribe. The active mode's task now leads. Adds Whisper Tiny/Base, which both sidecars already carry. Audio opts into community models via includeCommunity: non-unsloth TTS/ASR appear in search and trending ones below the unsloth rows, and Search Hub is restored for that case. Train: base model and dataset are search-as-you-type over the Hub with curated entries pinned; default dataset is Etherll/kaira (audio + text columns, no overrides). Panel is sectioned model/data/parameters with a run preview and its left edge tracks the header selector. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * studio/audio: drop the audio train panel, send Train to the Train page Audio fine-tuning was a second, thinner copy of a flow the Train page already owns. Removes the panel and its charts and Hub comboboxes; the Train pill now toasts that Unsloth trains TTS and STT there, given the right base and a dataset with audio plus transcript columns, and navigates to /studio. Also stops TTS picks dead-ending on Mac. Safetensors loads through MLX, which has no TTS branch, so the model loaded and every generation 501'd. A TTS pick with a GGUF build published now loads that instead and says why, since llama.cpp is the only backend carrying the snac/bicodec/dac decoders. Picks with no GGUF still get the 501, which now explains itself. * studio/audio: trim duplicated comments on the GGUF fallback * studio/audio: reword the Train page redirect toast * studio/audio: match the media pane heading treatment, drop the cross-page links Generate audio and Transcribe now use the same heading block as the Images and Video Create panes from #7986: text-xl with leading-none, an 18px icon on the heading line, and a text-xs line under it. Also drops the Images and Video links from the header. They belong between the two visual pages; audio is a different kind of output, so the row was noise here. * studio/audio: per-clip actions menu in History, drop the Audio cross-page links History rows were a single button with no per-row actions, so deleting one clip meant selecting it first and using the player's buttons. Each row now carries a dots menu (use text again, copy text, download WAV, delete), revealed on hover, focus or while open. The row becomes a shell div since the trigger is a button and cannot nest inside one. Downloading from a row fetches the clip bytes on demand: only the selected clip has them cached. Removes the Audio link from the Images and Video headers, matching the Audio page dropping its links to them. * studio/audio: address the Codex review findings Community models in the picker introduced most of these. - Route uncurated ASR picks to the STT sidecar. audioTaskFor returns null for a repo outside AUDIO_CATALOG, so community Whisper repos loaded into the TTS slot. Picks now carry their Hub pipeline tag and fall back to it. - Let community safetensors into Recommended. The curated-artifact clause in keep() can never pass for a community row, so third-party TTS and ASR checkpoints were browse-invisible. Community rows use the rest of the gate. - Feed the community listings into resultGgufIds, so a tag-only GGUF repo opens the variant expander instead of loading as safetensors. - Stop recording when the page goes inactive. The page stays mounted, so the unmount cleanup never ran and the mic stayed hot after navigating away. - Keep generated audio when the gallery write fails. Persistence is best-effort server-side and still returns the WAV; the page now plays it. - Guard gallery pagination with an in-flight flag. Repeated scrolls reused one offset and appended the same page, duplicating clips and React keys. - Select the mtmd engine for Qwen3-ASR in /v1/audio/transcriptions. Whisper ids are shared with the Transformers sidecar, so those keep the default. - Persist TTS clips via asyncio.to_thread, matching the image gallery routes. - Reword the MLX hint: only Orpheus publishes a GGUF build, so it now names the host as the general fix and GGUF as the conditional one. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * studio/audio: fix lifecycle and device inventory * Fix community audio routing and pagination * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Fix fresh audio review findings * Fix audio lifecycle review findings * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Fix remaining audio review findings * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Update Whisper cache inventory contract * Fix final audio convergence findings * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Cancel hidden audio model loads * Cancel hidden transcriptions and bound audio work * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Fix audio model discovery and gallery paging * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Fix concurrent audio cancellation and streaming * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Scope STT startup cancellation to request owner * Stabilize health auth test across hardware states * Scope audio STT lifecycle ownership * Close STT load cancellation race * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Hide unsupported filesystem ASR rows * Fix audio runtime residency edge cases * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Bound stalled audio cancellation * Scope audio load cancellation by request * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Fix audio model handoff routing * Trim redundant audio comments * Update the source and fixture contracts this branch changed The audio page renamed diffusionPageForTask to mediaPageForTask, added isAudioRoute to isChatLike, moved the recommendable-format gate into keep, folded the community listing into the recommended pager, forwards a scoped load_cancel_event through the GGUF loader, and emits cached Whisper repos as ASR rows for the Audio page. Point the exact-source and fixture assertions at the new shape; behaviour is unchanged. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * studio/audio: fix dictation residency, audio dataset decoding and the cancel paths Leo's run on Windows 11 with an RX 9060 XT reported two blockers on PR 7984, and a high-effort review of the branch found ten more defects. Both reports are addressed here. Dictation kept one model resident per engine, so a Transformers Whisper and a llama.cpp Qwen3-ASR sat in VRAM together for the whole 5-minute keep-alive, and a speech model loaded beside a dictation model the Audio page no longer needed. `stt_registry.load` now releases every other engine before loading, under a lock so two loads on different engines cannot interleave, and leaving Transcribe releases the sidecar that tab loaded. Eject and the mode transition share one release path. Audio datasets were unreadable whenever torchcodec cannot dlopen its FFmpeg libraries, which is the Windows default: `disable_torchcodec_if_broken` clears `datasets.config.TORCHCODEC_AVAILABLE`, and datasets 4.x then raises "To support decoding audio data, please install 'torchcodec'" for the format check and all six audio trainer paths. `utils/datasets/audio_decode` installs a soundfile decoder in that case, restoring the pre-4.0 `{"path", "array", "sampling_rate"}` contract those callers already read. `audio_array_and_rate` reads a cell from either backend, which also fixes the `.get("array")` reads that raised AttributeError against the torchcodec AudioDecoder on a working host. Cancelling a GGUF dictation request used to SIGTERM the shared whisper-server, so the next dictation paid a relaunch plus a model load. The sidecar now speaks http.client and shuts the socket instead, sharing `_close_connection_on_cancel` with the mtmd sidecar. `_transcribe_audio_result`'s CancelledError branch no longer calls the lock-taking `cancel_transcription` inline on the event loop. `refreshGallery` replaced the whole clip list with the newest page, collapsing a paginated History and moving the player to a different clip on every delete and generate. It merges the page into the list now and only reselects when the selected clip is gone. A superseded refresh returns the clips its own fetch saw, so a generation whose clip did persist is no longer told it was not saved. The chat picker routed cached repos tagged text-to-speech to /audio with no runtime-support gate, though the Audio page filters exactly those out, and forwarded `meta.pipelineTag` rather than the task that chose the route. Both now go through `audioPickIsRoutable` and `pickedTask`. Also: a drain-cancelled generation raises `AudioGenerationCancelledError` so an idle auto-unload reports 499 rather than a flattened 500; the load path's cancel handshake is bounded, since only the cancelling unload sets it and that unload needs a pool thread of its own; the cache inventory reuses the metadata it already probed; and the Audio page adopts the container-query layout Images and Video use, so the 408px rail stacks below 50rem instead of squeezing the preview. The review also flagged the TTS cancel drain tearing down the worker. Left alone: the stopping criteria is checked per token, not per decode step, and the pre-audio_started teardown is a deliberate invariant that `test_audio_tts_cancellation.py` asserts. Verified on CPU: 794 backend tests over the stt, audio, whisper, inventory and model sweeps, 1591 frontend node tests, typecheck. No GPU here, so the dictation residency and decode fixes still need a run on Leo's ROCm host. * studio/audio: drop the duplicate seed spread the merge left in recommendedMeta * Bound the scoped load cancel handshake _run_tracked_load_model_impl waited on cancel_complete with no timeout. Only /unload's finally sets it for a running attempt, so a disconnect or a shutdown between the cancel and that finally left nobody to set it: /load then parked forever while holding inference_lifecycle_gate, and since asyncio.to_thread runs on non-daemon executor threads the process could not exit either. Reproduced by dropping the handshake and watching both the request and the interpreter hang. Wait 15s, log, and release. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * studio/audio: correct the decode probe, dictation release order and gallery merge Follow-up to the previous commit, from a multi-agent review of its own diff. Seven defects it introduced, each with a reproduction. The decode fix did not fire in the API process. `datasets.config.TORCHCODEC_AVAILABLE` is `find_spec("torchcodec") is not None`, which is true for an installed torchcodec whose native libraries cannot dlopen, and only `unsloth.import_fixes` corrects that. The API process never imports unsloth, so the dataset format check still reached the broken decoder. `ensure_audio_decoding` now probes the import itself and clears the flag, so `datasets`' own gates agree with it. `np.mean(array, axis = tuple(range(array.ndim - 1)))` was copied from datasets' torchcodec shim, which yields (channels, frames); soundfile yields (frames, channels). Every stereo clip collapsed to one sample per channel and trained as near-silence with no warning. Now `axis = -1`, with a test asserting the frame count survives. `Audio.encode_example` needs torchcodec too, and the audio VLM path maps without `remove_columns`, so reading the decoded array writes it back through `cast_storage`. The gate passed and the run then died on the error it was meant to prevent, so encode is patched alongside decode. `MtmdSttSidecar.unload(wait = False)` called `RLock.locked()`, added in Python 3.14, on a 3.10 to 3.13 matrix. It raised, `stt_registry.unload` logged and swallowed it, and the llama-server kept its model: the exact doubling the release exists to prevent. The lock was also the wrong probe, since `transcribe` runs `_post_transcribe` outside it and counts `_active_requests` instead. That is what it checks now. The registry released other engines before the target's preflight, so a 409 for a model that is not downloaded cost the user the engine they were using. `_load_locked` orders preflight ahead of release for that reason; the registry now does too. On the frontend, `owned` tested residency rather than ownership. The activation resync adopts whatever a sidecar holds, including a model chat dictation loaded, so leaving Transcribe could unload it. An explicit `sttLoadedByThisPage` ref now gates the release, and the selection is forgotten only once the unload lands, so a failed unload still has an Eject to retry with. `mergeGalleryPage` stitched unconditionally. "Clear all" merged an empty page into the cache and left every deleted row on screen, and a cache with no ids in common with the page rendered a gap as contiguous with a cursor that could never reach it. It reports whether it stitched, and the cursor is only preserved when it did. Also reverted: `audio_array_and_rate` and its five trainer call sites. `unsloth_zoo.patch_torchcodec_audio_decoder` already gives the torchcodec AudioDecoder a `.get`, and the soundfile decoder returns a plain dict, so the pre-existing reads worked on both backends and the helper was an unrelated refactor. The chat picker now refuses an unrunnable speech pick with a message instead of falling through to a chat load that evicts the resident model. The Audio header pill takes the Images page's `px-3` below 68rem so it stops covering the model name. Tests that passed on deleted code are anchored, `audioPickIsRoutable` gets behavioural cases, and the backend CI job installs soundfile and librosa so the decode module no longer skips. Verified on CPU: 1149 backend tests over the stt, audio, whisper, inventory, monitor, load and admission sweeps, 1600 frontend node tests, typecheck, catalog:check, i18n strict. Two backend failures and 23 in tests/studio/install/test_rocm_support.py reproduce on a clean checkout. No GPU here, so the dictation residency and audio decode paths still need a run on Leo's ROCm host. * studio/audio: drop the duplicate fallthrough in the encode shim * Gate the Audio recorder on the browser capability check Safari and other WebKit builds ship no MediaRecorder, and Studio reached over plain http on a LAN address (-H 0.0.0.0) is not a secure context, so navigator.mediaDevices is undefined in every engine there. The composer already gates its microphone on StudioModelDictationAdapter.isSupported(); the Audio page did not, so Record was enabled and could only ever fail with 'Could not access the microphone', which blames the wrong thing. Reuse the same check, say why in the field hint, and leave file upload available so transcription still works on those hosts. * Fix eight review findings for PR #7984 Backend: - Propagate disconnect cancellation to the base64 JSON transcribe route, so a client that goes away no longer leaves the sidecar transcribing under its lock. - Serialize installation of the soundfile audio decoder. Two first-time callers could both pass the _installed check, and the loser captured the shim as _ORIGINAL_ENCODE, recursing into itself until RecursionError. - Match the GGUF audio read timeout to the exposed token limit instead of a fixed 300s, keeping 300s as the floor. Frontend: - Keep TTS generation running across route changes. Only unmount aborts, matching Images and Video and the note in routes/audio.tsx. - Distinguish a failed gallery refresh from a failed save, so a clip the server did persist is no longer reported as unsaved. - Drop server-deleted clips when merging gallery pages, and do not stitch a page that no longer overlaps the cache. - Preserve Hub evidence (base model, tags, library) when routing community audio picks, so a checkpoint whose family is only in its metadata still routes. - Refresh Audio residency after a global model eject. Tests: regression coverage for the decoder install race and the JSON transcribe request forwarding; restore Audio.encode_example in the decode fixture; skip the decode module without librosa; stub the torchcodec probe so the left-alone case holds on hosts without it; read trainer.py rather than importing the whole torch stack for a source-contract assertion. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Tighten comments in the audio code for PR #7984 Condense 51 multi-line docstrings and JSDoc blocks added by this PR across 24 files, dropping 43 lines while keeping what each one is there to say. Comments, docstrings and whitespace only: no code, signature or behaviour changes. * studio/audio: fetch Spark-TTS into the HF cache, and stop reporting an unreadable repo as non-audio snapshot_download(repo, local_dir = repo.split('/')[-1]) resolves against the process CWD. Under the desktop shell that is studio/src-tauri, so the model landed inside the Tauri crate, the dev watcher rebuilt on every file and killed the backend mid-load. It also bypasses hf_cache_settings, so the copy was invisible to the inventory and re-downloaded per CWD, while the trainer's local_files_only branch was already reading the cache. Dropped local_dir at all four Spark-TTS sites. detect_audio_type folded 'not an audio model' and 'could not read the repo' into a bare None, so a gated repo (401 on tokenizer_config.json) reached the Train page as a definitively non-audio model and the run was refused with 'This model does not support audio'. _detect_audio_from_tokenizer already tracked this correctly; the value was just discarded. Exposed it as detect_audio_type_checked, carried it to /api/models/config as audio_type_known, and the modality gate now blocks only on a known negative. A gated audio repo behaves like a gated text one: the run starts and fails on the real download error. The frontend flag is negative (audioCapabilityUnknown) so an absent value keeps the old blocking behaviour. Also: datasets 4.3 imports torchcodec.encoders at the top of Audio.encode_example before it inspects the value, so the soundfile shim delegating str/Path/bytes to the original re-raised the ImportError it exists to avoid. Those forms need no encoder and are handled directly now. (cherry picked from commit 1b6b6750eb1fad4e804f40091286432c7c0888a4) * studio/audio: summarise decoded audio cells in the dataset preview _serialize_preview_value compressed the undecoded {bytes, path} shape only. When torchcodec cannot load its FFmpeg libraries the soundfile fallback decodes instead and the dataset formatter returns {path, array, sampling_rate} with the waveform as a plain list, so the preview serialised one float per sample. Ten rows of a few seconds each is tens of MB of JSON; the client died with 'Maximum call stack size exceeded' before it could POST /api/train/start, which is why Spark-TTS training never reached the backend on a no-FFmpeg host. A decoded cell now collapses to '<audio, N samples @ R Hz, Ds>' the way a binary cell collapses. Also short-circuit detect_audio_type_checked on a falsy model name. Callers already passed None on every poll, which interpolated into the Hub URL and fetched /None/resolve/main/tokenizer_config.json every few seconds. Previously silent; the new not-definitive log made it visible. (cherry picked from commit 2ec97e54c917e946674fec3f89a3b038ce93b4da) * studio/audio: tag trained checkpoints with their codec and offer them on the Audio page A scan row carried no modality: scan_trained_models returns only (display_name, path, lora|merged) and LoRAInfo had no audio field. So a TTS checkpoint fine-tuned in Studio read as a text model everywhere -- the Audio page's task gate filtered it out, and chat sent it to the GGUF auto-switch, which cannot resolve a local adapter directory and answered 'is not downloaded on this server' for a model sitting in outputs/. /models/loras now reports audio_type, detected from the checkpoint's own tokenizer first (a merged export has one) and falling back to the base repo an adapter names. The Audio page feeds the TTS ones to the picker through the same additionalOnDeviceModels path Transcribe already uses for downloaded STT artifacts, so a checkpoint trained here is selectable where it was trained. Verified against the real run output: the trained Orpheus adapter detects as snac, a text model still reports None. (cherry picked from commit 6259cded9e05abc9c8ed792a0ae114ef6c283976) * studio/audio: label a trained checkpoint by name, not its directory renderAdditionalOnDeviceModelRow always labelled with model.id and linked it to the Hub. That reads fine for a repo id, but a checkpoint trained here is identified by its output directory, so the Audio picker showed two rows of truncated 'C:\Users\...' with a Hub link that goes nowhere. A local path now shows the model name with its base model as the meta line. (cherry picked from commit ba76272b409d5b6b351f182165712bd0c0941850) * studio/audio: stop the TTS watchdog killing a Transformers generation, and name checkpoints Three things from a field run of a trained Orpheus LoRA. Spark-TTS datasets could not be trained at all. The audio text-column allowlist is text/sentence/transcript/transcription/label; every svjack/SparkTTS_* set names the line to speak 'prompt', so detection found no text column, requires_manual_mapping came back True, and the mapping dialog left Continue disabled with no way to satisfy it. Orpheus's dataset uses 'text' and sailed through, which is the whole difference between the two. Added prompt and normalized_text (LJSpeech derivatives). Generation timed out at 120s. That bound only governs the Transformers subprocess path -- llama.cpp TTS never reaches it -- and was tuned against GGUF speeds, where the same clip returns in seconds. A safetensors LoRA needs minutes for it. The worker emits audio_started once and nothing until audio_done, so there is no progress signal to build a stall timeout on; raised the bound instead. A dead worker is already caught every second by _ensure_subprocess_alive, so this only has to bound a live wedged one. The load toast read 'Loading C:\Users\...\outputs\unsloth_orpheus-3b-0.1-ft_1786351654'. A trained checkpoint is identified by its directory, so it now shows the leaf with the training epoch stripped. (cherry picked from commit 00211b548c4a83087e0b889e4cbdd92d2b8ca1bc) * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Fix the chat runtime LoRA type for PR #7984 toLoraSummary reads lora.audio_type, which was added to BackendLoraInfo but not to this function's own parameter type, so the frontend build failed with TS2339 and the Tauri Linux job stopped before the Rust steps. * Fix six review findings for PR #7984 /v1/audio/speech asked for the chat default of 2048 new tokens, which the OpenAI CreateSpeech shape gives a client no way to raise, so any input past roughly half a minute of speech came back as a truncated WAV with HTTP 200. It now asks for AUDIO_GENERATION_MAX_TOKENS, the same ceiling the Audio page's slider uses. The gallery had no retention limit, so an automated client on that route could grow the Studio data directory until the disk filled. Oldest owned pairs beyond UNSLOTH_AUDIO_GALLERY_MAX_CLIPS (default 2000) are now pruned after a save. A TTS cancel arriving before audio_started armed the 5s drain deadline even though _cancel_generation is gated on the worker having started, so the window expired with no cancel ever sent and the teardown unloaded the model the user had just loaded. The pre-start wait now has its own 30s teardown budget, and the 5s drain is armed where the cancel is actually delivered. The worker's audio_error carried no cancelled flag, so a cancellation that sets the worker's shared event without the route's own event (an unload, a training admission, the GPU arbiter) surfaced as HTTP 500 rather than a cancellation. Every repo whose config sniffs as Whisper was un-hidden, but the can_chat guard was keyed on the seven curated ids, so a third-party or fine-tuned Whisper checkpoint stayed eligible for chat auto-load. The Audio page sent temperature on every request, which the backend reads as an explicit client override, so per-model recommendations (Spark-TTS 0.8, OuteTTS 0.4) never applied. It is now sent only once the slider has been moved. The page also loaded at exactly the max-token ceiling, leaving no context for the prompt itself; it now reserves room for both. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Fix two more review findings for PR #7984 The mtmd sidecar's wait=False unload read _active_requests without the lock and then acquired it, so a transcription claiming the slot in between had llama-server killed underneath it and lost the recording. Rechecked under the lock before releasing. Audio detection interpolated a local filesystem path into a Hub URL once the local read found no tokenizer_config.json. The /loras scan hits that for every adapter directory without its own tokenizer, and a transient failure is never cached, so each pass paid two 15s timeouts per checkpoint while blocking the event loop that called it. A local path now stops after the local read. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * studio/audio: load the Spark-TTS tokenizer from LLM/, not the repo root With the dataset gate fixed, a Spark run reaches pre_detect_and_load_tokenizer and dies there: unsloth/Spark-TTS-0.5B keeps only BiCodec/, config.yaml, src/ and wav2vec2-* at its repo root, so AutoTokenizer finds no vocab and raises 'Couldn't instantiate the backend tokenizer ... You need to have sentencepiece or tiktoken installed'. Both are installed; the message sends you after the wrong thing. _load_model already reads weights from LLM/. The tokenizer pre-detect now agrees, via subfolder, and only for a bicodec repo root -- a local checkpoint or an alias that already names LLM/ is left alone. Verified against the real repo: root raises, subfolder='LLM' returns Qwen2Tokenizer with 165158 tokens. * studio/audio: three Poseidon findings -- speech token budget, GGUF read timeout, checkpoint scan P1, /v1/audio/speech was pinned to the 2048 chat default. The route builds a ChatCompletionRequest with no max_tokens, so _tts_max_new_tokens fell through to 'or 2048' and CreateSpeech has no field a client could use to raise it. Anything past roughly half a minute of speech came back as a truncated WAV with HTTP 200 and no signal. Ask for AUDIO_GENERATION_MAX_TOKENS; the orchestrator clamps and scales its watchdog off the same value. My own regression: raising _AUDIO_GENERATION_TIMEOUT to 900s for the Transformers path also moved llama.cpp, which imports _audio_generation_timeout for its GGUF read timeout -- I had claimed llama.cpp never reaches it and was wrong. max(300.0, ...) went dead and every GGUF read got 900s minimum, up to 3600s. Since /audio/speech is in _INFERENCE_SUFFIXES a wedged server holds other_inference_request_count() up for that whole window, blocking idle auto-unload and 409-ing a training start. _audio_generation_timeout takes a base now: 900s subprocess, 300s GGUF. Also mine: _audio_type_of_checkpoint called detect_audio_type with no local_files_only, turning a filesystem scan into N Hub reads per poll, and a non-definitive miss is deliberately uncached so a gated or offline base re-fetched every time. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Resolve the audio timeout base at call time for PR #7984 base defaulted to _AUDIO_GENERATION_TIMEOUT in the signature, so it was bound once at import and reassigning the module constant afterwards had no effect. Resolved inside the function instead. Both backends keep their intended budgets: 900s scaling to 3600s for the Transformers subprocess, 300s scaling to 1200s for GGUF. * Fix four more review findings for PR #7984 /v1/audio/speech is reachable after any /api/inference/load, including the default max_seq_length=0 that becomes 2048, so asking for the full 8192 ceiling overflowed or truncated. Capped to what is left of the loaded context once the prompt is accounted for. A generation whose gallery refresh missed it selected an id that is not in clips, so the player fell through to the empty state and the audio could not be played or downloaded. My earlier fix for the mislabelling caused that. The response WAV is now kept as a fallback until the real record is observed, labelled as saved rather than unsaved. download_status() clears model once the worker thread stops, so a cancellation the user made while the Audio page was hidden matched nothing on return and the deferred load restarted the whole multi-GB download. All three sidecars now report cancelled_model alongside cancelled, which the page matches on. Kept separate from model so the Downloads panel does not start tracking a cancelled download. A locally trained Whisper checkpoint was tagged automatic-speech-recognition and routed to the Audio page, which hands the filesystem path to /audio/stt/load, where resolve_model_id takes only a curated key or an owner/model Hub id and 422s. Local checkpoints no longer get the ASR tag; TTS still routes, since that loads through the main slot, which accepts a local path. * Fix four more review findings for PR #7984 The mtmd sidecar's active-request guard only covered wait=False, but the training VRAM path unloads with wait=True, so llama-server was killed under a live transcription. A blocking unload now drains active requests for up to 30s first, then proceeds so training is not stalled by a long recording. /v1/audio/speech floored an over-context prompt at one output token and forwarded it anyway, failing deep in generation. It now returns 400 while the caller can still shorten the input. Clearing the gallery reset the module cache but not the React clips state, so a failed follow-up refresh left every cleared row rendered against a revoked object URL. Cleared synchronously on the DELETE. The Spark-TTS tokenizer helper treated an LLM/ child as proof the path was already the tokenizer directory, but a cache-pinned or offline snapshot root has one, which is exactly the case needing the subfolder. Only a path ending in LLM, or one carrying its own tokenizer_config.json, skips it now. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Fix three more review findings for PR #7984 selectClip nulls the fallback clip, which undid the setFallbackClip immediately before it, so a clip the server persisted but the refresh missed still rendered the empty state. My earlier fix for that case was a no-op. selectClip now takes keepFallback for the one caller that needs it. Deleting a clip left the row on screen against an already-revoked object URL when the follow-up refresh failed, since refreshGallery returns the cache without calling setClips. The row is dropped on the DELETE now, as clear-all does. A merged Spark-TTS export reached the non-LoRA BiCodec branch, which called snapshot_download on an absolute path and then looked for an LLM/ child that a merged export does not have. It now loads the LLM from the export directory and resolves BiCodec assets from the base model recorded in export_metadata.json, mirroring the processor fallback already in this file. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Fix five more review findings for PR #7984 Switching STT engines loaded the new sidecar while the old one was still resident and released it only afterwards, so a switch could OOM on a device with room for either model alone. The other engines are now released before the allocation, but only once the checkpoint is known to be on disk, so a 409 for a model that was never downloaded still cannot cost the user the engine they were using. _tts_max_new_tokens ignored the prompt, so a Max tokens slider near the ceiling plus a long prompt overflowed the context the page loads with. It now subtracts the prompt from the loaded context, covering both the Studio and OpenAI routes. UNSLOTH_AUDIO_GALLERY_MAX_CLIPS documented that a non-numeric value disables pruning, but the parser restored the 2000-clip default, which would then delete the oldest recordings an operator had asked to keep. A client disconnecting mid-decode was only noticed after PyAV reached EOF or the 30-minute cap. The cancel event is polled in the frame loop now. A microphone recording had no duration or size bound and no timeslice, so an over-long take was buffered whole and uploaded only to be refused. It stops at the sidecar's own limits and reports why. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Fix five more review findings for PR #7984 The GGUF and MTMD sidecars still called the now-cancellable decoder without the event, so only the Transformers path actually stopped on a disconnect. The recorder's byte cap was 96 MiB against the raw route's 25 MiB STT_AUDIO_RAW_MAX_BYTES, so a dense codec could still build a recording that was refused with 413. It mirrors the raw limit now and stops before appending the chunk that would cross it. The TTS prompt reserve used len(prompt) // 3, which under-counts CJK and emoji badly, which is exactly the input that then overflows the context. It asks the loaded tokenizer where one is reachable, and otherwise estimates by character class rather than a flat ratio. Trained TTS checkpoints were offered on macOS even though MLX has no TTS decoder, so selecting one always returned "not supported on the MLX backend yet". Only GGUF exports are listed there now, matching how the catalog rows are filtered. The transcript download revoked its blob URL immediately after the synthetic click, which races browsers that resolve that navigation asynchronously. Deferred, as the gallery download already does. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Fix the CI backend failure and five review findings for PR #7984 The (Python 3.11) Backend tests job was failing on two counts this branch owns. scan_loras probed detect_audio_type without an hf_token, which test_security_gate_consistency forbids because a token-less probe misclassifies a gated model and poisons a token-keyed cache; the route takes the token as a dependency now and threads it through. The health-gate test stubbed _hardware_snapshot as a two-tuple, and main has since added chat_only_detail, so health_check raised IndexError reading snapshot[2] after the merge. Review findings: - /v1/audio/speech preflighted with len(input) // 3 while the budget helper used the tokenizer-aware estimate, so dense text passed the check and was then floored to one output token. Both use the same estimate now, and no budget left is a 400 rather than a one-token clip. - The audio routing evidence map held only remote search results, so a cached community Whisper row picked from the chat picker was judged on its id alone and refused routing to the page that does list it. Cached rows are included now. - Leaving Transcribe fired the sidecar release and returned, so a following TTS load allocated while the sidecar still held its model. The load waits for that teardown. - An audio dataset carrying both an instruction-like prompt column and a real transcript was mapped by schema order, which silently trains ASR against the instructions. Transcript names are matched first, prompt and normalized_text only as a fallback. Also merged current main. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Fix three more review findings for PR #7984 The merged Spark export's BiCodec fetch ran snapshot_download without the request token, so a private or gated base 401'd while the load that followed it would have authenticated fine. Only /v1/audio/speech rejected an over-context prompt; /api/inference/audio/generate floored the budget at one token and generated a clip too short to hold codec tokens. The guard moved into _generate_tts_wav, the core both routes share, so they cannot diverge again. The two route tests that covered it were retargeted at the helper and the shared core, since the route tests fake that core. The response fallback was kept when a refresh missed a persisted clip, as intended, but never cleared once the record arrived. Deleting the now-visible clip then made the fallback reappear from a stale data URL, labelled as saved. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Skip the Transformers TTS cancellation test without the training stack core.inference.inference imports peft transitively, which the backend-test CI job does not install, so this test failed the whole job with ModuleNotFoundError instead of reporting a skip for something that cannot run there. It is the only backend failure in CI that is not also present on main. * Fix /loras audio probe re-walking the cache on every poll for PR #7984 Measured before/after from two isolated installs at the merge base and the head. With 50 trained checkpoints, GET /api/models/loras went 6.0ms -> 26.5ms steady state, +340%, and it runs on the event loop, so it delayed unrelated requests too. The per-checkpoint audio probe answers non-definitively for an adapter directory without its own tokenizer and for a base repo that is not downloaded, and a non-definitive answer is deliberately never cached, so both repeated on every poll. - Remember an offline miss for 60s instead of re-probing. Bounded rather than permanent because both cases can become answerable without a restart: the base gets downloaded, or a training run finishes writing its tokenizer. - Key that on the raw name, before the casing resolution, since resolving a repo id that is not cached walks every HF cache directory, which is the cost itself. - Drop the per-row "could not determine" log to debug when offline. It was one line per checkpoint per poll, and offline it is the ordinary answer. - Run the scan in a worker thread. It was already blocking before the probe was added; the probe made the block long enough to matter. Steady state is now 6.0ms -> 7.3ms, +0.027ms per checkpoint. The latency tail is unchanged: over 400 samples p99 is 103ms before and 121ms after, max 264ms and 260ms, which is this box, not the PR. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Reserve prompt overhead, bound the gallery by bytes, own STT by identity for PR #7984 Three review findings, all confirmed: - The TTS budget was context minus the RAW text, but no backend generates from that: llama_cpp's _TTS_PROMPTS wraps it in codec delimiters and the Transformers path builds its own prompt, so zero headroom meant those tokens pushed prompt plus max_new_tokens back over the context. Reserve 32 tokens for the wrapper. - The gallery cap counted clips, so 2000 maximum-length WAVs was still tens of gigabytes on a route an API client drives. Added a byte quota alongside it, whichever binds first, keeping the newest clip so a single oversized request does not read as a silent failure. - STT ownership was a boolean, so when another surface replaced the sidecar's model while Audio was inactive the activation resync adopted it and Eject unloaded a model this page never loaded. Store the model and engine and require both to match current residency before releasing. Backend 877 passed for the audio, gallery and STT suites; frontend 1716 passed; typecheck clean. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Fix STT engine mismatch, routed pick loss, cached load id and Spark alias for PR #7984 Four review findings, all confirmed: - Keying STT ownership on model AND engine, which I added last commit, broke the fallback case: a "gguf" pick on a host without whisper-server is served by the Transformers fallback and comes back resident under that engine, so the compare never matched and the sidecar was never freed. Key on the model alone. The registry keeps one model resident, and the unload resolves the serving engine server-side already (_resolve_serving_stt_engine). - A pick arriving while a cancelled TTS load was still settling hit the in-flight guard and was dropped, and the route effect had already cleared ?model=, so nothing retried it. Queue the loser and replay it when the load settles. - meta.loadId was discarded, so a row cached in a non-active HF cache was sent as its display repo id: it failed to load offline, or downloaded again into the active cache. Thread it through as the load target, as chat-page.tsx does. - A merged BiCodec export records its base as the registry alias "Spark-TTS-0.5B/LLM", which names a load subdirectory rather than a repo, so snapshot_download rejected it. Resolve it the way the trainer does. Extracted to core/inference/spark_tts_paths.py, a dependency-light leaf, so the mapping is testable without the Unsloth stack. Backend 1303 passed across the audio, gallery, STT, LoRA, capability and Spark suites; frontend 1734 passed; typecheck clean. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Settle a text tokenizer without parsing it, for PR #7984 The cold /api/models/loras path: with 50 checkpoints the first call after a restart was 172ms against 9ms before the PR, and profiling put 84 percent of the scan in the per-checkpoint audio probe. Most of that was json.loads on tokenizer_config files that were never going to match anything. A pattern can only match if its marker text appears in the file at all, so the raw text is scanned for the markers first and an ordinary text checkpoint is settled without a parse. Two details that matter: - The markers cannot be derived from _AUDIO_TOKEN_PATTERNS, which is lambdas, so a codec added there without a marker here would silently stop being detected. A test pins the pattern set and drives every pattern through the marker scan, and it fails when a codec is added. - A marker miss counts as "read" only when the text ends in a closing brace. Without that, a training run part-way through writing its tokenizer would become a definitive "not audio" and be cached for the life of the process, where before it stayed unknown because json.loads raised on the truncated text. Also stops the snac count summing all 28k of Orpheus's codes to answer a question settled by the first 10,001. First call 172ms -> 74ms, of which 27ms is now the scan itself, measured in process on the same fixtures. Steady state is unchanged at +1.4ms. Backend 1152 passed for the detection suites. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Fix Spark-TTS capability detection and a broken whisper runtime for PR #7984 From the Windows/ROCm report on this PR. Spark-TTS training was blocked by the modality gate. Root cause is capability detection, not the gate: the Train page passes the registry alias "Spark-TTS-0.5B/LLM", which names a load subdirectory rather than a repo, so the probe fetched a repo that does not exist, got a 404 on every candidate path, and read that as a DEFINITIVE "not an audio model" rather than "not a repo id". unsloth/Spark-TTS-0.5B -> ('bicodec', True) correct Spark-TTS-0.5B/LLM -> (None, True) wrong, and definitive Resolved through load_scan_target first, the same way routes/training.py already resolves it for the trainer's own preflight. A whisper.cpp build that starts, answers GET /, and then dies on the first inference kept reporting as available: the binary and every linked library are present, so slim_runtime_intact() is satisfied and _resolve_serving_stt_engine never fell back, which left every recording 501-ing next to a loaded chip. Only inference can prove this case, so a failure there now marks the engine unavailable and the existing Transformers fallback takes over. A cancel closes that socket deliberately and is excluded, and a later success clears the flag. Also report the dictation device as rocm rather than cuda on ROCm. Torch keeps the "cuda" device name for HIP, which is right for the API and reads as a bug on an AMD card. Backend: 1324 passed across the STT, audio, capability and model route suites. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Drop Llasa from the audio picker: Studio cannot decode XCodec2, for PR #7984 Swept all 13 curated audio models against a running Studio rather than reading the catalog. Eleven are fine. Two are not: unsloth/Llasa-1B is_audio=False audio_type=None known=True unsloth/Llasa-3B is_audio=False audio_type=None known=True Llasa speaks XCodec2: 65,536 <|s_N|> tokens, confirmed from its tokenizer_config. That is in neither _AUDIO_TOKEN_PATTERNS nor AudioCodecManager, which decodes snac, csm, bicodec and dac only. So the row loaded and then failed at generation with "loaded but is not a supported TTS model", which is the same shape as the Orpheus defect this PR was opened to fix. The policy module already said so and contradicted itself: "the main-slot TTS backend decodes only the four codec families below", above a list of five. Removed Llasa from both the curated catalog and that community family list, so a searched Llasa repo is not admitted either. Studio can still TRAIN Llasa (unsloth_Llasa-3B.yaml is untouched); this catalog only feeds the Generate picker. Re-add both together with an xcodec2 decoder. An existing test asserted a community Llasa row WAS runnable, encoding the same wrong assumption. Corrected it, with the live reading recorded next to it. The other two entries that do not report as audio are the Qwen3-ASR GGUF pair, and that one is expected: a GGUF repo has no tokenizer_config.json at its root, and those models are served by the mtmd sidecar, which routes by catalog engine rather than by this flag. Left alone. Frontend 1736 passed, typecheck and catalog check clean. * Consolidate: one alias resolver, and name the STT tests after their subject Looked for duplication rather than assuming it. There is less than expected: the three STT sidecars share only ~116 near-identical lines between ggml and mtmd and ~26 across all three, and their big methods (_run, load, transcribe) are 25 to 40 percent similar, so a shared base would churn delicate concurrency code to save about 100 lines. Not worth it. Across 309 STT and audio tests exactly one pair is a near-duplicate, and that pair is the same check for two different engines, so it should stay. This PR is large because it does a lot, not because it repeats itself. Two things were worth consolidating. I had added core/inference/spark_tts_paths.py with its own copy of the "Spark-TTS-0.5B/LLM" to "unsloth/Spark-TTS-0.5B" mapping, while the capability probe and routes/training.py both resolve it with load_scan_target. Three copies of one mapping is how they drift. The export path now uses load_scan_target too, and the module and its test file are gone. The test that replaces them reads inference.py as text rather than importing it, because that import pulls the whole Unsloth stack, which was what made a second copy tempting. Four test files were named after the review rounds that produced them, which told a reader when they were written and nothing about what they cover. 105 tests, none duplicated: test_stt_review_fixes_3 + _4 -> test_stt_mtmd_sidecar (43 tests; the mtmd sidecar had no file of its own) test_stt_review_fixes -> merged into test_stt_ggml_sidecar (its subject) test_stt_review_fixes_2 -> test_stt_install_and_snapshot_validation 105 tests before, 105 after, five files down to three, and 1009 passed across the STT, audio, Spark and capability suites. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Fix nine review findings across the audio page and its backend Backend: - audio_decode: pick the token belonging to the repository a URL points at, the way datasets.Audio.decode_example does. Concatenated or interleaved streaming splits carry one token per source repo, so taking an arbitrary value sent one repo's credential to another repo's host. - stt_ggml_sidecar fallback: a GGUF pick downloads one .bin, so when the runtime turns out to be broken at inference time the Transformers engine it is redirected to has no snapshot and every retry raised SttModelNotDownloadedError. Fetch the equivalent snapshot in the background so the promised fallback is real; /audio/stt/status reports its progress. - TTS budget: recheck the prompt against the context after an idle-evicted model is restored. With nothing loaded there is no context to measure, so the first request after an eviction reached generation over-context and came back as a one-token clip. - mtmd unload: the under-lock recheck of _active_requests only guarded wait=False, so a blocking unload could reap llama-server underneath a transcription that started during the acquire. Drain outside the lock and retry, bounded by the same window (draining under the lock would block the request being waited on). - TTS prompt estimate: count every non-ASCII character as a token. The cut at U+2E7F billed Arabic, Cyrillic, Hebrew and the Indic scripts at the Latin third-of-a-token rate, so a long prompt in any of them passed the guard and overflowed during generation. Frontend: - Cancel a deactivating load with the target it was actually started with, not the repo id, so a load keyed on a path or a gguf file is really cancelled. - Replay a queued TTS pick only once the page is active again, and drain the queue on activation. - Abort the TTS handoff when releasing the STT sidecar fails, instead of loading TTS on top of a resident sidecar. - Gallery merge: a complete first page (has_more=false) is everything the server holds, so drop the cached tail rather than rendering clips another client deleted or the size cap pruned. * Move the imports left mid-file by the test merge to the top of the module * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Finish the merge: drop the expander prop main replaced, and type loadTarget #8383 moved GGUF partition identity to the backend and removed filenamePrefix from GgufVariantExpander, so the four call sites this branch touches keep pipelineTag and drop the prop that no longer exists. loadTarget was added to the pendingTtsLoad value but not to the ref's type, which only the project build (tsc -b) catches. * Hold speech generation until the transcribe sidecar release settles Switching straight from Transcribe to Speak with a speech model already resident needs no load, so the gate in the load path never ran and Generate could allocate beside the dictation model. Generate now waits on the same release, and a release that failed puts the page back in Transcribe instead of showing Speak while the sidecar is still resident. * Skip the two torch-dependent STT tests on a runner without torch Both drive a path that reaches `import torch` before the behaviour under test, so on a bare cross-platform runner they failed for a missing wheel rather than for anything they assert. * Pin the language list the GGUF language check reads in its test Without Transformers the helper returns None and the check is skipped, so the request fell through to the download guard and the test asserted nothing on any runner that lacks the dependency. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Give the whisper.cpp discovery tests the filenames Windows actually uses The launcher looks for whisper-server.exe there, and the slim guard checks the libraries the marker names verbatim, so fixtures that hardcoded whisper-server and libggml.so.0 were invisible to both: six tests failed on a Windows runner for the spelling rather than for anything they assert. The product code already handled both platforms. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Keep dictation to one resident engine, and scope releases to the claimed model - Implicit loads on the transcribe routes went straight to the sidecar, and only the registry releases the other engines, so an API client alternating between Qwen3-ASR and Whisper through /v1/audio/transcriptions held both until their independent idle timers fired. Routed through the shared lifecycle, which is a no-op once the model is resident. - Unload now takes the model the caller claims and compares it under the sidecar's own lock. Ownership is decided by the caller, so another surface can switch the same engine before a queued Eject or mode transition lands, and an unscoped release tore down a model the caller never owned. Threaded through the registry, the orchestrator, the route and the Audio page. - A whisper.cpp pick on a host without whisper-server is served and loaded through Transformers, but the residency resolver read only the gguf block, so the refresh that completes the load found nothing and the Transcribe controls stayed disabled until the page was revisited. Same fallback sttEngineStatusFor already applies. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Claim the generation slot before awaiting the transcribe release The Generate button only disables on busy, so awaiting the release first let every click during a slow unload through: each resumed into its own generateAudio while generateAbort tracked only the last, and either finally cleared the busy state out from under the other. --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> Co-authored-by: shimmyshimmer <danielhanchen@gmail.com> Co-authored-by: Etherll <61019402+Etherll@users.noreply.github.com> Co-authored-by: danielhanchen <elliegouldingstuff@gmail.com> Co-authored-by: LeoBorcherding <borchborchmail@gmail.com>
216 lines
8.7 KiB
Python
216 lines
8.7 KiB
Python
# SPDX-License-Identifier: AGPL-3.0-only
|
|
# Copyright 2026-present the Unsloth AI Inc. team. All rights reserved. See /studio/LICENSE.AGPL-3.0
|
|
|
|
"""Infra-only model detection shared by the model routes and the hub
|
|
inventory. Lives directly under ``utils`` (not ``utils.models``) so the hub
|
|
cache scanner can import it without pulling in ``utils/models/__init__.py``,
|
|
which eagerly loads the model-config/checkpoint stack, and without importing
|
|
``routes.models`` (import-time side effects, would cycle)."""
|
|
|
|
from __future__ import annotations
|
|
|
|
import json
|
|
import re
|
|
from pathlib import Path
|
|
from typing import Optional
|
|
|
|
# Hub repo id shape ("owner/name", no leading separator); anything else is
|
|
# treated as a local filesystem path.
|
|
_HF_REPO_ID_RE = re.compile(r"^[A-Za-z0-9][\w.\-]*/[\w.\-]+$")
|
|
|
|
# The llama.cpp install-validation probe repo. Always hidden.
|
|
_PROBE_REPO_ID = "ggml-org/models"
|
|
# The probe's on-disk filename. Carries the ".gguf" so it stays specific and
|
|
# does not hide unrelated repos like ``user/stories260K-finetune-GGUF``.
|
|
_PROBE_FILENAME = "stories260k.gguf"
|
|
# Keep previously cached defaults hidden after settings changes.
|
|
_DEFAULT_EMBEDDING_REPO_IDS = {
|
|
"unsloth/bge-small-en-v1.5",
|
|
"unsloth/bge-small-en-v1.5-GGUF",
|
|
}
|
|
# Local copies do not always retain the repo id. Keep a narrow basename
|
|
# fallback for Studio's static default embedder only; configured custom repos
|
|
# remain exact-match-only.
|
|
_DEFAULT_EMBEDDING_PATH_BASENAMES = {"bge-small-en-v1.5"}
|
|
# Curated dictation checkpoints (STT, never chat), hidden from the chat
|
|
# inventory and pickers: Transformers safetensors repos (unsloth/whisper-*) and
|
|
# their GGUF companions (unslothai/whisper-*-GGUF). Custom checkpoints are caught
|
|
# by config below, but the GGUF companions carry a raw .bin (no config.json), so
|
|
# they must be listed here by id or they leak into chat pickers. The Qwen3-ASR
|
|
# GGUFs are listed for the same reason: llama.cpp will happily load one as a
|
|
# chat model, where it only answers with transcripts.
|
|
_HIDDEN_STT_REPO_IDS = frozenset(
|
|
{
|
|
"unsloth/whisper-tiny",
|
|
"unsloth/whisper-base",
|
|
"unsloth/whisper-small",
|
|
"unsloth/whisper-large-v3-turbo",
|
|
"unsloth/whisper-large-v3",
|
|
"unslothai/whisper-tiny-GGUF",
|
|
"unslothai/whisper-base-GGUF",
|
|
"unslothai/whisper-small-GGUF",
|
|
"unslothai/whisper-large-v3-turbo-GGUF",
|
|
"unslothai/whisper-large-v3-GGUF",
|
|
"unslothai/Qwen3-ASR-0.6B-GGUF",
|
|
"unslothai/Qwen3-ASR-1.7B-GGUF",
|
|
}
|
|
)
|
|
_HIDDEN_STT_REPO_IDS_LOWER = frozenset(repo_id.lower() for repo_id in _HIDDEN_STT_REPO_IDS)
|
|
|
|
|
|
def is_curated_stt_repo_id(value: str | None) -> bool:
|
|
"""True only for Studio's exact curated STT Hub repositories.
|
|
|
|
Still hidden from chat, but task-scoped inventory consumers need the real cache rows
|
|
so the Audio page need not reimplement size, format, variants and lifecycle.
|
|
"""
|
|
return bool(value and value.strip().lower() in _HIDDEN_STT_REPO_IDS_LOWER)
|
|
|
|
|
|
def _config_is_whisper(path: Path) -> bool:
|
|
"""True if a config.json declares a Whisper model."""
|
|
try:
|
|
with open(path, "r", encoding = "utf-8") as file:
|
|
config = json.load(file)
|
|
except Exception:
|
|
return False
|
|
if not isinstance(config, dict):
|
|
return False
|
|
model_type = config.get("model_type")
|
|
if isinstance(model_type, str) and model_type.strip().lower() == "whisper":
|
|
return True
|
|
architectures = config.get("architectures")
|
|
return isinstance(architectures, list) and any(
|
|
isinstance(name, str) and name == "WhisperForConditionalGeneration"
|
|
for name in architectures
|
|
)
|
|
|
|
|
|
def _path_is_whisper_model(value: str) -> bool:
|
|
"""Inspect an existing local model path's config; never hides name-only matches."""
|
|
if _HF_REPO_ID_RE.fullmatch(value.strip()):
|
|
return False
|
|
path = Path(value).expanduser()
|
|
try:
|
|
if path.is_file():
|
|
path = path.parent
|
|
candidates = [path / "config.json"]
|
|
snapshots = path / "snapshots"
|
|
if snapshots.is_dir():
|
|
candidates.extend(child / "config.json" for child in snapshots.iterdir())
|
|
except OSError:
|
|
return False
|
|
return any(_config_is_whisper(candidate) for candidate in candidates)
|
|
|
|
|
|
def _safe_resolve(path: Path) -> Optional[str]:
|
|
"""resolve() to a string, or None when the path is inaccessible."""
|
|
try:
|
|
return str(path.resolve())
|
|
except OSError:
|
|
return None
|
|
|
|
|
|
def _existing_resolved_path(value: str) -> Optional[str]:
|
|
"""Resolve an existing local path."""
|
|
path = Path(value).expanduser()
|
|
try:
|
|
if not path.exists():
|
|
return None
|
|
except OSError:
|
|
return None
|
|
return _safe_resolve(path)
|
|
|
|
|
|
def _path_contains_repo_id(value: str, repo_ids: set[str]) -> bool:
|
|
"""Match exact repo-derived path segments."""
|
|
parts = [part for part in value.lower().replace("\\", "/").split("/") if part]
|
|
for repo_id in repo_ids:
|
|
owner, name = repo_id.split("/", 1)
|
|
if f"models--{owner}--{name}" in parts:
|
|
return True
|
|
if any(
|
|
parts[index] == owner and parts[index + 1] == name for index in range(len(parts) - 1)
|
|
):
|
|
return True
|
|
return False
|
|
|
|
|
|
def _path_basename_is_default_embedder(value: str) -> bool:
|
|
"""Match a default embedder folder or a suffixed local weight filename."""
|
|
normalized = value.lower().replace("\\", "/").rstrip("/")
|
|
basename = normalized.rsplit("/", 1)[-1]
|
|
return any(
|
|
basename == needle
|
|
or any(basename.startswith(f"{needle}{separator}") for separator in ("-", "_", "."))
|
|
for needle in _DEFAULT_EMBEDDING_PATH_BASENAMES
|
|
)
|
|
|
|
|
|
def is_hidden_model(*values: str | None) -> bool:
|
|
"""True if any id/path is the RAG embedding model (the effective embedder
|
|
or its GGUF companion repo), the llama.cpp install validation probe
|
|
(ggml-org/models / stories260K), or a curated/custom Whisper dictation
|
|
model, so pickers hide them (GGUF and non-GGUF). None are usable chat
|
|
models; the probe can be cached as a side effect of installing the prebuilt
|
|
llama-server and otherwise sorts smallest, so it would be auto-selected.
|
|
|
|
Hub repo ids are matched EXACTLY (case-insensitive full "owner/name"), so a
|
|
custom embedder with a generic basename like "org/model" cannot substring
|
|
hide unrelated cached repos such as "user/model-chat" or "org/model-GGUF".
|
|
Existing paths take precedence over the identical ``owner/name`` repo
|
|
shape. Cache and LM Studio paths use exact repo-derived segments. Local
|
|
copies of the static default embedder also use a boundary-aware basename
|
|
fallback; configured custom repos never do."""
|
|
from core.rag import config as rag_config
|
|
|
|
hidden_repo_ids = {
|
|
_PROBE_REPO_ID.lower(),
|
|
*(repo_id.lower() for repo_id in _DEFAULT_EMBEDDING_REPO_IDS),
|
|
*_HIDDEN_STT_REPO_IDS_LOWER,
|
|
}
|
|
exact_paths: list[str] = []
|
|
for model in {
|
|
rag_config.EMBEDDING_MODEL,
|
|
rag_config.default_gguf_repo(),
|
|
rag_config.effective_embedding_model(),
|
|
rag_config.effective_gguf_repo(),
|
|
}:
|
|
existing_path = _existing_resolved_path(model)
|
|
if existing_path:
|
|
exact_paths.append(existing_path.lower())
|
|
elif _HF_REPO_ID_RE.match(model):
|
|
hidden_repo_ids.add(model.lower())
|
|
else:
|
|
resolved = _safe_resolve(Path(model).expanduser())
|
|
if resolved:
|
|
exact_paths.append(resolved.lower())
|
|
for v in values:
|
|
if not v:
|
|
continue
|
|
low = v.lower()
|
|
if _HF_REPO_ID_RE.match(v):
|
|
# A repo id ("owner/name"): match the hidden set exactly. It is
|
|
# never a filesystem path, so skip the path/filename checks.
|
|
if low in hidden_repo_ids:
|
|
return True
|
|
continue
|
|
# Anything else is treated as a filesystem path (the cached snapshot
|
|
# path, or a local model id). Match the probe by its exact filename and
|
|
# any configured local-path embedder by exact resolved path. Split on
|
|
# both separators so a Windows-style path ("...\\stories260K.gguf") is
|
|
# matched even when this runs on a POSIX interpreter (and vice versa).
|
|
if low.replace("\\", "/").rsplit("/", 1)[-1] == _PROBE_FILENAME:
|
|
return True
|
|
if _path_basename_is_default_embedder(v):
|
|
return True
|
|
if _path_contains_repo_id(v, hidden_repo_ids):
|
|
return True
|
|
# Custom Whisper checkpoints keep no curated repo id, so match by config.
|
|
if _path_is_whisper_model(v):
|
|
return True
|
|
if exact_paths:
|
|
resolved = _safe_resolve(Path(v).expanduser())
|
|
if resolved and resolved.lower() in exact_paths:
|
|
return True
|
|
return False
|