mirror of
https://github.com/LostRuins/koboldcpp.git
synced 2026-08-23 07:15:39 +00:00
# Conflicts: # .github/workflows/release.yml # .github/workflows/server.yml # examples/model-conversion/requirements.txt # examples/model-conversion/scripts/causal/run-casual-gen-embeddings-org.py # examples/speculative-simple/README.md # examples/speculative-simple/speculative-simple.cpp # ggml/src/ggml-opencl/ggml-opencl.cpp # requirements/requirements-convert_hf_to_gguf.txt # requirements/requirements-convert_lora_to_gguf.txt # scripts/hip/gcn-cdna-vgpr-check.py # tests/test-backend-ops.cpp # tests/test-chat.cpp # tools/imatrix/imatrix.cpp # tools/mtmd/CMakeLists.txt |
||
|---|---|---|
| .. | ||
| README.md | ||
| tts.cpp | ||
llama.cpp TTS
This is a tool to demonstrate audio generation capability in llama.cpp via libmtmd. It was added via PR #26254
Note: this tool used to serve as a demo for OuteTTS, but it was converted to a more model-agnostic tool.
Common usage
Simple usage:
llama-tts -hf ggml-org/Qwen3-TTS-12Hz-1.7B-Base-GGUF -p "Hello world" --output out.wav
Common params:
- Sampling params such as
--top-k,--top-p,--temp, etc. -n <number_of_frames>limits the output length, e.g.-n 500. Note that how many milliseconds each frame represents varies by model- Core inference params such as
-ngl,-b,-ub, etc.
Qwen3-TTS
Available params:
--tts-langcan bezh,en,de,it,pt,es,ja,ko,fr,ru(default:en)--tts-speaker-fileshould point to a speaker reference audio file (wav, mp3)
Example usage:
llama-tts -hf ggml-org/Qwen3-TTS-12Hz-1.7B-Base-GGUF \
-p "Hello world" \
--tts-lang english \
--tts-speaker-file speaker.mp3 \
--output out.wav
Pocket TTS
Available params:
--tts-speaker-fileshould point to a speaker reference audio file (wav, mp3). It is required, the model produces almost no audio without it- Note:
langis not used, the language is a property of the weights
Example usage:
llama-tts -m pocket-tts.gguf \
-mm mmproj-pocket-tts.gguf \
-p "Hello world" \
--tts-speaker-file speaker.mp3 \
--output out.wav
Note for GGUF conversion:
The upstream repository holds one complete model per language under languages/, next to a set of shared files at the root. Convert one of the languages/<name> directories, not the root directory:
python convert_hf_to_gguf.py path/to/pocket-tts/languages/english --outfile pocket-tts.gguf
python convert_hf_to_gguf.py path/to/pocket-tts/languages/english --mmproj --outfile mmproj-pocket-tts.gguf