mirror of
https://github.com/LostRuins/koboldcpp.git
synced 2026-05-20 09:25:53 +00:00
* mtmd : add MERaLiON-2 multimodal audio support Adds support for A*STAR's MERaLiON-2 audio-language model (3B and 10B) to the multimodal framework. Architecture: - Whisper large-v2 encoder for audio feature extraction - Gated MLP adaptor: ln_speech -> frame stack (x15) -> Linear+SiLU -> GLU -> out_proj - Gemma2 3B / 27B decoder The mmproj GGUF is generated via convert_hf_to_gguf.py --mmproj on the full MERaLiON-2 model directory (architecture: MERaLiON2ForConditionalGeneration). The decoder is converted separately as a standard Gemma2 model after stripping the text_decoder. weight prefix. New projector type: PROJECTOR_TYPE_MERALION Supports tasks: speech transcription (EN/ZH/MS/TA), translation, spoken QA. Model: https://huggingface.co/MERaLiON/MERaLiON-2-3B https://huggingface.co/MERaLiON/MERaLiON-2-10B * simplify comments in meralion adaptor * meralion: use format_tensor_name, ascii arrows in comments |
||
|---|---|---|
| .. | ||
| cogvlm.cpp | ||
| conformer.cpp | ||
| deepseekocr.cpp | ||
| dotsocr.cpp | ||
| gemma4v.cpp | ||
| glm4v.cpp | ||
| hunyuanocr.cpp | ||
| internvl.cpp | ||
| kimik25.cpp | ||
| kimivl.cpp | ||
| llama4.cpp | ||
| llava.cpp | ||
| minicpmv.cpp | ||
| mobilenetv5.cpp | ||
| models.h | ||
| nemotron-v2-vl.cpp | ||
| paddleocr.cpp | ||
| pixtral.cpp | ||
| qwen2vl.cpp | ||
| qwen3vl.cpp | ||
| siglip.cpp | ||
| step3vl.cpp | ||
| whisper-enc.cpp | ||
| youtuvl.cpp | ||