mirror of
https://github.com/LostRuins/koboldcpp.git
synced 2026-08-30 02:34:00 +00:00
* Get started with Onyx
* Add architecture
* Skip keys handled in super()
* Loading tensors
* Shorten
* Graph
* Apply suggestion from @pcuenca
* Remove norm now embedding in transformers weights
* Add eot
* Explicit output_multiplier
* Handle post_norm_eps
* No super call; unhardcode eot.
The pattern `self._set_vocab_gpt2()` seems preferred throughout the
codebase, and it allows `set_vocab()` to be called from a different part
of the Python class hierarchy: the drafter model converter that we may
need eventually.
* Register for drafting
* DFlash: inherit rope type from the linked target.
Another option would be to store it in the gguf file itself.
* mmproj conversion
Note: some fields to be renamed after the implementation works. We are
keeping compatibility with the reference Meta gguf for testing purposes.
* "clip" header declarations
* Load mmproj
* Pre-processing
* Graph
* Go back to using delimiters.
Otherwise our generations are worse.
Transformers does not use them. We need to trace inputs to verify
whether they are equivalent.
* downsample_factor -> merge_size
* Add vision graph
lol, forgot from a previous commit
* Additional renames, align with llama.cpp / transformers
* Prefer _size instead of independent _h and _w
* Fix token layout
Co-authored-by: Young Han <younghan@fb.com>
* onyx: bring the chat parser onto the onyx branch
common/chat.cpp on this branch has no Onyx handling, so a converted model
serves malformed chat: the assistant preamble leaks into content
("to=self<|message|>...") and tool calls fail with
HTTP 500 "The model produced output that does not match the expected
peg-native format"
common_chat_params_init_onyx exists on onyx-fair-patch, added there by
8bb73dd3d. It was never on this branch, so this is not a regression --
the two lines developed independently.
The code here is taken verbatim from that commit. It is the clean side of
`git merge origin/onyx-fair-patch`: chat.cpp is one of the files that
merges without conflict. The full merge is not viable -- it produces 13
conflicts, including add/add on conversion/onyx.py and src/models/onyx.cpp
where the q_norm-folding and metadata-scale approaches contradict each
other, and #4/#7 are stacked on this branch's side of that.
Verified on this branch: builds with 0 errors, converts an Onyx checkpoint,
and serving it gives "4" for "What is 2+2?" plus a correct
get_weather {"city":"Paris"} tool call, where the unported branch gives the
two failures above.
No converter or runtime changes are included, so this should not interact
with the q_norm work.
Co-authored-by: Beto de Paola <betodepaola@meta.com>
* Less params, bilinear pos-emb interpolation as a graph op instead of CPU
* Map to symbolic V_MMPROJ instead of strings
* Make a couple params explicit
* Patchify via build_inp()
* No param for rope_theta
* Small cleanup
* Restore blank line
* Unpermute, to adapt to the latest transformers checkpoint
* Apply norm after token embeddings
This follows the latest transformers approach.
* Remove duplicated function
* build_vit
* onyx: use the model rope theta on sliding-window layers
* DFlash: conversion from transformers drafter
* Revert rope_type derivation from target
NOTE: this breaks compatibility with Meta's distributed DFlash GGUFs, as
the Q/K are stored in "NEOX" (rotated half) format, like in
transformers.
* Apply suggestion from @pcuenca
* Set model type
* Remove comment that will become obsolete
* Hardcode post_norm_rms_eps instead of new param
* Derive SWA+RoPE pattern from gguf array or scalar
* Fix model type <-> number of layers
* Reorder
* Rename
* Fix typo
* DFlash: seed the draft KV cache from multimodal embedding batches
`common_speculative_impl_draft_dflash::process()` returned early on any batch carrying embeddings, so an image prefill never had its target-layer features fused through the DFlash encoder and injected into the draft's KV cache. That left a hole spanning the image's positions, and the next injection at a post-image position failed to initialize its batch:
```
decoding image batch 1/1, n_tokens_batch = 256
decode: failed to initialize batch
llama_decode: failed to decode, ret = -1
process: llama_decode(ctx_dft) failed rc=-1 (n_tokens=17, offset=0)
srv decode: failed to process speculative batch
```
Every image request with `--spec-type draft-dflash` failed with HTTP 500. Text-only was unaffected, since those batches carry token ids and were let through.
Restore the earlier condition, which admits a batch that is either tokens or embeddings and skips only the degenerate neither/both cases. The rest of `process()` is already layout-agnostic -- it gathers features via `llama_get_embeddings_layer_inp()` and indexes `batch_in.pos[]` / `batch_in.seq_id[]`, none of which assume token ids -- so this is the whole fix.
Validated against `muse-glimmer-30B-bf16.gguf` + `mmproj-muse-glimmer-30B-bf16.gguf` + a DFlash draft head, on an image describe-the-shapes request:
- before: HTTP 500, `failed to process speculative batch`
- after: HTTP 200, draft acceptance 0.34012 (167 accepted / 491 generated), mean len 3.04
Output equivalence holds, which is the property that matters: at temperature 0 the drafted response is byte-identical to the same request served with no draft attached (1213/1213 chars), so the draft is drafting correctly through the image context rather than merely not crashing.
* Conversion: prefer rewrite to mapping
* Revert "Conversion: prefer rewrite to mapping"
This reverts commit a92d0ac584d315e876741e85b6dad3dbc8b23bf7.
* fix lint
* sliding_window metadata is not optional
* disable state save/load
* Apply suggestion from @pcuenca
---------
Co-authored-by: Young Han <younghan@fb.com>
Co-authored-by: Beto de Paola <betodepaola@meta.com>
Co-authored-by: Daniel Han <michaelhan2050@gmail.com>
Co-authored-by: ruanrms <ruanslv@gmail.com>
Co-authored-by: Xuan Son Nguyen <son@huggingface.co>
Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co>
|
||
|---|---|---|
| .. | ||
| afmoe.cpp | ||
| apertus.cpp | ||
| arcee.cpp | ||
| arctic.cpp | ||
| arwkv7.cpp | ||
| baichuan.cpp | ||
| bailingmoe.cpp | ||
| bailingmoe2.cpp | ||
| bert.cpp | ||
| bitnet.cpp | ||
| bloom.cpp | ||
| chameleon.cpp | ||
| chatglm.cpp | ||
| clip.cpp | ||
| codeshell.cpp | ||
| cogvlm.cpp | ||
| cohere2.cpp | ||
| cohere2moe.cpp | ||
| command-r.cpp | ||
| dbrx.cpp | ||
| deci.cpp | ||
| deepseek.cpp | ||
| deepseek2.cpp | ||
| deepseek2ocr.cpp | ||
| deepseek4.cpp | ||
| deepseek32.cpp | ||
| delta-net-base.cpp | ||
| dflash.cpp | ||
| dots1.cpp | ||
| dream.cpp | ||
| eagle3.cpp | ||
| ernie4-5-moe.cpp | ||
| ernie4-5.cpp | ||
| eurobert.cpp | ||
| exaone-moe.cpp | ||
| exaone.cpp | ||
| exaone4.cpp | ||
| falcon-h1.cpp | ||
| falcon.cpp | ||
| gemma-embedding.cpp | ||
| gemma.cpp | ||
| gemma2.cpp | ||
| gemma3.cpp | ||
| gemma3n.cpp | ||
| gemma4-assistant.cpp | ||
| gemma4.cpp | ||
| glm-dsa.cpp | ||
| glm4-moe.cpp | ||
| glm4.cpp | ||
| gpt2.cpp | ||
| gptneox.cpp | ||
| granite-hybrid.cpp | ||
| granite-moe.cpp | ||
| granite-switch.cpp | ||
| granite.cpp | ||
| grok.cpp | ||
| grovemoe.cpp | ||
| hunyuan-dense.cpp | ||
| hunyuan-moe.cpp | ||
| hunyuan-vl.cpp | ||
| hy-v3.cpp | ||
| internlm2.cpp | ||
| jais.cpp | ||
| jais2.cpp | ||
| jamba.cpp | ||
| jina-bert-v2.cpp | ||
| jina-bert-v3.cpp | ||
| kimi-linear.cpp | ||
| laguna.cpp | ||
| lfm2.cpp | ||
| lfm2moe.cpp | ||
| llada-moe.cpp | ||
| llada.cpp | ||
| llama-embed.cpp | ||
| llama.cpp | ||
| llama4.cpp | ||
| maincoder.cpp | ||
| mamba-base.cpp | ||
| mamba.cpp | ||
| mamba2.cpp | ||
| mellum.cpp | ||
| mimo2.cpp | ||
| minicpm.cpp | ||
| minicpm3.cpp | ||
| minimax-m2.cpp | ||
| minimax-m3.cpp | ||
| mistral3.cpp | ||
| mistral4.cpp | ||
| models.h | ||
| modern-bert.cpp | ||
| mpt.cpp | ||
| muse-glimmer.cpp | ||
| nanbeige.cpp | ||
| nemotron-h-moe.cpp | ||
| nemotron-h.cpp | ||
| nemotron.cpp | ||
| neo-bert.cpp | ||
| nomic-bert-moe.cpp | ||
| nomic-bert.cpp | ||
| olmo.cpp | ||
| olmo2.cpp | ||
| olmoe.cpp | ||
| openai-moe.cpp | ||
| openelm.cpp | ||
| orion.cpp | ||
| paddleocr.cpp | ||
| pangu-embed.cpp | ||
| phi2.cpp | ||
| phi3.cpp | ||
| phimoe.cpp | ||
| plamo.cpp | ||
| plamo2.cpp | ||
| plamo3.cpp | ||
| plm.cpp | ||
| qwen.cpp | ||
| qwen2.cpp | ||
| qwen2moe.cpp | ||
| qwen2vl.cpp | ||
| qwen3.cpp | ||
| qwen3moe.cpp | ||
| qwen3next.cpp | ||
| qwen3tts.cpp | ||
| qwen3vl.cpp | ||
| qwen3vlmoe.cpp | ||
| qwen35.cpp | ||
| qwen35moe.cpp | ||
| refact.cpp | ||
| rnd1.cpp | ||
| rwkv6-base.cpp | ||
| rwkv6.cpp | ||
| rwkv6qwen2.cpp | ||
| rwkv7-base.cpp | ||
| rwkv7.cpp | ||
| seed-oss.cpp | ||
| smallthinker.cpp | ||
| smollm3.cpp | ||
| stablelm.cpp | ||
| starcoder.cpp | ||
| starcoder2.cpp | ||
| step35.cpp | ||
| t5.cpp | ||
| t5encoder.cpp | ||
| talkie.cpp | ||
| wavtokenizer-dec.cpp | ||
| xverse.cpp | ||