koboldcpp/scripts
Max Krasnyansky 58367713a6
hexagon: new HMX-optimized GATED_DELTA_NET (#29199)
* hex-gdn: start putting together HMX support for GDN

* hex-gdn: working hmx but not-pipelined and slow for now

* hex-gdn: re-write vtcm layout handling and prep for pipelining

* hex-gdn: starting to pipeline hmx and dmas

* hex-gdn: add hvx threading for most pipeline stages

* hex-gdb: add detailed trace events

* hex-gdn: vectorize expfs and use aligned hvx reads/writes

* hex-gnd: vectorize the rest of expf

* hex-gdn: optimize tail processing (pad partial chunks)

* hex-gdb: avoid float up/down casts in hot loops

* hex-fa: remove float up/down casts from inner loops

* hex-gdn: do exp() in f16 to improve HVX utilization

* hex-gdn: optimize tiler

* hex-hmx: bump hmx-queue to 128 and dispatch all GDN gemms at once

* hex-gdn: further pipeline improvements

* hex-gdn: optimize gdn prep stage

* hex-gdn: yet more tweaks to optimize GND_SOLVE task and pipeline

* hex-gdn: improve accuracy and optmize gdn-prep further

* hex-gdn: fix rebase conflict

* hex-bufs: revert max_bufsize enforcement, it is enough to just enforce max_vmem

* hex-scripts: improved inspect script to avoid false alarms in reg spill detector

* hex-fa: improve inline softmax with in-reg VKQ32 accum

* hex-fa: minor improvement for dma pipeline in hvx kernel

* hex-fa: reduce ddr reads by 20-30% during token gen

* hex-gdn: proper alignment for hvx vtcm spads
2026-09-21 14:49:52 -07:00
..
apple
hip CI: hip-quality-check: ignore spill added in bfdc32183d (#28909) 2026-09-14 21:25:22 +02:00
jinja ci : bump ty to 0.0.78 (#28548) 2026-09-07 21:10:06 +02:00
snapdragon hexagon: new HMX-optimized GATED_DELTA_NET (#29199) 2026-09-21 14:49:52 -07:00
bench-models.sh common: migrate the deprecated --mmap/--no-mmap to --load-mode (#26934) 2026-08-15 16:35:53 +08:00
build-info.sh
build-profile.ps1 cmake : add PCH and unity build to improve build times (#28091) 2026-09-11 13:01:29 +02:00
build-profile.sh cmake : add PCH and unity build to improve build times (#28091) 2026-09-11 13:01:29 +02:00
ccache-clear.sh ci : apply ccache-clear with older/min/dry-run to all ccache jobs (#27602) 2026-08-24 10:49:20 +03:00
check-apiabi-compat.sh ci : add API/ABI check to make-release workflow [no ci] (#28947) 2026-09-17 09:55:16 +02:00
check-release-apiabi.sh ci : add API/ABI check to make-release workflow [no ci] (#28947) 2026-09-17 09:55:16 +02:00
check-requirements.sh
compare-commits.sh
compare-llama-bench.py args: refactor mlock/mmap/directio into load-mode (#20834) 2026-07-23 20:32:56 +08:00
compare-logprobs.py scripts: update corpus of compare-logprobs (#19326) 2026-02-25 12:57:34 +01:00
create_ops_docs.py
debug-test.sh refactor : remove libcurl, use OpenSSL when available (#18828) 2026-01-14 18:02:47 +01:00
gen-authors.sh
gen-unicode-data.py ci : bump ty to 0.0.26 (#21156) 2026-03-30 09:29:15 +02:00
get-flags.mk
get-hellaswag.sh scripts : update get-hellaswag.sh and get-winogrande.sh (#20542) 2026-03-14 11:21:50 +01:00
get-pg.sh
get-wikitext-2.sh scripts : improve get-wikitext-2.sh (#19952) 2026-03-02 15:40:49 +01:00
get-winogrande.sh scripts : update get-hellaswag.sh and get-winogrande.sh (#20542) 2026-03-14 11:21:50 +01:00
get_chat_template.py
git-bisect-run.sh llama: end-to-end tests (#19802) 2026-03-08 12:30:21 +01:00
git-bisect.sh llama: end-to-end tests (#19802) 2026-03-08 12:30:21 +01:00
hf.sh
install-oneapi.bat
make-release-checks.sh ci : add API/ABI check to make-release workflow [no ci] (#28947) 2026-09-17 09:55:16 +02:00
make-release-desc.sh ci : add nightly-tag.txt to make-release (#27485) 2026-08-21 13:20:44 +03:00
make-release-summary.txt llama.cpp : bump version to 0.3.0 (#27696) 2026-08-25 12:42:21 +03:00
pr2wt.sh pr2wt : use ssh/https remote in worktree depending on base (#27800) 2026-08-27 16:27:17 +03:00
release.sh scripts : add release.sh for release preparation (#27497) 2026-08-21 14:51:26 +03:00
serve-static.js refactor : remove libcurl, use OpenSSL when available (#18828) 2026-01-14 18:02:47 +01:00
server-bench.py server-bench : add speed-bench for speculative decoding benchmarking (#23869) 2026-05-29 23:09:47 +02:00
server-test-function-call.py common/autoparser: fixes for newline handling / forced tool calls (#22654) 2026-05-04 13:18:11 +02:00
server-test-model.py Autoparser - complete refactoring of parser architecture (#18675) 2026-03-06 21:01:00 +01:00
server-test-parallel-tc.py chat: fix parallel_tool_calls default setting based on model capabilities, add tests for parallel tool calls and structured outputs (#22217) 2026-04-22 18:10:56 +02:00
server-test-structured.py parser: fix structured output bug (#22302) 2026-04-24 23:19:55 +02:00
sync-ggml-am.sh
sync-ggml.last sync : ggml 2026-09-14 16:45:33 +03:00
sync-ggml.sh
sync_vendor.py vendor : update cpp-httplib to 0.57.1 (#29239) 2026-09-21 22:50:20 +02:00
tool_bench.py ci : bump ty to 0.0.78 (#28548) 2026-09-07 21:10:06 +02:00
tool_bench.sh
ui-assets.cmake ui : add cache (#28802) 2026-09-12 16:09:46 +02:00
verify-checksum-models.py
wc2wt.sh scripts : allow wc2wt with an existing branch (#23189) 2026-05-18 08:57:28 +03:00