mirror of
https://github.com/LostRuins/koboldcpp.git
synced 2026-10-03 11:35:46 +00:00
* scripts : add initial profiling script (wip) * src : add precompile headers (PCH) for models.h * common : add common.h as PCH * ggml : add PCH for ggml-impl.h * mtmd : use PCH for models.h * scripts : add script to build with Server/Tools/Tests * server : add PCH for common.h * docs: add profiling progress notes (wip) * ggml : add exclude for GCC + SVE on ARM Refs: https://github.com/ggml-org/llama.cpp/actions/runs/33393906061/job/99493756214?pr=28091 * ggml : attempt to fix use of std::hardware_destructive_inference_size Refs: https://github.com/ggml-org/llama.cpp/actions/runs/33396221677/job/99501265689?pr=28091 * squash! ggml : attempt to fix use of std::hardware_destructive_inference_size Add a version check for GCC 12 to conditionally apply the `-Winterference-size` pragma. * editorconfig : exclude profiling reports dir This directory will not be included in the merge later and this commit can be ignore at that point. Just fixing to keep CI happy. * ggml : skip PCH for gcc on non-x86 architectures * tests : add PCH for peg-parser/tests.h There are 7 peg-parser tests that can share one PCH instead of then each parsing the full tests.h. * common : add PCH for chat.h * docs : update linux build profiling full results Just updating after a number of PCH additions. These are not exact figures and will vary a bit from run to run, but they give a general idea of the performance impact of PCH. * cmake : introduce unity build for models This commit introduces a unity build for the models to improve compilation time. The improvements were roughly the following: ```console +------------------------+-----+------------+------------+------------+ | Build | TUs | Frontend | Backend | Total | +------------------------+-----+------------+------------+------------+ | Full, master | 396 | 811.0 s | 692.2 s | 1,503.2 s | | Full, with PCH | 405 | 380.0 s | 664.7 s | 1,044.7 s | | Full, with PCH + UB | 264 | 357.7 s | 635.7 s | 993.4 s | +------------------------+-----+------------+------------+------------+ TU = Translation Unit. Full = includes Server, Tools, and Tests. PCH = precompiled headers. UB = unity build for models. ``` * docs : update linux profiling table with unitiy build results * docs : update mac profiling results to include unity build [no ci] * docs: remove profiling reports * scripts : merge build profile scripts into one script I was lazy before and just copied the first script to enable Tests, Server, and Tools. This now merges them into a single script. * Revert "editorconfig : exclude profiling reports dir" [no ci] This reverts commit 2922a12118a0730d2f7632bcba265b44a0856c59. * src : rename ggml_view_2d_slice to gemma3n_view_2d_slice This is to be consistent with the rename in gemma4.cpp which was required to avoid a name clash. * cmake : add build profile script for windows [no ci] This commit adds a port of the scripts/build-profile.sh script to windows powershell. This was developed on Windows on ARM but should work on X64 as well but needs to be tested there as well. |
||
|---|---|---|
| .. | ||
| models | ||
| CMakeLists.txt | ||
| llama-adapter.cpp | ||
| llama-adapter.h | ||
| llama-arch.cpp | ||
| llama-arch.h | ||
| llama-batch.cpp | ||
| llama-batch.h | ||
| llama-chat.cpp | ||
| llama-chat.h | ||
| llama-context.cpp | ||
| llama-context.h | ||
| llama-cparams.cpp | ||
| llama-cparams.h | ||
| llama-ext.h | ||
| llama-grammar.cpp | ||
| llama-grammar.h | ||
| llama-graph.cpp | ||
| llama-graph.h | ||
| llama-hparams.cpp | ||
| llama-hparams.h | ||
| llama-impl.cpp | ||
| llama-impl.h | ||
| llama-io.cpp | ||
| llama-io.h | ||
| llama-kv-cache-dsa-iswa.cpp | ||
| llama-kv-cache-dsa-iswa.h | ||
| llama-kv-cache-dsa.cpp | ||
| llama-kv-cache-dsa.h | ||
| llama-kv-cache-dsv4.cpp | ||
| llama-kv-cache-dsv4.h | ||
| llama-kv-cache-iswa.cpp | ||
| llama-kv-cache-iswa.h | ||
| llama-kv-cache-msa.cpp | ||
| llama-kv-cache-msa.h | ||
| llama-kv-cache.cpp | ||
| llama-kv-cache.h | ||
| llama-kv-cells.h | ||
| llama-memory-hybrid-idx.cpp | ||
| llama-memory-hybrid-idx.h | ||
| llama-memory-hybrid-iswa.cpp | ||
| llama-memory-hybrid-iswa.h | ||
| llama-memory-hybrid.cpp | ||
| llama-memory-hybrid.h | ||
| llama-memory-recurrent.cpp | ||
| llama-memory-recurrent.h | ||
| llama-memory.cpp | ||
| llama-memory.h | ||
| llama-mmap.cpp | ||
| llama-mmap.h | ||
| llama-model-loader.cpp | ||
| llama-model-loader.h | ||
| llama-model-saver.cpp | ||
| llama-model-saver.h | ||
| llama-model.cpp | ||
| llama-model.h | ||
| llama-quant.cpp | ||
| llama-quant.h | ||
| llama-sampler.cpp | ||
| llama-sampler.h | ||
| llama-version.h.in | ||
| llama-vocab.cpp | ||
| llama-vocab.h | ||
| llama.cpp | ||
| unicode-data.cpp | ||
| unicode-data.h | ||
| unicode.cpp | ||
| unicode.h | ||