koboldcpp

mirror of https://github.com/LostRuins/koboldcpp.git synced 2025-09-11 01:24:36 +00:00

Author	SHA1	Message	Date
slaren	7fc50c051a	cuBLAS: use host pinned memory and dequantize while copying (#1207 ) * cuBLAS: dequantize simultaneously while copying memory * cuBLAS: use host pinned memory * cuBLAS: improve ggml_compute_forward_mul_mat_f16_f32 with pinned memory * cuBLAS: also pin kv cache * fix rebase	2023-04-29 02:04:18 +02:00
Stephan Walter	36d19a603b	Remove Q4_3 which is no better than Q5 (#1218 )	2023-04-28 23:10:43 +00:00
Evan Jones	1481a9cf25	llama : add session file format and saved sessions in main (#1169 )	2023-04-28 18:59:37 +03:00
0cc4m	7296c961d9	ggml : add CLBlast support (#1164 ) * Allow use of OpenCL GPU-based BLAS using ClBlast instead of OpenBLAS for context processing * Improve ClBlast implementation, avoid recreating buffers, remove redundant transfers * Finish merge of ClBlast support * Move CLBlast implementation to separate file Add buffer reuse code (adapted from slaren's cuda implementation) * Add q4_2 and q4_3 CLBlast support, improve code * Double CLBlast speed by disabling OpenBLAS thread workaround Co-authored-by: Concedo <39025047+LostRuins@users.noreply.github.com> Co-authored-by: slaren <2141330+slaren@users.noreply.github.com> * Fix device selection env variable names * Fix cast in opencl kernels * Add CLBlast to CMakeLists.txt * Replace buffer pool with static buffers a, b, qb, c Fix compile warnings * Fix typos, use GGML_TYPE defines, improve code * Improve btype dequant kernel selection code, add error if type is unsupported * Improve code quality * Move internal stuff out of header * Use internal enums instead of CLBlast enums * Remove leftover C++ includes and defines * Make event use easier to read Co-authored-by: Henri Vasserman <henv@hot.ee> * Use c compiler for opencl files * Simplify code, fix include * First check error, then release event * Make globals static, fix indentation * Rename dequant kernels file to conform with other file names * Fix import cl file name --------- Co-authored-by: Concedo <39025047+LostRuins@users.noreply.github.com> Co-authored-by: slaren <2141330+slaren@users.noreply.github.com> Co-authored-by: Henri Vasserman <henv@hot.ee> Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>	2023-04-28 17:57:16 +03:00
Concedo	95bbd46019	Merge branch 'master' into concedo_experimental # Conflicts: # .devops/tools.sh # README.md	2023-04-27 16:12:00 +08:00
Georgi Gerganov	574406dc7e	ggml : add Q5_0 and Q5_1 quantization (#1187 ) * ggml : add Q5_0 quantization (cuBLAS only) * ggml : fix Q5_0 qh -> uint32_t * ggml : fix q5_0 histogram stats * ggml : q5_0 scalar dot product * ggml : q5_0 ARM NEON dot * ggml : q5_0 more efficient ARM NEON using uint64_t masks * ggml : rename Q5_0 -> Q5_1 * ggml : adding Q5_0 mode * quantize : add Q5_0 and Q5_1 to map * ggml : AVX2 optimizations for Q5_0, Q5_1 (#1195) --------- Co-authored-by: Stephan Walter <stephan@walter.name>	2023-04-26 23:14:13 +03:00
Ásgeir Bjarni Ingvarsson	87a6f846d3	Allow setting the rng seed after initialization. (#1184 ) The llama_set_state_data function restores the rng state to what it was at the time llama_copy_state_data was called. But users may want to restore the state and proceed with a different seed.	2023-04-26 22:08:43 +02:00
Concedo	93a8e00dfa	Merge branch 'master' into concedo # Conflicts: # flake.nix	2023-04-26 18:01:35 +08:00
Georgi Gerganov	7a32fcb3b2	ggml : add Q8_0 quantization format (rename the old one to Q8_1) (ARM NEON) (#1179 ) * ggml : add Q8_0 quantization format (rename the old one to Q8_1) * tests : fix test-quantize-fns * ggml : finalize Q8_0 implementation * ggml : use q4_0_q8_0 and q4_2_q8_0 * ggml : fix Q8_0 dot product bug (ARM) * ggml : Q8_0 unroll x2 * ggml : fix bug - using wrong block type * ggml : extend quantize_fns_t with "vec_dot_type" * ggml : fix Q8_0 to use 255 values out of 256 * ggml : fix assert using wrong QK4_2 instead of QK4_3	2023-04-25 23:40:51 +03:00
Concedo	235daf4016	Merge branch 'master' into concedo # Conflicts: # .github/workflows/build.yml # README.md	2023-04-25 20:44:22 +08:00
Georgi Gerganov	957c8ae21d	llama : increase scratch buffer size for 65B (ref #1152 ) Temporary solution	2023-04-24 18:47:30 +03:00
Concedo	e58f1d1336	Merge branch 'master' into concedo_experimental	2023-04-24 19:43:17 +08:00
Georgi Gerganov	c4fe84fb0d	llama : refactor get / set state + remove redundant kv cache API (#1143 )	2023-04-24 07:40:02 +03:00
Concedo	8e615c8245	Merge branch 'master' into concedo_experimental # Conflicts: # README.md	2023-04-24 12:20:08 +08:00
Georgi Gerganov	e4422e299c	ggml : better PERF prints + support "LLAMA_PERF=1 make"	2023-04-23 18:15:39 +03:00
Concedo	7c60441d71	Merge branch 'master' into concedo # Conflicts: # .github/workflows/build.yml # CMakeLists.txt	2023-04-22 23:46:14 +08:00
Stephan Walter	c50b628810	Fix CI: ARM NEON, quantization unit tests, editorconfig (#1122 )	2023-04-22 10:54:13 +00:00
Concedo	1b7aa2b815	Merge branch 'master' into concedo # Conflicts: # .github/workflows/build.yml # CMakeLists.txt # Makefile	2023-04-22 16:22:08 +08:00
Georgi Gerganov	872c365a91	ggml : fix AVX build + update to new Q8_0 format	2023-04-22 11:08:12 +03:00
Concedo	1ea0e15292	Merge branch 'master' into concedo # Conflicts: # llama.cpp	2023-04-22 16:07:27 +08:00
xaedes	b6e7f9b09e	llama : add api for getting/setting the complete state: rng, logits, embedding and kv_cache (#1105 ) * reserve correct size for logits * add functions to get and set the whole llama state: including rng, logits, embedding and kv_cache * remove unused variables * remove trailing whitespace * fix comment	2023-04-22 09:21:32 +03:00
Concedo	cee018960e	Merge branch 'master' into concedo_experimental	2023-04-22 00:19:50 +08:00
xaedes	8687c1f258	llama : remember and restore kv cache data pointers (#1104 ) because their value is stored in buf and overwritten by memcpy	2023-04-21 18:25:21 +03:00
Concedo	82d74ca1a6	Merge branch 'master' into concedo # Conflicts: # .github/workflows/build.yml	2023-04-21 16:24:30 +08:00
Georgi Gerganov	d40fded93e	llama : fix comment for "output.weight" tensor	2023-04-21 10:24:02 +03:00
Georgi Gerganov	12b5900dbc	ggml : sync ggml (add GPT-NeoX RoPE implementation)	2023-04-20 23:32:59 +03:00
Kawrakow	38de86a711	llama : multi-threaded quantization (#1075 ) * Multi-threading quantization. Not much gain for simple quantizations, bit it will be important for quantizations that require more CPU cycles. * Multi-threading for quantize-stats It now does the job in ~14 seconds on my Mac for Q4_0, Q4_1 and Q4_2. Single-threaded it was taking more than 2 minutes after adding the more elaborate version of Q4_2. * Reviewer comments * Avoiding compiler confusion After changing chunk_size to const int as suggested by @ggerganov, clang and GCC starting to warn me that I don't need to capture it in the lambda. So, I removed it from the capture list. But that makes the MSVC build fail. So, making it a constexpr to make every compiler happy. * Still fighting with lambda captures in MSVC --------- Co-authored-by: Iwan Kawrakow <iwan.kawrakow@gmail.com> Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>	2023-04-20 20:42:27 +03:00
Georgi Gerganov	e0305ead3a	ggml : add Q4_3 quantization (#1082 )	2023-04-20 20:35:53 +03:00
Concedo	be1222c36e	Merged the upstream cublas feature,	2023-04-19 20:45:37 +08:00
slaren	8944a13296	Add NVIDIA cuBLAS support (#1044 )	2023-04-19 11:22:45 +02:00
Concedo	f662a9a230	Merge branch 'master' into concedo # Conflicts: # .github/workflows/build.yml # .github/workflows/docker.yml # CMakeLists.txt # Makefile # README.md	2023-04-19 16:34:51 +08:00
Georgi Gerganov	77a73403ca	ggml : add new Q4_2 quantization (ARM only) (#1046 ) * ggml : Q4_2 ARM * ggml : add ggml_is_quantized() * llama : update llama_type_name() with Q4_2 entry * ggml : speed-up q4_2 - 4 threads: ~100ms -> ~90ms - 8 threads: ~55ms -> ~50ms * ggml : optimize q4_2 using vmlaq_n_f32 + vmulq_n_f32	2023-04-18 23:54:57 +03:00
Concedo	ac61e34d5f	Merge branch 'master' into concedo_experimental # Conflicts: # CMakeLists.txt # README.md	2023-04-18 17:38:10 +08:00
slaren	315a95a4d3	Add LoRA support (#820 )	2023-04-17 17:28:55 +02:00
Arik Poznanski	efd05648c8	llama : well-defined static initialization of complex objects (#927 ) * Replaced static initialization of complex objects with a initialization on first use. This prevents an undefined behavior on program run, for example, crash in Release build, works in Debug build * replaced use of auto with exact type to avoid using -std=c++14 * Made the assessors functions for static maps be static const	2023-04-17 17:41:53 +03:00
Ivan Komarov	f266259ad9	Speedup the AVX-512 implementation of ggml_vec_dot_q4_0() (#933 )	2023-04-17 15:10:57 +02:00
Concedo	96fb12cfa2	Merge branch 'master' into concedo	2023-04-16 21:59:05 +08:00
Georgi Gerganov	3173a62eb9	stdout : vertical align outputs for better readibility	2023-04-16 13:59:27 +03:00
nanahi	2d3481c721	Fix msys2 build error and warnings (#1009 )	2023-04-16 11:13:42 +02:00
Concedo	d00b865eb1	Merge branch 'master' into concedo # Conflicts: # .devops/full.Dockerfile # Makefile # flake.nix	2023-04-15 11:33:43 +08:00
Pavol Rusnak	c56b715269	Expose type name from ggml (#970 ) Avoid duplication of type names in utils Co-authored-by: Håkon H. Hitland <haakon@likedan.net>	2023-04-14 20:05:37 +02:00
Concedo	a819f22cac	Merge branch 'master' into concedo # Conflicts: # CMakeLists.txt # Makefile # README.md # flake.nix	2023-04-14 21:40:33 +08:00
Georgi Gerganov	9190e8eac8	llama : merge llama_internal.h into llama.h Hide it behind an #ifdef	2023-04-13 18:04:45 +03:00
Concedo	f4257a8eef	Merge branch 'master' into concedo	2023-04-12 23:25:45 +08:00
Concedo	1bd5992da4	clean and refactor handling of flags	2023-04-12 23:25:31 +08:00
Stephan Walter	e7f6997f89	Don't crash on ftype (formerly f16) == 4 (#917 )	2023-04-12 15:06:16 +00:00
Concedo	9245c7d7d0	Merge branch 'master' into concedo	2023-04-11 23:38:15 +08:00
Stephan Walter	3e6e70d8e8	Add enum llama_ftype, sync ggml_type to model files (#709 )	2023-04-11 15:03:51 +00:00
comex	2663d2c678	Windows fixes (#890 ) Mostly for msys2 and mingw64 builds, which are different from each other and different from standard Visual Studio builds. Isn't Windows fun? - Define _GNU_SOURCE in more files (it's already used in ggml.c for Linux's sake). - Don't use PrefetchVirtualMemory if not building for Windows 8 or later (mingw64 doesn't by default). But warn the user about this situation since it's probably not intended. - Check for NOMINMAX already being defined, which it is on mingw64. - Actually use the `increment` variable (bug in my `pizza` PR). - Suppress unused variable warnings in the fake pthread_create and pthread_join implementations for Windows. - (not Windows-related) Remove mention of `asprintf` from comment; `asprintf` is no longer used. Fixes #871.	2023-04-11 15:19:54 +02:00
Concedo	f53238f570	Merged the upstream updates for model loading code, and ditched the legacy llama loaders since they were no longer needed.	2023-04-10 12:00:34 +08:00

1 2 3

144 commits