Commit graph

5 commits

Author SHA1 Message Date
Atream
25cee5810e add balance-serve, support concurrence 2025-03-31 22:55:32 +08:00
Azure-Tang
e5b001d76f Update readme; Format code; Add example yaml. 2025-03-14 14:25:52 -04:00
Azure-Tang
ed8437413b merge main; Add torch q8 linear 2025-03-14 05:52:07 -04:00
Atream
5ec33d046d optimize gguf dequant, save mem, support Q2_K
use marlin for lm_head, lm_head only calc last token for prefill
extend context window to 19K for DeepSeek-V3/R1 within 24GB VRAM
2025-02-22 06:13:01 +00:00
anyanqilin
2d67016d14 wjh-change 2024-11-04 14:02:19 +08:00