vrr/kvcache-ai-ktransformers

mirror of https://github.com/kvcache-ai/ktransformers.git synced 2025-09-05 20:19:51 +00:00

Author	SHA1	Message	Date
hrz6976	2c3dcd9774	Add a lock to server inference()	2025-02-13 10:05:22 +00:00
Azure	c4d9bc6670	support KExpertsMarlin backend	2025-02-07 05:57:40 +00:00
Azure	907251c743	done support deepseekv3	2025-02-04 15:53:38 +00:00
Azure	476b1d8dc6	support deepseekv3; runable but have precition problem	2025-01-31 08:27:24 +00:00
liam	dd1d8667f3	✨: refactor local_chat and fix message slice bug in server	2024-11-04 14:02:19 +08:00
chenxl	b9f0819a86	None for load config	2024-08-22 15:52:25 +00:00
TangJingqi	170b7a6001	fix server don't accept yaml path as param; fix server static cache device problem	2024-08-21 14:19:43 +08:00
Atream	412055d450	[feature] experts can be injected using CPUInfer [fix] fix ktransformers interface when use new CUDAGraphRunner [fix] fix YAML and optimize logic, the top rule has the highest priority	2024-08-14 16:10:54 +08:00
chenxl	18c42e67df	Initial commit	2024-07-27 16:06:58 +08:00