kvcache-ai-ktransformers

mirror of https://github.com/kvcache-ai/ktransformers.git synced 2026-04-28 03:39:48 +00:00

History

Jiaqi Liao fcf8882075 Some checks are pending Book-CI / test-2 (push) Waiting to run Details Book-CI / test (push) Waiting to run Details Book-CI / test-1 (push) Waiting to run Details Deploy / deploy (macos-latest) (push) Waiting to run Details Deploy / deploy (ubuntu-latest) (push) Waiting to run Details Deploy / deploy (windows-latest) (push) Waiting to run Details [Feature] Add avx-based kimi-k2 support (#1656 ) * support Kimi-K2-Thinking original weight fix amx kernel bug * update k2 avx kernel. * feat: add CPUInfer write buffer task * [feat]: add kimi k2 cpu write buffer support - Implement write_weights_to_buffer function in k2-moe.hpp for extracting GPU expert weights - Fix down (w2) weight column-wise slicing for different TP configurations - Support three TP scenarios: cpu_tp == gpu_tp, cpu_tp > gpu_tp, cpu_tp < gpu_tp - Add comprehensive test cases for weight extraction validation - Ensure compatibility with Kimi model's MoE architecture * [fix]: correct write_weight_scale_to_buffer expert offset calculation Fixed the bug in write_weight_scale_to_buffer_task where expert offsets in GPU buffers were incorrectly calculated. Changed from using per_expert_gpu sizes to using full gpu_tp sizes, ensuring correct memory layout for multi-expert scenarios. Also added benchmark scripts for k2 moe and write buffer operations, and cleaned up debug output in test files. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com> * [feat]: add write buffer wrapper * [fix] fix comment --------- Co-authored-by: ouqingliang <1692110604@qq.com> Co-authored-by: Claude <noreply@anthropic.com>		2025-12-02 16:01:07 +08:00
..
.gitignore	add kt-kernel	2025-10-12 05:13:00 +00:00
bench_attention.py	update kt-kernel	2025-11-03 15:19:52 +08:00
bench_attention_torch.py	update kt-kernel	2025-11-03 15:19:52 +08:00
bench_k2_moe_amx.py	[Feature] Add avx-based kimi-k2 support (#1656 )	2025-12-02 16:01:07 +08:00
bench_k2_write_buffer.py	[Feature] Add avx-based kimi-k2 support (#1656 )	2025-12-02 16:01:07 +08:00
bench_linear.py	update kt-kernel	2025-11-03 15:19:52 +08:00
bench_linear_torch.py	add kt-kernel	2025-10-12 05:13:00 +00:00
bench_mla.py	update kt-kernel	2025-11-03 15:19:52 +08:00
bench_mlp.py	update kt-kernel	2025-11-03 15:19:52 +08:00
bench_mlp_torch.py	add kt-kernel	2025-10-12 05:13:00 +00:00
bench_moe.py	update kt-kernel	2025-11-03 15:19:52 +08:00
bench_moe_amx.py	update kt-kernel	2025-11-03 15:19:52 +08:00
bench_moe_amx_k.py	update kt-kernel	2025-11-03 15:19:52 +08:00
bench_moe_kernel.py	update kt-kernel	2025-11-03 15:19:52 +08:00
bench_moe_kernel_tiling.py	update kt-kernel	2025-11-03 15:19:52 +08:00
bench_moe_kml.py	update kt-kernel	2025-11-03 15:19:52 +08:00
bench_moe_torch.py	add kt-kernel	2025-10-12 05:13:00 +00:00
compare_moe_performance.py	update kt-kernel	2025-11-03 15:19:52 +08:00
Makefile	update kt-kernel	2025-11-03 15:19:52 +08:00
multi_bench_moe.py	update kt-kernel	2025-11-03 15:19:52 +08:00
upload-bench-json.py	add kt-kernel	2025-10-12 05:13:00 +00:00