[Feature] Add avx-based kimi-k2 support (#1656)

* support Kimi-K2-Thinking original weight fix amx kernel bug * update k2 avx kernel. * feat: add CPUInfer write buffer task * [feat]: add kimi k2 cpu write buffer support - Implement write_weights_to_buffer function in k2-moe.hpp for extracting GPU expert weights - Fix down (w2) weight column-wise slicing for different TP configurations - Support three TP scenarios: cpu_tp == gpu_tp, cpu_tp > gpu_tp, cpu_tp < gpu_tp - Add comprehensive test cases for weight extraction validation - Ensure compatibility with Kimi model's MoE architecture * [fix]: correct write_weight_scale_to_buffer expert offset calculation Fixed the bug in write_weight_scale_to_buffer_task where expert offsets in GPU buffers were incorrectly calculated. Changed from using per_expert_gpu sizes to using full gpu_tp sizes, ensuring correct memory layout for multi-expert scenarios. Also added benchmark scripts for k2 moe and write buffer operations, and cleaned up debug output in test files. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com> * [feat]: add write buffer wrapper * [fix] fix comment --------- Co-authored-by: ouqingliang <1692110604@qq.com> Co-authored-by: Claude <noreply@anthropic.com>
2026-04-28 20:00:06 +00:00 · 2025-12-02 16:01:07 +08:00 · 2025-12-02 16:01:07 +08:00 · fcf8882075
commit fcf8882075
parent c2b8c60c4e
12 changed files with 2649 additions and 34 deletions
--- a/kt-kernel/python/utils/init.py
+++ b/kt-kernel/python/utils/init.py
@ -4,13 +4,15 @@
 Utilities for kt_kernel package.
 """

-from .amx import AMXMoEWrapper
+from .amx import AMXMoEWrapper, RAWAMXMoEWrapper
 from .llamafile import LlamafileMoEWrapper
-from .loader import SafeTensorLoader, GGUFLoader
+from .loader import SafeTensorLoader, GGUFLoader, CompressedSafeTensorLoader

 __all__ = [
    "AMXMoEWrapper",
+    "RAWAMXMoEWrapper",
    "LlamafileMoEWrapper",
    "SafeTensorLoader",
+    "CompressedSafeTensorLoader",
    "GGUFLoader",
 ]