Li, Zonghang
|
fbbc30c950
|
Merge branch 'speculative' into dev
|
2025-06-16 13:27:36 +04:00 |
|
Li, Zonghang
|
dfb1feb54e
|
update README
|
2025-06-16 12:09:07 +04:00 |
|
Li, Zonghang
|
45de284f3d
|
Merge branch 'fix' into speculative
|
2025-06-14 18:57:17 +04:00 |
|
Li, Zonghang
|
f38cfc625c
|
Merge branch 'fix' into dev
|
2025-06-14 18:56:36 +04:00 |
|
Li, Zonghang
|
b5ccd62135
|
fix n_gpu_layers allocation errors
|
2025-06-14 18:55:53 +04:00 |
|
Li, Zonghang
|
0a535cbdc1
|
Merge branch 'speculative' of github.com:Lizonghang/prima.cpp into speculative
|
2025-06-13 13:31:12 +04:00 |
|
Li, Zonghang
|
c9cae626cf
|
speculative: free sockets and send stop signal when inference ends
|
2025-06-13 13:30:29 +04:00 |
|
Li, Zonghang
|
2687ef3126
|
speculative: free sockets and send stop signal when inference ends
|
2025-06-13 11:25:42 +04:00 |
|
Li, Zonghang
|
dc875bbef9
|
fix speculative decoding
|
2025-06-13 08:18:12 +04:00 |
|
Zonghang Li
|
ba29717613
|
add feature: keep the forwarder if its previous device cannot directly connect to its next device.
feat: nodes attempt connections during topology rebuild while preserving forwarders
|
2025-06-12 16:57:35 +04:00 |
|
DeEMO
|
d4618de991
|
fix: block when free socket
|
2025-06-12 12:26:10 +00:00 |
|
DeEMO
|
2039e3b0c1
|
fix: send and recv meta
|
2025-06-12 12:26:10 +00:00 |
|
DeEMO
|
d6c8d322cd
|
fix try_connect
|
2025-06-12 12:26:10 +00:00 |
|
DeEMO
|
d1b97f798e
|
support reconnection
|
2025-06-12 12:26:09 +00:00 |
|
Zonghang Li
|
e50b3aa473
|
Merge pull request #27 from Lizonghang/lizh_dev
Fix seq_id mismatch between the head and worker devices.
|
2025-06-11 17:12:08 +04:00 |
|
Li, Zonghang
|
3e6d831930
|
fix seq_id mismatch between head and worker devices
|
2025-06-11 17:10:21 +04:00 |
|
Li, Zonghang
|
fb9b1f2b00
|
reformat llama.cpp
|
2025-06-09 13:04:22 +04:00 |
|
Li, Zonghang
|
fbf853341b
|
add endpoint /v1/cancel
|
2025-06-07 11:34:38 +04:00 |
|
Zonghang Li
|
c8af1be27e
|
Merge pull request #24 from Lizonghang/lizh_dev
Fix batch decoding and dynamic batching.
|
2025-06-07 01:02:10 +04:00 |
|
Li, Zonghang
|
22a6ddef13
|
fix batch decoding and dynamic batching
|
2025-06-07 00:53:56 +04:00 |
|
Lizonghang
|
e56be76bdf
|
assume only a single seq_id per token is needed
|
2025-06-07 00:42:44 +04:00 |
|
Lizonghang
|
d8aea899d1
|
fix n_seq_id and seq_id
|
2025-06-06 23:58:03 +04:00 |
|
Lizonghang
|
a1a2238831
|
add batch_all.n_seq_id and batch_all.seq_id to sync_meta
|
2025-06-06 23:36:53 +04:00 |
|
Lizonghang
|
68ecc8509d
|
add batch_all.logits to sync_meta
|
2025-06-06 22:58:48 +04:00 |
|
Lizonghang
|
500e066a2f
|
fix batch decoding and dynamic batching
|
2025-06-06 16:53:22 +04:00 |
|
Lizonghang
|
ef1e10101e
|
add test for IQ1 and doc for device selection
|
2025-06-04 15:12:00 +04:00 |
|
Lizonghang
|
27756ee182
|
fix: enable rolling back set assignment when all devices are assigned to M4 but no feasible solutions
|
2025-06-04 15:11:29 +04:00 |
|
Li, Zonghang
|
6439090920
|
reformat code
|
2025-06-03 23:53:24 +04:00 |
|
Li, Zonghang
|
b6fdbd541b
|
Merge branch 'dev' of github.com:Lizonghang/prima.cpp into dev
|
2025-06-03 18:20:17 +04:00 |
|
Lizonghang
|
9f0ec78a4b
|
Merge branch 'dev' of github.com:Lizonghang/prima.cpp into dev
|
2025-06-03 18:18:53 +04:00 |
|
Li, Zonghang
|
a01fafd126
|
Merge branch 'main' into dev
|
2025-06-03 17:56:47 +04:00 |
|
Li, Zonghang
|
1b3b6a506f
|
fix: add warm-up in profiling to prevent init delay
|
2025-06-03 17:10:09 +04:00 |
|
Li, Zonghang
|
b30f749e5e
|
fix n_embd cannot be divided by quantized block size
|
2025-06-03 14:06:31 +04:00 |
|
Li, Zonghang
|
e25c739ecf
|
Merge pull request #19 from yezhizi/feat/auto-exit
feat: ranks with only 1 layer auto-exit and rebuild topology
|
2025-05-20 02:04:50 +08:00 |
|
Li, Zonghang
|
7b0ededd24
|
Merge branch 'dev' into feat/auto-exit
|
2025-05-20 02:04:14 +08:00 |
|
Lizonghang
|
421b3deca5
|
fix llama-cli pos sync
|
2025-05-19 18:08:27 +04:00 |
|
Lizonghang
|
c54a6a0132
|
fix context shifting
|
2025-05-19 16:58:35 +04:00 |
|
DeEMO
|
34eaa8224d
|
fix: handle socket closure and connection in llama_rebuild_topo
Signed-off-by: DeEMO <yzzxrx@gmail.com>
|
2025-05-19 09:22:35 +00:00 |
|
DeEMO
|
8b61cb2fa4
|
fix: adapt the new topo
Signed-off-by: DeEMO <yzzxrx@gmail.com>
|
2025-05-19 09:22:29 +00:00 |
|
DeEMO
|
df16b1876f
|
refactor: add zmq helper to generate message
Signed-off-by: DeEMO <yzzxrx@gmail.com>
|
2025-05-19 09:22:24 +00:00 |
|
DeEMO
|
0ad009a2f4
|
fix: update serialization and deserialization for next_ip in device_info
Signed-off-by: DeEMO <yzzxrx@gmail.com>
|
2025-05-19 09:22:16 +00:00 |
|
DeEMO
|
4b36aef157
|
fix some bugs
Signed-off-by: DeEMO <yzzxrx@gmail.com>
|
2025-05-19 09:22:08 +00:00 |
|
DeEMO
|
cc46aa9828
|
update rank and n_world
Signed-off-by: DeEMO <yzzxrx@gmail.com>
|
2025-05-19 09:22:02 +00:00 |
|
DeEMO
|
fdd6694633
|
add topo rebuild
Signed-off-by: DeEMO <yzzxrx@gmail.com>
|
2025-05-19 09:21:53 +00:00 |
|
DeEMO
|
26bb86c09b
|
Add tune_layer_allocation
Signed-off-by: DeEMO <yzzxrx@gmail.com>
|
2025-05-19 09:21:22 +00:00 |
|
Lizonghang
|
07c4966a80
|
reduce fio data size to 1gb to speed up profiling
|
2025-05-14 21:26:01 +04:00 |
|
Lizonghang
|
2cc01483fd
|
support server mode
|
2025-05-14 18:28:46 +04:00 |
|
Lizonghang
|
ebd09fc83c
|
Merge branch 'dev'
|
2025-05-14 14:19:53 +04:00 |
|
Lizonghang
|
258fb2d06b
|
add QA: How to manually profile a device
|
2025-05-14 14:19:20 +04:00 |
|
Lizonghang
|
2fbc0c8da3
|
fix: reset -ngl to 0 when GPU is not used and reformat code
|
2025-05-14 13:27:20 +04:00 |
|