分支列表

默认分支

0ba6b6aab7 · perf(mla): add opt-in packed Q decode for FP8 and FP16 (#1512) · 最后更新于 2026-09-12 09:04:01 +00:00

分支列表

97690af405 · test: document input pinning and record the kimi-k3 validation results · 最后更新于 2026-08-21 15:31:15 +00:00    frostbyte_neo

263
12

23c25aa28b · perf(attn_res): NVIDIA online-v2 fused kernel + combine load reorder with PDL · 最后更新于 2026-08-20 22:28:25 +00:00    frostbyte_neo

268
1

07212f2f82 · ci: align KVV evals with max effort · 最后更新于 2026-08-20 21:54:59 +00:00    frostbyte_neo

272
6

def2a87ab9 · refactor(rl): remove vLLM RL control plane, consolidate on SGLang dialect · 最后更新于 2026-08-20 21:53:56 +00:00    frostbyte_neo

267
1

338e38c889 · perf(k3): fuse the multi-token MoE front · 最后更新于 2026-08-20 15:37:54 +00:00    frostbyte_neo

268
1

de2a8035f4 · perf(k3): extend packed top-k to decode batches · 最后更新于 2026-08-20 13:12:46 +00:00    frostbyte_neo

269
1

eeb4a52ecf · perf(k3): expose aligned state checkpoints · 最后更新于 2026-08-20 13:12:46 +00:00    frostbyte_neo

269
1

bd7c3d35eb · Merge branch 'main' into k3-decode-gemv-route · 最后更新于 2026-08-20 02:50:49 +00:00    frostbyte_neo

277
3

d5fee37923 · perf(kimi3): enable batched AttnRes runtime fusion · 最后更新于 2026-08-19 20:09:16 +00:00    frostbyte_neo

284
16

ab93343c54 · perf(kimi3): shard the TP MoE final projection · 最后更新于 2026-08-19 20:09:08 +00:00    frostbyte_neo

284
15

9ddf36898f · perf(kimi3): join TP MoE reductions · 最后更新于 2026-08-19 20:08:59 +00:00    frostbyte_neo

284
14

43cd6034a4 · perf(comm): extend fused AttnRes through M16 · 最后更新于 2026-08-19 20:08:50 +00:00    frostbyte_neo

284
13

10863f1956 · perf(comm): add two-stage producer-direct reduction · 最后更新于 2026-08-19 20:08:36 +00:00    frostbyte_neo

284
12

377b28925f · perf(amd): tune TP-local A16W4 decode · 最后更新于 2026-08-19 20:08:22 +00:00    frostbyte_neo

284
11

bcebe08a43 · feat(moe): select gfx950 TP A8W4 SiTU · 最后更新于 2026-08-19 20:08:15 +00:00    frostbyte_neo

284
7

07fded4624 · feat(amd): add gfx950 A8W4 SiTU core · 最后更新于 2026-08-19 20:08:08 +00:00    frostbyte_neo

284
6

2c32487d76 · feat(amd): prepare gfx950 SiTU prefill · 最后更新于 2026-08-19 20:07:39 +00:00    frostbyte_neo

284
5

3a351c31d7 · perf(gemm): tune Kimi latent projection by batch · 最后更新于 2026-08-19 20:07:26 +00:00    frostbyte_neo

284
4

f36484c664 · perf(gemm): extend Kimi AttnRes projection to M4 · 最后更新于 2026-08-19 20:07:20 +00:00    frostbyte_neo

284
2

7759d332ae · perf(kda): consume strided decode projections · 最后更新于 2026-08-19 20:07:10 +00:00    frostbyte_neo

284
1