frostbyte_neo
  • 加入于 2026-07-22
84478e96d2 test(kda): vectorize MTP reference batch
ef9df14e69 test(kda): validate MTP schedules against FP32 recurrence
比较 2 提交 »
2026-09-10 00:35:15 +00:00
8e8fffde5b fix(moe): keep mixed-dtype gfx950 sigmoid top-k eligible
e7a83cb047 style: format gfx1250 sigmoid top-k registration
f00ea3ca0c perf(moe): add gfx1250 fused sigmoid top-k
比较 3 提交 »
2026-09-10 00:35:13 +00:00
9e827e828a fix(ci): update MI450 sim node id for projected MLA decode
be9dd1def5 style: format projected MLA files
ab4221df4d perf(amd): extend gfx1250 projected MLA decode
比较 3 提交 »
2026-09-10 00:35:12 +00:00
ff22b1ade8 style: format gfx1250 latent-input decode
14d0f35685 perf(moe): port gfx1250 M=1 latent-input decode
比较 2 提交 »
2026-09-10 00:35:11 +00:00
f8f6f2b54a fix(kimi3): tolerate missing gfx1250 AttnRes and snapshot graph replay
9439b5482d style: format gfx1250 linear AttnRes files
47cc483db8 perf(kimi3): port linear AttnRes partials to gfx1250
比较 3 提交 »
2026-09-10 00:35:10 +00:00
2026-09-10 00:35:08 +00:00
2026-09-10 00:35:08 +00:00
0422a3932e perf(cache): reduce cache block zeroing overhead (#1453)
dbd89e78fc perf(amd): Optimize GLM-5.3-Flash DSA decode on MI350X (#1433)
f5e0b0e571 fix(kda): preserve NVIDIA MTP launch schedule (#1467)
e67a4033c4 perf(moe): speed up gfx1250 MXFP4 prefill routing (#1463)
4d601173de refactor(kernel): move moe tactics script under benchmark (#1464)
比较 7 提交 »
2026-09-10 00:35:07 +00:00
b8be9fa658 docs(cache): remove Iris allocation note
edb9cc8f62 fix(kimi-k3): prepare reachable Iris states
b900924169 feat(kimi-k3): reserve Iris buffers before KV cache
f454577af1 fix(comm): preserve zero producer-direct capacity
比较 16 提交 »
2026-09-10 00:35:06 +00:00
6e4dd1f6d2 perf(comm): tune TP4 Iris decode launches
b640457632 perf(mhc): tune GLM-5.3-Flash reduction grid
19bceca275 perf(mhc): tune GLM-5.3-Flash prenorm projection
c1835d3cff perf(amd): reuse sigmoid routing kernels across row counts
d566362d7c perf(amd): fill small-route expert blocks directly
比较 12 提交 »
2026-09-10 00:35:05 +00:00
b640457632 perf(mhc): tune GLM-5.3-Flash reduction grid
19bceca275 perf(mhc): tune GLM-5.3-Flash prenorm projection
c1835d3cff perf(amd): reuse sigmoid routing kernels across row counts
d566362d7c perf(amd): fill small-route expert blocks directly
98fc0a58f6 perf(amd): tune GFX950 BF16 MoE reduction tiles
比较 11 提交 »
2026-09-10 00:35:05 +00:00
a4d2d9fc2f Merge branch 'main' into Max191/glm53-flash-stack-05-moe-0902
dbd89e78fc perf(amd): Optimize GLM-5.3-Flash DSA decode on MI350X (#1433)
f5e0b0e571 fix(kda): preserve NVIDIA MTP launch schedule (#1467)
e67a4033c4 perf(moe): speed up gfx1250 MXFP4 prefill routing (#1463)
c1835d3cff perf(amd): reuse sigmoid routing kernels across row counts
比较 13 提交 »
2026-09-10 00:35:04 +00:00
c74b37d377 Ensure acos(h)/asin(h) respect special cases from applicable standards (#1173)
2026-09-10 00:34:49 +00:00