Default Branch

master
android / android_build (push) Has been cancelled
ios / ios_build (push) Has been cancelled
linux / linux_buil_test (push) Has been cancelled
macos / macos_buil_test (push) Has been cancelled
windows / windows_build_test (push) Has been cancelled

797e7d189d · [LLM:Feature] Add one-command iOS LLM benchmark and fix Metal4 tensor API on iOS 26.5 · Updated 2026-07-24 09:51:42 +00:00

Branches

7a7dcf52a8 · [LLM:Feature] Add one-command iOS LLM benchmark and fix Metal4 tensor API on iOS 26.5 · Updated 2026-07-24 09:50:31 +00:00    frostbyte_neo

1
1

714f25b19e · [Metal:Bugfix] Decide flash-attn prefill before KV alloc and fail KV cache expansion safely · Updated 2026-07-24 08:29:15 +00:00    frostbyte_neo

2
1

014906d7cd · [LLM:Bugfix] Do not use static global variables. · Updated 2026-07-23 06:36:26 +00:00    frostbyte_neo

3
1
feature/pymnn-ci-dependency-pins
pymnn-linux / pymnn_linux_buil_test (push) Has been cancelled
pymnn-macos / pymnn_macos_buil_test (push) Has been cancelled
pymnn-windows / pymnn_windows_buil_test (push) Has been cancelled

6c76c146bd · [Infra:Bugfix] Pin PyMNN CI test dependencies · Updated 2026-07-22 13:17:17 +00:00    frostbyte_neo

4
1
feature/opencl-sharedgather-md5
android / android_build (push) Has been cancelled
ios / ios_build (push) Has been cancelled
linux / linux_buil_test (push) Has been cancelled
macos / macos_buil_test (push) Has been cancelled
windows / windows_build_test (push) Has been cancelled

68af277df6 · [OpenCL:Chore] Update shared_gather_buf md5 in OpenCLProgramMd5Map after kernel fix · Updated 2026-07-22 11:28:56 +00:00    frostbyte_neo

5
1
feature/opencl-attention-bugfix
android / android_build (push) Has been cancelled
ios / ios_build (push) Has been cancelled
linux / linux_buil_test (push) Has been cancelled
macos / macos_buil_test (push) Has been cancelled
windows / windows_build_test (push) Has been cancelled

487a7317c4 · [OpenCL:Bugfix] Fix maskless attention setArg overflow (#4653), LinearAttention OOB writes, shared runtime gpuMode clobbering, int loop staging size and shared_gather build error · Updated 2026-07-22 09:59:45 +00:00    frostbyte_neo

7
1
feature/pymnn-python314-release
pymnn-linux / pymnn_linux_buil_test (push) Has been cancelled
pymnn-macos / pymnn_macos_buil_test (push) Has been cancelled
pymnn-windows / pymnn_windows_buil_test (push) Has been cancelled

7db91e9aa4 · [Infra:Release] Update PyMNN packaging for 3.6.1 release · Updated 2026-07-22 08:33:29 +00:00    frostbyte_neo

8
1

6b1b7ede7b · [Tools:Feature] llm_bench: add '-fa' meaning whether to use flash attention(default 1, using flash attention) · Updated 2026-07-22 08:25:45 +00:00    frostbyte_neo

7
1

037d02341c · [Infra:Release] Bump version to 3.6.1 · Updated 2026-07-22 02:29:14 +00:00    frostbyte_neo

9
1

3071344803 · [Metal:Bugfix] prefill_qkv_tensor missing ATTENTION_C4 output branch · Updated 2026-07-22 02:08:22 +00:00    frostbyte_neo

10
1
feature/cuda_realdiv_bugfix
android / android_build (push) Has been cancelled
ios / ios_build (push) Has been cancelled
linux / linux_buil_test (push) Has been cancelled
macos / macos_buil_test (push) Has been cancelled
windows / windows_build_test (push) Has been cancelled

a559ed4ebb · [CUDA:Bugfix] Fix int32 REALDIV not dispatched in BinaryBlit · Updated 2026-07-21 11:49:17 +00:00    frostbyte_neo

14
1

e266398532 · [Metal:Speed] Fused Q4/Q8 dequant+GEMM kernel + M_TILE=64 variant for LLM prefill on M5 · Updated 2026-07-21 11:33:33 +00:00    frostbyte_neo

11
1

825abe9c6e · * [Metal:Speed] LLM prefill/decode kernel + shape/RoPE infrastructure + Metal Op Profile. · Updated 2026-07-21 11:30:42 +00:00    frostbyte_neo

12
1

4d9ef932d3 · [LLM:Bugfix] Fix silent all-zero Q4 weights when quantizing large tensors on MPS · Updated 2026-07-21 11:27:38 +00:00    frostbyte_neo

13
1

f10ea652d7 · [Metal:Speed] Fused Q4/Q8 dequant+GEMM kernel + M_TILE=64 variant for LLM prefill on M5 · Updated 2026-07-21 11:24:32 +00:00    frostbyte_neo

13
1

4315807449 · [Metal:Speed] Fused Q4/Q8 dequant+GEMM kernel + M_TILE=64 variant for LLM prefill on M5 · Updated 2026-07-21 11:15:51 +00:00    frostbyte_neo

13
2

ee53bc97d8 · [Metal:Speed] Fused Q4/Q8 dequant+GEMM kernel + M_TILE=64 variant for LLM prefill on M5 · Updated 2026-07-21 10:57:08 +00:00    frostbyte_neo

13
1

66f706d936 · [LLM:Bugfix] Fix silent all-zero Q4 weights when quantizing large tensors on MPS · Updated 2026-07-21 10:54:10 +00:00    frostbyte_neo

13
1

1a98c3e712 · [Metal:Speed] Fused Q4/Q8 dequant+GEMM kernel + M_TILE=64 variant for LLM prefill on M5 · Updated 2026-07-21 10:51:48 +00:00    frostbyte_neo

13
1

7cf2f54b4d · [OpenCL:Bugfix] Fix opencl kernel compile bug for attention c4 · Updated 2026-07-20 06:43:03 +00:00    frostbyte_neo

15
1