1. 04 Feb, 2026 2 commits
  2. 02 Feb, 2026 1 commit
  3. 23 Jan, 2026 2 commits
  4. 15 Jan, 2026 1 commit
  5. 30 Dec, 2025 1 commit
  6. 26 Dec, 2025 1 commit
  7. 23 Dec, 2025 2 commits
  8. 15 Dec, 2025 1 commit
  9. 12 Dec, 2025 1 commit
  10. 25 Nov, 2025 1 commit
  11. 14 Nov, 2025 1 commit
  12. 07 Nov, 2025 1 commit
  13. 06 Nov, 2025 1 commit
  14. 05 Nov, 2025 1 commit
  15. 03 Nov, 2025 1 commit
  16. 30 Oct, 2025 1 commit
  17. 17 Oct, 2025 1 commit
  18. 24 Sep, 2025 1 commit
  19. 17 Sep, 2025 1 commit
  20. 16 Sep, 2025 1 commit
    • Chenggang Zhao's avatar
      Canonicalize TMA usages (#410) · 2012e310
      Chenggang Zhao authored
      * Remove redundant TMA flushes
      
      * Less barrier initialization overhead
      
      * Simplify `elect_one_sync`
      
      * Use `elect_one_sync` instead of lanes
      
      * Minor fix
      
      * Polish testing prints
      
      * Refactor for internode kernels
      
      * Better performance
      2012e310
  21. 15 Sep, 2025 1 commit
  22. 01 Sep, 2025 1 commit
  23. 28 Aug, 2025 1 commit
  24. 25 Aug, 2025 2 commits
  25. 07 Aug, 2025 2 commits
    • Chenggang Zhao's avatar
      Fix compilation · dd14b36d
      Chenggang Zhao authored
      dd14b36d
    • Zhean Xu's avatar
      Support 10-bit LogFMT Combine (#345) · c5facf5c
      Zhean Xu authored
      
      
      * independent logfmt_simulate function
      
      * draft: logfmt low latency combine
      
      * Minor bug fixes
      
      * Fix non-logfmt bugs
      
      * Fix logfmt bugs
      
      * Fix logfmt bugs
      
      * Minor fix
      
      * Minor fix
      
      * Clean code
      
      * Clean code
      
      * Use fewer regs
      
      * Use two warp groups
      
      * Correct shared memory size
      
      * Minor fix
      
      * Minor fix
      
      * More rigorous tests
      
      * Clean code
      
      * Use more SMs
      
      * Use different unroll factor for send & recv
      
      * Update csrc/kernels/internode_ll.cu
      Co-authored-by: default avatarCopilot <175728472+Copilot@users.noreply.github.com>
      
      * Update csrc/kernels/internode_ll.cu
      Co-authored-by: default avatarCopilot <175728472+Copilot@users.noreply.github.com>
      
      * Some renaming
      
      * Some comments of tests
      
      * Format `logfmt_encode`
      
      * More lints
      
      * Some refactors on sends
      
      * Fix testing
      
      * Fix bugs
      
      * Renaming
      
      * Use the full warp
      
      * Unify combine recv
      
      * Lint
      
      * Lint
      
      * Support 2560
      
      * Fix meta buffer dtype
      
      * Better encode calls
      
      * Better amin/max writes
      
      * Extra sync
      
      * Read `topk_idx` by once
      
      * Better specialization
      
      * Read weights by once
      
      * Rename
      
      * Bug fixed
      
      * Some renaming
      
      * Fix local memory usage for sending
      
      * Fix local memory usage for receiving
      
      * Less writes
      
      * Optimize performance
      
      * Optimize performance
      
      * Better performance
      
      * Optimize performance
      
      * Fix rounding
      
      * Manually unroll
      
      * Fix bench
      
      ---------
      Co-authored-by: default avatarCopilot <175728472+Copilot@users.noreply.github.com>
      Co-authored-by: default avatarChenggang Zhao <chenggangz@deepseek.com>
      c5facf5c
  26. 31 Jul, 2025 1 commit
  27. 30 Jul, 2025 1 commit
  28. 14 Jul, 2025 1 commit
  29. 10 Jul, 2025 1 commit
    • Chenggang Zhao's avatar
      Support 10-bit LogFMT (simulated version) (#284) · 1cf85fb2
      Chenggang Zhao authored
      
      
      * Add LogFMT interface
      
      * Update comments
      
      * Add simulated code
      
      * Fix comments
      
      * Change to 128 channels
      
      * Add notes
      
      * Optimize performance
      
      * optimize simulate logfmt 10bit
      
      * Minor fix
      
      * Stronger low latency tests
      
      * Minor fix
      
      * Stronger low latency tests for logfmt
      
      * Optimize logfmt simulate: lg2/ex2 ptx, step_inv
      
      * Minor fix
      
      * Minor fix
      
      * Add non-logfmt bench
      
      * Fix value=0 for logfmt
      
      * Optimize performance
      
      * Refactor tests
      
      ---------
      Co-authored-by: default avatarZhean Xu <xza@deepseek.com>
      1cf85fb2
  30. 02 Jul, 2025 2 commits
  31. 23 Jun, 2025 1 commit
  32. 16 Jun, 2025 2 commits
  33. 12 Jun, 2025 1 commit