1. 21 Oct, 2024 3 commits
    • carlushuang's avatar
      opt loading · 009cce41
      carlushuang authored
      009cce41
    • Po Yen Chen's avatar
      [CK_TILE] Optimize fmha splitkv & splitkv combine kernels (#1577) · 95e722a3
      Po Yen Chen authored
      * Use smaller width for lse_accum dist tensor
      
      * Update pipeline comment
      
      * Fix wrong distribution for lse_accum
      
      * Remove duplicate dim in lse_accum dist encoding
      
      * Decide fmha splitkv combine kernel kBlockSize by kM0
      
      * Remove assumption of MPerThread=1
      
      * Add log<4> & log<8> specialization
      
      * Enlarge occupancy array
      
      * Fix vector size for small tile
      
      * Add support for kMaxSplits=8
      
      * Re-format gemm.hpp
      
      * Use 16x16x16 warp gemm for fwd_splitkv
      
      * Centralize policy code changes
      
      * Leave fp8/bf8 tile settings unchanged
      95e722a3
    • carlushuang's avatar
      remove duplicated define · 24bebf15
      carlushuang authored
      24bebf15
  2. 20 Oct, 2024 4 commits
  3. 16 Oct, 2024 6 commits
  4. 15 Oct, 2024 2 commits
  5. 14 Oct, 2024 5 commits
  6. 12 Oct, 2024 4 commits
  7. 10 Oct, 2024 1 commit
    • Thomas Ning's avatar
      Ck tile gemm cshuffle & CK Tile GEMM restructure (#1535) · 6f27bc98
      Thomas Ning authored
      
      
      * ake the cshuffle compilable
      
      * modify Mhe reference on gpu and cpu. Correaccess of cshuffle
      
      * fix the cpu reference code
      
      * Complete the in tile shuffle logic
      
      * restructure the kernel template input
      
      * change the naming pattern of ck_tile gemm pipeline
      
      * Re-format files using remod.py
      
      * Solve the fmha conflict with gemm
      
      * Comment Addressed from Carlus
      
      ---------
      Co-authored-by: default avatarPo Yen, Chen <PoYen.Chen@amd.com>
      6f27bc98
  8. 09 Oct, 2024 1 commit
  9. 08 Oct, 2024 2 commits
  10. 07 Oct, 2024 3 commits
  11. 04 Oct, 2024 2 commits
  12. 02 Oct, 2024 1 commit
  13. 01 Oct, 2024 2 commits
  14. 27 Sep, 2024 1 commit
  15. 26 Sep, 2024 1 commit
  16. 25 Sep, 2024 1 commit
  17. 22 Sep, 2024 1 commit