Commits · 3c305feebd0da0cef7d8c862a869a9ce0d2aaf1c · gaoqiong / composable_kernel_ROCM

"include/vscode:/vscode.git/clone" did not exist on "1a2a167914c69ebd7ecf7015027d5bac57516d89"

30 Oct, 2024 3 commits

support fused dynamic-quant · 3c305fee
carlushuang authored Oct 30, 2024

3c305fee

[Ck tile] support rmsnorm and related fusion (#1605) · 3d609534

rocking authored Oct 30, 2024

* Add reduce2d new api

* Prevent user use cross warp reduction

* Fix bug of std caculation

* Add rmsnorm2d

* Add rmsnorm small example

* Remove static assert to prevent compile fail

* Add script to test performance and correctness

* Add missing cmake change

* refine naming

* refine example of rmsnorm

* Fix bug of rmsnorm

* Refine naming

* Fix cmake

* clang format

* Refine pipeline name

* Add add_rmsnorm2d_rdquant kernel

* Add reduce op

* host verification

* Fix bug of one pass pipeline

* Refine tile size

* Add two pass pipeline

* Rename two pass to three pass

* Fix bug of kSaveX == false

* Add instance library

* Add test script

* Fix bug of x verification

* Add save_x to trait

* Add README

* Move reduce2d into reduce folder

* Fix bug of welford when number of m warp > 1

* remove reduncant comment

* 1. move 06_rmsnorm2d to 10_rmsnorm2d
2. move 07_add_rmsnorm2d_rdquant to 11_add_rmsnorm2d_rdquant

* clang format and add missing header

* Add host validation of add + layernorm2d + rsquant

* Revert "Add host validation of add + layernorm2d + rsquant"

This reverts commit 936cb457978b928b90eff89a08fcdb7dc8bbed67.

* Remove deprecated flag

3d609534

[CK_TILE] Add fmha fwd headdim96 support (#1608) · 86322218

Qianfeng authored Oct 30, 2024



* Add ceil_to_qualified_tile_length()

* Rename kK0BlockLength to kQKHeaddim

* Add kSubQKHeaddim concept to support headdim96

* Fix in math.hpp to avoid using __half interfaces

* Add LdsBufferSequence instance for headdim96

* Update in fmha_fwd/fmha_fwd_splitkv codegen to support hd96 testing

* Disable hd96 instance generation in codegen fmha_fwd and fmha_fwd_splitkv to save compiling time

* Reformat one file

* Fix text alignment in fmha_fwd_splitkv.py

---------
Co-authored-by: Po Yen Chen <PoYen.Chen@amd.com>

86322218

29 Oct, 2024 3 commits
- dynamic-quant ready · 9e063018
  carlushuang authored Oct 30, 2024
  
  9e063018
- [CK_TILE] add scatter_gather (#1609) · 4d7e063a
  valarLip authored Oct 29, 2024
  
  4d7e063a
- [CK_TILE] add generic_permute (#1607) · 9fbd72e9
  valarLip authored Oct 29, 2024
  
  9fbd72e9
28 Oct, 2024 6 commits
- format and update n4096 · e2935465
  carlushuang authored Oct 28, 2024
  
  e2935465
- use non-raw for loading · 62557466
  carlushuang authored Oct 28, 2024
  
  62557466
- format · 865edada
  carlushuang authored Oct 28, 2024
  
  865edada
- update format · e90cf8f9
  carlushuang authored Oct 28, 2024
  
  e90cf8f9
- update some description and fix format · 3141f9f9
  carlushuang authored Oct 28, 2024
  
  3141f9f9
- fix format · 6451d7fa
  carlushuang authored Oct 28, 2024
  
  6451d7fa
27 Oct, 2024 1 commit
- add prenorm/postnorm support, refactor using generate.py · 4b59b5c9
  carlushuang authored Oct 27, 2024
  
  4b59b5c9
26 Oct, 2024 2 commits

topk_softmax (#1592) · b098b71b

carlushuang authored Oct 26, 2024

* topk_softmax

* remove some file

* fix atomix linear_offset

* address various comment, and change sfc get_index api to static(tuple)

b098b71b

[CK_TILE] More fmha splitkv optimizations (#1588) · 54f0e6f4

Po Yen Chen authored Oct 26, 2024

* Use pre-defined constants for readability

* Use vector write for o_acc tensor

* Remove no-longer used policy method

* Deprecate no-longer used policy/pipeline

* Specify gemm0/gemm1 block warps separately in codegen

* Fix wrong ps_idx creation logic

* Add single-warp block gemm

* Supoprt single-warp gemm0

* Make MakeCBlockTile() as static method

* Use MakeCBlockTile() to get underlying tile distribution

* Use kNumGemm1Warps to compute # threads for gemm1

* Put normal case in the if clause

* Refine fmha splitkv block mapping

* Refine & fix the lse_acc/o_acc layout

* Fix wrong LDS size for K tile

* Use kK0=64 for hdim=128,256 fmha splitkv kernels

* Use kK1=64 for hdim=32,64,128 fmha splitkv kernels

* Undo kK0/kK1 changes

* Use more reasonable GetAlignmentV() computation

* Using store_tile() in fmha splitkv kernel epilogue

54f0e6f4

25 Oct, 2024 1 commit
- hot_fix epsilon pos (#1597) · 9183ce69
  dummycoderfe authored Oct 25, 2024
```
Co-authored-by: dummycoderfe <noplydummmycoder@163.com>
```
  9183ce69
22 Oct, 2024 1 commit

update layernorm (#1570) · 0394f8a7

ltqin authored Oct 22, 2024

* port layernorm

* change warp_welford.hpp

* Update warpshuffle

* 1. Add save mean and save std back
2. Move construction of tensor_view and tile_window to operator()

* refine welford max count calculation

* unify layernorm api

* Rename file

* Remove save mean and inv std

* Revert "refine welford max count calculation"

This reverts commit 02236580

.

* Fix order of parameter

* refine welford max count calculation again

* Remove fp32 instances

* Fix bug of padding

* refactor api

* Support bf16

* Extract common function

* Refine arg of operator()

* Add kMThreadPerBlock to template parameter

* clang format

* Refine variable name

* Refine file name

* remove redundant line

* refactor layernorm2d pipeline and add block-per-block utility

* fix name

* rename more

* add more block-per-tile instance

* remove duplicated define

* update instance for 2048, 1024 case

* support up to 2048 now

* opt loading

* add n1536

* Add two pass pipeline

* format

* Fix incorrect type

* parallel compilation

* Use smaller N

* fix 2p pass

* Support Repeat_M in distribution

* Refine nameing

* Add reduce example

---------
Co-authored-by: letaoqin <letaoqin@amd.com>
Co-authored-by: aska-0096 <haocwang@amd.com>
Co-authored-by: rocking <ChunYu.Lai@amd.com>
Co-authored-by: carlushuang <carlus.huang@amd.com>

0394f8a7

21 Oct, 2024 1 commit

[CK_TILE] Optimize fmha splitkv & splitkv combine kernels (#1577) · 95e722a3

Po Yen Chen authored Oct 21, 2024

* Use smaller width for lse_accum dist tensor

* Update pipeline comment

* Fix wrong distribution for lse_accum

* Remove duplicate dim in lse_accum dist encoding

* Decide fmha splitkv combine kernel kBlockSize by kM0

* Remove assumption of MPerThread=1

* Add log<4> & log<8> specialization

* Enlarge occupancy array

* Fix vector size for small tile

* Add support for kMaxSplits=8

* Re-format gemm.hpp

* Use 16x16x16 warp gemm for fwd_splitkv

* Centralize policy code changes

* Leave fp8/bf8 tile settings unchanged

95e722a3

16 Oct, 2024 1 commit

[CK_TILE] Improve headdim96 performance for fmha-bwd (#1573) · 14c3cfb1

Qianfeng authored Oct 16, 2024



* Add kQKHeaddimForGemmN and kVHeaddimForGemmN in order to support headdim 96

* Remove the using of MakeKRegBlockDescriptor and MakeVRegBlockDescriptor

* Fix in bwd_piple_default_policy

* Remove kQKHeaddim and rename kQKHeaddimForGemmN to kQKHeaddim in the bwd kernel and pipelines

* Replace kVHeaddimForGemmN by kVHeaddim and kDoDvHeaddim

* Update to hd96 tile settings

* Add smoke test scripts for fmha-bwd hd96

* Revert "Add smoke test scripts for fmha-bwd hd96"

This reverts commit 7ca7e1a93dc65eb99ce3ff4e82693589830e42a2.

* Remove hd96 tile settings in fmha_bwd codegen to save compiling

* Fix lost code line in bwd_pipeline_default_policy

* Merge kDoDvHeaddim/kPadHeadDimDoDv to kVHeaddim/kPadHeadDimV and remove TileFmhaBwdTraits

* Rename KRegSliceBlockDescriptor/VRegSliceBlockDescriptor to KRegBlockDescriptor/VRegBlockDescriptor

* tiny adjustments

---------
Co-authored-by: Po Yen Chen <PoYen.Chen@amd.com>
Co-authored-by: danyao12 <Dan.Yao@amd.com>

14c3cfb1

15 Oct, 2024 2 commits
- [CK_TILE] Add block universal gemm pipeline policy (#1557) · d02a92cc
  Bartłomiej Kocot authored Oct 15, 2024
```
* [CK_TILE] Add block universal gemm pipeline policy

* Fixes

* fixes2

* Fixes3

* fixeS
```
  d02a92cc
- Apply ROCm 6.2 WA to ROCm 6.3 and later (#1563) · 9868fd02
  Po Yen Chen authored Oct 15, 2024
  
  9868fd02
14 Oct, 2024 1 commit
- decouple the calling from gemm_pipeline (#1571) · 35c1777d
  Thomas Ning authored Oct 14, 2024
```
* decouple the calling from gemm_pipeline

* clang format
```
  35c1777d
10 Oct, 2024 1 commit

Ck tile gemm cshuffle & CK Tile GEMM restructure (#1535) · 6f27bc98

Thomas Ning authored Oct 10, 2024



* ake the cshuffle compilable

* modify Mhe reference on gpu and cpu. Correaccess of cshuffle

* fix the cpu reference code

* Complete the in tile shuffle logic

* restructure the kernel template input

* change the naming pattern of ck_tile gemm pipeline

* Re-format files using remod.py

* Solve the fmha conflict with gemm

* Comment Addressed from Carlus

---------
Co-authored-by: Po Yen, Chen <PoYen.Chen@amd.com>

6f27bc98

08 Oct, 2024 2 commits

[CK_TILE] Update example README files & fix script compatibility issue (#1548) · 0c094daa

Po Yen Chen authored Oct 08, 2024

* Fix text alignment of ArgParser::print()

* Update example README files

* Clarify make-ck-dev.sh <arch> usage

* Only keep some of the argument from '-?' output

* Undo command line output changes in README

* Only keep existing argument on doc and update description

* Fix text alignment

* Make cmake-ck-*.sh compatible with 'sh' command

0c094daa

[CK_TILE] Simplify the codes in splitkv_combine pipeline (#1549) · 74d68e3b

Qianfeng authored Oct 08, 2024



* Simplify the codes in splitkv_combine pipeline

* Always set kPadSeqLenK=true for fmha splitkv kernels

* Change in Oacc Alignment and TileDistribution to be more adaptable to tile sizes

---------
Co-authored-by: Po Yen Chen <PoYen.Chen@amd.com>

74d68e3b

07 Oct, 2024 2 commits

[CK_TILE] Fix conv param multiple definition (#1550) · cc8f466a
Bartłomiej Kocot authored Oct 07, 2024
```
Co-authored-by: Po Yen Chen <PoYen.Chen@amd.com>
```
cc8f466a

[Ck tile] Support layernorm one pass (#1512) · 0023f01a

rocking authored Oct 07, 2024



* Fix compile error

* Add one pass pipeline

* Extract creating tile_window to operator()

* clang format

* reduce duplicated code

* do not hardcode

* Support padding in layernorm

---------
Co-authored-by: Po Yen Chen <PoYen.Chen@amd.com>

0023f01a

04 Oct, 2024 1 commit

Adding seed and offset pointer support to the philox random number generator. (#1523) · c24fae23

kylasa authored Oct 04, 2024



* Adding seed and offset pointer support to the philox random number generator.

* Separating seed and offset pointer checks with different condition statements.

* Changes include, adding support for device seed and offset pointers, union is used to store seed/offset values and device pointers to minimize device SGPRs.

* Correcting a typo in the readme file

* Re-format files using remod.py

* Use STL type for API parameters

* Use simpler struct design for drop_seed & drop_offset

* Undo unnecessary changes

* Sync kargs style for fmha_fwd.hpp/.cpp

* Use templated union to reduce code

* Use structured binding to make code more readable

---------
Co-authored-by: Sudhir Kylasa <sukylasa@amd.com>
Co-authored-by: Po Yen Chen <PoYen.Chen@amd.com>

c24fae23

01 Oct, 2024 2 commits

[CK_TILE] add missing vector header (#1537) · 8e4c3fb1

Illia Silin authored Oct 01, 2024



* add missing vector header

* Re-format header using remod.py

---------
Co-authored-by: Po Yen, Chen <PoYen.Chen@amd.com>

8e4c3fb1

[CK_TILE] Change output accum tensor layout of fmha fwd split-kv & combine kernels (#1527) · a1c07e8d

Po Yen Chen authored Oct 01, 2024

* Use same layout for o_acc and o tensor

* Use better param names in partitioner

* Remove redundant kargs 'max_seqlen_q'

* Use better param names in splitkv kernel

* Add comment for additional kernel arguments

* Sync empty loop early return logics between pipelines

* Pass more arguments to cmake in scripts

* Align backslashes

* Fix wrong o_acc tensor view strides

* Change o_acc layout if o_perm=0

* Handle whole row masked via attn_bias

* Use use vector width = 1 for o_acc

* Use more even split sizes

a1c07e8d

27 Sep, 2024 1 commit

[CK_TILE] Image to Column kernel (#1532) · de3e3b64

Bartłomiej Kocot authored Sep 27, 2024

* [CK_TILE] Image to Column kernel

* Fixes

* Vector loads and stores

* Fixes

* Fixes

* change test dir name

de3e3b64

26 Sep, 2024 1 commit
- [CK_TILE] Fix compiler related FA bwd issues (#1530) · 9d69a099
  Dan Yao authored Sep 27, 2024
```
* add barriers

* tail bias barriers

* adjust bf16/hd256 tol

* continue adjust bf16/hd256 tol
```
  9d69a099
25 Sep, 2024 1 commit
- Fix compilation errors with Clang20.0. (#1533) · 42e6dcea
  Illia Silin authored Sep 25, 2024
```
* fix clang20 compilation errors for gfx90a

* fix clang20 compilation errors for gfx11 targets
```
  42e6dcea
22 Sep, 2024 1 commit
- Early return if seqlen_k=0 on group mode (#1524) · 770d2b77
  Po Yen Chen authored Sep 22, 2024
  
  770d2b77
18 Sep, 2024 1 commit

Ck tile gemm padding dim (#1516) · 694c3001

Thomas Ning authored Sep 18, 2024

* Support the N dimension padding

* Finished the padding feature for different dimension of K

694c3001

14 Sep, 2024 1 commit

Ck tile GPU verification sample develop & Add the CK TILE GEMM to the CI/CD test (#1505) · 844f5a17

Thomas Ning authored Sep 14, 2024



* Finished the feature of gpu verification

* Add the ck_tile_gemm test in the CI CD

* add the include of tensor_layou in reference_gemm

* Comment Addressed

* split ck_tile fhma and gemm tests into separate stages

* restructure the reference gemm

* restructure a new reference_gemm api that could read the device mem

---------
Co-authored-by: carlushuang <carlus.huang@amd.com>
Co-authored-by: illsilin <Illia.Silin@amd.com>

844f5a17

10 Sep, 2024 1 commit
- [CK_TILE] FA bwd repair (#1502) · d09572e8
  Dan Yao authored Sep 11, 2024
```
* fix fa bwd

* revert kernelBlockSize in gemm_kernel.hpp
```
  d09572e8
07 Sep, 2024 1 commit

Ck tile gemm example (#1488) · caacd388

Thomas Ning authored Sep 07, 2024



* Checkpoint: Finished with the tile example & kernel verification, working on the different matrix layout

* Finished the Matrix Layout feature set up. Note: Need to modify the inner block to solve the shuffle problem in the future.

* Fix: Clang Format, API fixed from fmha

* fix with better naming convention

* revert back the pipeline code of fmha

* Fixed: Addressed the comments and merge the GEMM shape of GEMM Operator and FMHA Operator to one.

* clang format with the reference_gemm file

* convert the clang format with the remod.py

* Changed the format and variable name of the kernel gemm_shape and partitioner

---------
Co-authored-by: thomasning <thomasning@banff-cyxtera-s70-4.ctr.dcgpu>

caacd388

30 Aug, 2024 2 commits
- [CK_TILE] float -> bf16 inline asm rtn (#1482) · b8addae2
  Dan Yao authored Aug 30, 2024
```
* asm rtn

* add asm rtn macro

* reorder macro

---------
Co-authored-by: carlushuang <carlus.huang@amd.com>
```
  b8addae2
- Enable scratch memory workaround on ROCm 6.2 (#1486) · 461ec98d
  Po Yen Chen authored Aug 30, 2024
```
Co-authored-by: carlushuang <carlus.huang@amd.com>
```
  461ec98d