Commits · cedccd59c94cb0c74e7ec0d0f6c791aed081febc · gaoqiong / composable_kernel_ROCM

23 Oct, 2024 1 commit
- [POST MERGE PR] Enable grouped conv bwd wei bf16 NGCHW (#1594) · cedccd59
  Bartłomiej Kocot authored Oct 23, 2024
  
  cedccd59
22 Oct, 2024 3 commits

Explicit cast values to half (#1593) · 4d5248e2
Jatin Chaudhary authored Oct 22, 2024
```
Co-authored-by: Illia Silin <98187287+illsilin@users.noreply.github.com>
```
4d5248e2
Enable grouped conv bwd wei bf16 NGCHW (#1589) · 82fc5383
Bartłomiej Kocot authored Oct 22, 2024
```
* Enable grouped conv bwd wei bf16 NGCHW

* fixes

* fixes

* Fixes

* fixes

* fixes

* Fixes
```
82fc5383

ltqin authored Oct 22, 2024

* port layernorm

* change warp_welford.hpp

* Update warpshuffle

* 1. Add save mean and save std back
2. Move construction of tensor_view and tile_window to operator()

* refine welford max count calculation

* unify layernorm api

* Rename file

* Remove save mean and inv std

* Revert "refine welford max count calculation"

This reverts commit 02236580

.

* Fix order of parameter

* refine welford max count calculation again

* Remove fp32 instances

* Fix bug of padding

* refactor api

* Support bf16

* Extract common function

* Refine arg of operator()

* Add kMThreadPerBlock to template parameter

* clang format

* Refine variable name

* Refine file name

* remove redundant line

* refactor layernorm2d pipeline and add block-per-block utility

* fix name

* rename more

* add more block-per-tile instance

* remove duplicated define

* update instance for 2048, 1024 case

* support up to 2048 now

* opt loading

* add n1536

* Add two pass pipeline

* format

* Fix incorrect type

* parallel compilation

* Use smaller N

* fix 2p pass

* Support Repeat_M in distribution

* Refine nameing

* Add reduce example

---------
Co-authored-by: letaoqin <letaoqin@amd.com>
Co-authored-by: aska-0096 <haocwang@amd.com>
Co-authored-by: rocking <ChunYu.Lai@amd.com>
Co-authored-by: carlushuang <carlus.huang@amd.com>

0394f8a7

21 Oct, 2024 5 commits

Update default stride (#1576) · 3f710930

Rostyslav Geyyer authored Oct 21, 2024

* Update default stride value to -1

* Fix format

* Revert "Fix format"

This reverts commit ae0c3649

.

---------
Co-authored-by: Harisankar Sadasivan <135730918+hsadasiv@users.noreply.github.com>

3f710930

added link to documentation (#1578) · 794f2d64
spolifroni-amd authored Oct 21, 2024

794f2d64

Bump rocm-docs-core from 1.8.2 to 1.8.3 in /docs/sphinx (#1587) · d0565e33

dependabot[bot] authored Oct 21, 2024

Bumps [rocm-docs-core](https://github.com/ROCm/rocm-docs-core) from 1.8.2 to 1.8.3.
- [Release notes](https://github.com/ROCm/rocm-docs-core/releases)
- [Changelog](https://github.com/ROCm/rocm-docs-core/blob/develop/CHANGELOG.md)
- [Commits](https://github.com/ROCm/rocm-docs-core/compare/v1.8.2...v1.8.3

)

---
updated-dependencies:
- dependency-name: rocm-docs-core
  dependency-type: direct:production
  update-type: version-update:semver-patch
...
Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>

d0565e33

Ck profiler instance support (#1575) · 560917b1

Thomas Ning authored Oct 21, 2024

* The draft on ckProfiler instance add

* support the ck profiler instance with same data types

* add a small feature on the M and N variable switch.

* Partially solve the incorrect result problem

* fix based on ci cd

560917b1

[CK_TILE] Optimize fmha splitkv & splitkv combine kernels (#1577) · 95e722a3

Po Yen Chen authored Oct 21, 2024

* Use smaller width for lse_accum dist tensor

* Update pipeline comment

* Fix wrong distribution for lse_accum

* Remove duplicate dim in lse_accum dist encoding

* Decide fmha splitkv combine kernel kBlockSize by kM0

* Remove assumption of MPerThread=1

* Add log<4> & log<8> specialization

* Enlarge occupancy array

* Fix vector size for small tile

* Add support for kMaxSplits=8

* Re-format gemm.hpp

* Use 16x16x16 warp gemm for fwd_splitkv

* Centralize policy code changes

* Leave fp8/bf8 tile settings unchanged

95e722a3

18 Oct, 2024 2 commits
- disable bad instance detected on MI308CPX (#1584) · a285d6f9
  Haocong WANG authored Oct 18, 2024
  
  a285d6f9
- add the lsr-drop-solution=1 compiler flag (#1582) · 88e6fa7f
  Illia Silin authored Oct 18, 2024
  
  88e6fa7f
16 Oct, 2024 1 commit

[CK_TILE] Improve headdim96 performance for fmha-bwd (#1573) · 14c3cfb1

Qianfeng authored Oct 16, 2024



* Add kQKHeaddimForGemmN and kVHeaddimForGemmN in order to support headdim 96

* Remove the using of MakeKRegBlockDescriptor and MakeVRegBlockDescriptor

* Fix in bwd_piple_default_policy

* Remove kQKHeaddim and rename kQKHeaddimForGemmN to kQKHeaddim in the bwd kernel and pipelines

* Replace kVHeaddimForGemmN by kVHeaddim and kDoDvHeaddim

* Update to hd96 tile settings

* Add smoke test scripts for fmha-bwd hd96

* Revert "Add smoke test scripts for fmha-bwd hd96"

This reverts commit 7ca7e1a93dc65eb99ce3ff4e82693589830e42a2.

* Remove hd96 tile settings in fmha_bwd codegen to save compiling

* Fix lost code line in bwd_pipeline_default_policy

* Merge kDoDvHeaddim/kPadHeadDimDoDv to kVHeaddim/kPadHeadDimV and remove TileFmhaBwdTraits

* Rename KRegSliceBlockDescriptor/VRegSliceBlockDescriptor to KRegBlockDescriptor/VRegBlockDescriptor

* tiny adjustments

---------
Co-authored-by: Po Yen Chen <PoYen.Chen@amd.com>
Co-authored-by: danyao12 <Dan.Yao@amd.com>

14c3cfb1

15 Oct, 2024 3 commits

Build codegen as standalone (#1556) · 10158b0f

Paul Fultz II authored Oct 15, 2024



* Build codegen as standalone

* Add exception for device tests

* Use local filesystem header

* add a codegen test CI stage and daily build

---------
Co-authored-by: illsilin <Illia.Silin@amd.com>
Co-authored-by: Illia Silin <98187287+illsilin@users.noreply.github.com>

10158b0f

[CK_TILE] Add block universal gemm pipeline policy (#1557) · d02a92cc
Bartłomiej Kocot authored Oct 15, 2024
```
* [CK_TILE] Add block universal gemm pipeline policy

* Fixes

* fixes2

* Fixes3

* fixeS
```
d02a92cc
Apply ROCm 6.2 WA to ROCm 6.3 and later (#1563) · 9868fd02
Po Yen Chen authored Oct 15, 2024

9868fd02

14 Oct, 2024 3 commits

Add custom type vector support (#1333) · 4cf70b36

Rostyslav Geyyer authored Oct 14, 2024



* Add non_native_vector_type

* Add a test

* Add non-native vector type

* Fix CTOR

* Fix non-native vector type of 1

* Fix CTORs

* Use vector_type to cover non-native implementation as well

* Update the test

* Format

* Format

* Fix copyright years

* Remove BoolVecT so far

* Add AsType test cases

* Update assert error message

* Remove redundant type

* Update naming

* Add complex half type with tests

* Add tests for vector reshaping

* Add missing alignas

* Update test/data_type/test_custom_type.cpp
Co-authored-by: Adam Osewski <19374865+aosewski@users.noreply.github.com>

* Compare custom types to built-in types

* Add default constructor test

* Add an alignment test

---------
Co-authored-by: Illia Silin <98187287+illsilin@users.noreply.github.com>
Co-authored-by: Adam Osewski <19374865+aosewski@users.noreply.github.com>
Co-authored-by: Po Yen Chen <PoYen.Chen@amd.com>

4cf70b36

Add transpose scale amax example (#1547) · f21cda25
Bartłomiej Kocot authored Oct 14, 2024
```
* Add transpose scale amax example

* fixes

* Tune reduce instance
```
f21cda25
decouple the calling from gemm_pipeline (#1571) · 35c1777d
Thomas Ning authored Oct 14, 2024
```
* decouple the calling from gemm_pipeline

* clang format
```
35c1777d

12 Oct, 2024 1 commit
- Implement GetWorkSpaceSize from BaseOperator. (#1564) · 29d384d0
  Adam Osewski authored Oct 12, 2024
  
  29d384d0
11 Oct, 2024 1 commit
- [CI] remove the --rm docker container flags (#1568) · 11444e4c
  Illia Silin authored Oct 11, 2024
  
  11444e4c
10 Oct, 2024 4 commits

only build tests and examples if user sets GPU_TARGETS (#1565) · f46a9eee
Illia Silin authored Oct 10, 2024

f46a9eee
removed API usage header (#1566) · 14c52bef
spolifroni-amd authored Oct 10, 2024

14c52bef
Fix default stride value (#1559) · d18fc079
Rostyslav Geyyer authored Oct 10, 2024

d18fc079

Ck tile gemm cshuffle & CK Tile GEMM restructure (#1535) · 6f27bc98

Thomas Ning authored Oct 10, 2024



* ake the cshuffle compilable

* modify Mhe reference on gpu and cpu. Correaccess of cshuffle

* fix the cpu reference code

* Complete the in tile shuffle logic

* restructure the kernel template input

* change the naming pattern of ck_tile gemm pipeline

* Re-format files using remod.py

* Solve the fmha conflict with gemm

* Comment Addressed from Carlus

---------
Co-authored-by: Po Yen, Chen <PoYen.Chen@amd.com>

6f27bc98

09 Oct, 2024 3 commits
- fix the target selection logic (#1561) · 2e1165c1
  Illia Silin authored Oct 09, 2024
  
  2e1165c1
- remove gfx12 targets from daily builds with rocm6.2 (#1560) · cfac9497
  Illia Silin authored Oct 09, 2024
  
  cfac9497
- Fixes small memory leak from missing hipEventDestroy (#1554) · ceaed8e0
  Christopher Millette authored Oct 09, 2024
  
  ceaed8e0
08 Oct, 2024 3 commits

Add a gpu gemm reference kernel (#1528) · aa932445

Rostyslav Geyyer authored Oct 08, 2024



* Add a gpu gemm reference kernel

* Switch to gpu reference in gemm examples

* Remove redundant arguments

* Update all related examples

* Update more examples

* Try less threads per block

* Try even less threads per block

* Add support for all matrix layouts

* Increase block size

* Clean up

* Remove hardcoded strides

* Clean up

* Try a column-major case

* Revert back to row-major

* Run both CPU and GPU veriffication

---------
Co-authored-by: Po Yen Chen <PoYen.Chen@amd.com>

aa932445

[CK_TILE] Update example README files & fix script compatibility issue (#1548) · 0c094daa

Po Yen Chen authored Oct 08, 2024

* Fix text alignment of ArgParser::print()

* Update example README files

* Clarify make-ck-dev.sh <arch> usage

* Only keep some of the argument from '-?' output

* Undo command line output changes in README

* Only keep existing argument on doc and update description

* Fix text alignment

* Make cmake-ck-*.sh compatible with 'sh' command

0c094daa

[CK_TILE] Simplify the codes in splitkv_combine pipeline (#1549) · 74d68e3b

Qianfeng authored Oct 08, 2024



* Simplify the codes in splitkv_combine pipeline

* Always set kPadSeqLenK=true for fmha splitkv kernels

* Change in Oacc Alignment and TileDistribution to be more adaptable to tile sizes

---------
Co-authored-by: Po Yen Chen <PoYen.Chen@amd.com>

74d68e3b

07 Oct, 2024 4 commits

add a CK_USE_CODEGEN build argument to enable codegen (#1552) · 7733ae16
Illia Silin authored Oct 07, 2024
```
* add a CK_USE_CODEGEN build argument to enable codegen

* fix cmake codegen logic
```
7733ae16

Fix build logic using GRU_ARCHS. (#1536) · 7d8ea5f0

Illia Silin authored Oct 07, 2024

* update build logic with GPU_ARCHS

* fix the GPU_ARCHS build for codegen

* unset GPU_TARGETS when GPU_ARCHS are set

7d8ea5f0

[CK_TILE] Fix conv param multiple definition (#1550) · cc8f466a
Bartłomiej Kocot authored Oct 07, 2024
```
Co-authored-by: Po Yen Chen <PoYen.Chen@amd.com>
```
cc8f466a

[Ck tile] Support layernorm one pass (#1512) · 0023f01a

rocking authored Oct 07, 2024



* Fix compile error

* Add one pass pipeline

* Extract creating tile_window to operator()

* clang format

* reduce duplicated code

* do not hardcode

* Support padding in layernorm

---------
Co-authored-by: Po Yen Chen <PoYen.Chen@amd.com>

0023f01a

04 Oct, 2024 3 commits

Adding seed and offset pointer support to the philox random number generator. (#1523) · c24fae23

kylasa authored Oct 04, 2024



* Adding seed and offset pointer support to the philox random number generator.

* Separating seed and offset pointer checks with different condition statements.

* Changes include, adding support for device seed and offset pointers, union is used to store seed/offset values and device pointers to minimize device SGPRs.

* Correcting a typo in the readme file

* Re-format files using remod.py

* Use STL type for API parameters

* Use simpler struct design for drop_seed & drop_offset

* Undo unnecessary changes

* Sync kargs style for fmha_fwd.hpp/.cpp

* Use templated union to reduce code

* Use structured binding to make code more readable

---------
Co-authored-by: Sudhir Kylasa <sukylasa@amd.com>
Co-authored-by: Po Yen Chen <PoYen.Chen@amd.com>

c24fae23

Codegen build (#1526) · b545de17

arai713 authored Oct 04, 2024

* updating codegen build for MIOpen access: adding .cmake for codegen component

(cherry picked from commit 652a7c04)

* updating CMake

(cherry picked from commit a685822e)

b545de17

Fix grouped gemm check to avoid overflow (#1545) · 6b54d2fa
Bartłomiej Kocot authored Oct 04, 2024

6b54d2fa

02 Oct, 2024 2 commits

Fix compilation errors generated by forthcoming Clang changes (#1544) · aeb7c91f

macurtis-amd authored Oct 02, 2024

Without this change, the following diagnostic is generated:
  a template argument list is expected after a name prefixed by the template
  keyword [-Wmissing-template-arg-list-after-template-kw]

See C++17 spec [temp.names] p5.

aeb7c91f

Add generating mha static library for gfx90a (#1540) · 294cb823
BrianHarrisonAMD authored Oct 02, 2024
```
* Add generating mha static library for gfx90a

* Update comment to reflect changes
```
294cb823

01 Oct, 2024 1 commit
- re-enable the FMHA performance monitoring (#1539) · 11b7a4db
  Illia Silin authored Oct 01, 2024
  
  11b7a4db