Commits · 80f0da377f0bc80ef717e62c18a454e7f987dcf6 · gaoqiong / composable_kernel

01 Dec, 2023 6 commits
- fix build (#46) · 80f0da37
  Chao Liu authored Dec 01, 2023
  
  80f0da37
- Ck tiled main fixing for building xformers (#44) · ca2105a4
  Qianfeng authored Dec 02, 2023
```
* Add include/ck/config.h to support xformers c++ extension building

* Disable exp() and log() overloading for half_t to support xformers C++ extension building

* config.h.default

---------
Co-authored-by: Chao Liu <chao.liu2@amd.com>
```
  ca2105a4
- format (#45) · 99c9d3b7
  Chao Liu authored Dec 01, 2023
  
  99c9d3b7
- Merge remote-tracking branch 'upstream/develop' into merge_upstream_1129 · e94e60e8
  Chao Liu authored Dec 01, 2023
  
  e94e60e8
- fix bug · 87c7888e
  Chao Liu authored Nov 30, 2023
  
  87c7888e
- update cmake · 13542e66
  Chao Liu authored Nov 30, 2023
  
  13542e66
30 Nov, 2023 5 commits
- [CI] Update Jenkinsfile (#1073) · c7d5c772
  Jun Liu authored Nov 30, 2023
  
  c7d5c772
- Fixed GroupedGemmFixedNK with hipGraph (#1065) · 49df1dc5
  zjing14 authored Nov 30, 2023
```
* fixed examples; add async_mem_set

* add stream to all deviceOp using SetWorkspace

---------
Co-authored-by: Jing Zhang <jizha@amd.com>
```
  49df1dc5
- Merge remote-tracking branch 'upstream/develop' into merge_upstream_1129 · 0bdbd358
  Chao Liu authored Nov 30, 2023
  
  0bdbd358
- Merge remote-tracking branch 'upstream/develop' into merge_upstream_1129 · d27e0691
  Chao Liu authored Nov 30, 2023
```
also fix regression
```
  d27e0691
- Introduce wrapper for layout (#1054) · 8ff845f2
  Bartłomiej Kocot authored Nov 30, 2023
```
* Introduce wrapper for layout

* Extend functionality

* Fix for getLength

* Comment fixes

* Add comments and remove not needed getters
```
  8ff845f2
29 Nov, 2023 1 commit

Disable transpose device op for MI300 (#1050) · a2969aa8

arai713 authored Nov 29, 2023



* added working example for 5D input using 1D kernel

* example with 5D input tensor and 2d kernel - not working: issues with arguments

* added updated version of 3d device op - changed descriptors/dims

* added example file to check kernel

* fixed descriptor and isSupportedArgument stride problem

* added and modified kernel for 3d - updated tids/loop

* adding some more 5d example files

* fixed some issues

* changes made for testing

* working version: fixed error in stride for A, still a bit inefficient

* cleaned up formatting/comments

* updating formatting

* more formatting fixes

* fixing cmake, adding back gpu targets in cmake script

* adding client example

* added instances for client example

* fixed errors in client example

* implemented client ex with device_elementwise.hpp and device_elementwise_3d_impl.hpp

* removed extra files

* minor formatting and naming fixes

* adding test files and profiler

* fixing minor error

* minor fix

* removed unneccesary comments, renamed files

* updated instance list for client example, added different layout example

* removing instances

* fixed error in instance generation

* remove comments

* update profiler and client example tensor layouts

* fixed errors in test/profiler

* updated vector dim access to enable vector load

* updated test/profiler files

* updated example with 1d kernel

* updating profiler

* renamed files

* disabled device op for MI300

* skip  elementwise_permute_2d on gfx94x

* Update CMakeLists.txt

* fixing CMake - disabling some GPU targets

---------
Co-authored-by: Jing Zhang <jizha@amd.com>
Co-authored-by: Jing Zhang <jizhan@amd.com>
Co-authored-by: zjing14 <zhangjing14@gmail.com>

a2969aa8

28 Nov, 2023 3 commits

recover default niter (#1064) · ae5e5181
zjing14 authored Nov 28, 2023

ae5e5181

Split the static library into several files. (#1044) · 7965d66a

Illia Silin authored Nov 28, 2023

* spolit the static library into several

* update lib paths and fix client example

* do not use device_mha_operarions for client examples

* use appropriate libs to link to client examples

* remove the gpu/transpose path from the list

* try fixing clinet examples 3,4,9

* add necessary libs for client examples

* fix the layernorm client example

* fix the client examples 23 and 24

* fix typo

* add interface library and refresh clang format

7965d66a

Switch default f8 conversion to stochastic rounding (#1048) · 6ef034f6

Rostyslav Geyyer authored Nov 27, 2023

* Switch default f8 conversion to stochastic rounding

* Refactor f8-related type_converts

* Add an element-wise op

6ef034f6

27 Nov, 2023 2 commits
- Add missing check for K padding in XDL GEMM (#1056) · 60ecfd73
  Bartlomiej Wroblewski authored Nov 27, 2023
  
  60ecfd73
- Fix cluster length arrange order in fp16 GEMM example (#1055) · bfecc193
  Bartlomiej Wroblewski authored Nov 27, 2023
  
  bfecc193
25 Nov, 2023 1 commit

Add basic support for direct loads from global to LDS (#999) · 627054b9

Bartlomiej Wroblewski authored Nov 25, 2023

* Add basic support for direct loads from global to LDS

* Clean the code and comments

* Add support for fp16

* Add comments

* Add check for thread cluster lengths

* Align non-direct-load fp16 example

* Small fixes

* Extend IsSupported to check for supported GPU gens

* Build examples only on the supported HW

* Do not throw when instance not supported in 04 example

* Review: Apply review suggestions

* Review: small fix

* Review: small fix

627054b9

21 Nov, 2023 1 commit
- Merge with (not the latest) upstream CK (#32) · 0a7174ad
  Chao Liu authored Nov 21, 2023
```
* fix build for old ck examples

* fix build for old ck
```
  0a7174ad
17 Nov, 2023 1 commit

Improve 4k gemm perf (#1047) · e8cddfdc

zjing14 authored Nov 17, 2023



* improve 4k gemm perf

* add f8 instances

* format

---------
Co-authored-by: Jing Zhang <jizha@amd.com>

e8cddfdc

16 Nov, 2023 2 commits

[Hotfix] Remove unsed profile_transpose.cpp (#1046) · e1fa0091
Chao Liu authored Nov 16, 2023

e1fa0091

Bump rocm-docs-core from 0.26.0 to 0.27.0 in /docs/sphinx (#1023) · 61cce232

dependabot[bot] authored Nov 15, 2023

Bumps [rocm-docs-core](https://github.com/RadeonOpenCompute/rocm-docs-core) from 0.26.0 to 0.27.0.
- [Release notes](https://github.com/RadeonOpenCompute/rocm-docs-core/releases)
- [Changelog](https://github.com/RadeonOpenCompute/rocm-docs-core/blob/develop/CHANGELOG.md)
- [Commits](https://github.com/RadeonOpenCompute/rocm-docs-core/compare/v0.26.0...v0.27.0

)

---
updated-dependencies:
- dependency-name: rocm-docs-core
  dependency-type: direct:production
  update-type: version-update:semver-minor
...
Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>

61cce232

15 Nov, 2023 4 commits

increase warm-up to 10 iter (#28) · 496be40e
Chao Liu authored Nov 15, 2023

496be40e

Fmha pr 2 (#26) · 3753c4bc

carlushuang authored Nov 16, 2023

* support hdim=64/128 in same example code

* support v transpose

* revert gemm.cpp, not intent to modify it

* remove useless code

* fix a bug for swizzle C encoding, no perf change

* optimize LDS encoding

* update LDS layout

* clean up code

3753c4bc

Log CDEBlockTransferScalarPerVector_NPerBlock in conv fwd multiD xdl (#1042) · 1fefd82e

Bartłomiej Kocot authored Nov 15, 2023

* Log CDEBlockTransferScalarPerVector_NPerBlock in conv_fwd_multi_d_xdl implementation

* Log CDEBlockTransferScalarPerVector_NPerBlock in conv fwd multiD xdl

1fefd82e

Fix check for conv Fwd Filter1x1Pad0 (#1040) · 3ef3102f
Bartłomiej Kocot authored Nov 15, 2023
```
* Fix check for conv Fwd Filter1x1Pad0

* Fix check for conv Fwd Filter1x1Pad0
```
3ef3102f

14 Nov, 2023 1 commit

Introduce multiABD api and deprecate multiD (#1035) · f2398f61

Bartłomiej Kocot authored Nov 14, 2023

* Introduce multiABD api and deprecate multiD

* Replace multiD with multiABD

* Mark structures as deprecated

* Change doxygen deprecated to note to avoid warnings

f2398f61

13 Nov, 2023 2 commits

Add conv bwd weight client example (#1005) · 5356c4a9

Rostyslav Geyyer authored Nov 13, 2023

* Add conv bwd weight client example

* Update instance selector

* Fake the conversion

* Bring the conversion back

5356c4a9

Hip tensor permute (#1002) · 454cf7bd

arai713 authored Nov 13, 2023

* adding files for F32 example

* adding functioning implementation with scalar multiplication and unary operator support

* added fp 16 type check in unary square

* updating scalar multiplication as an operator

* functioning version with scalar operator

* changing strides for col major

* updated column major implementation

* working column major implementation

* cleaned up comments, rearranged/renamed files

454cf7bd

11 Nov, 2023 1 commit

add more instances for bfp16 gemm (#1036) · 600fc000

zjing14 authored Nov 11, 2023



* add more instances for bfp16

* reduce the gemm input values to prevent round-off errors

---------
Co-authored-by: Jing Zhang <jizha@amd.com>
Co-authored-by: illsilin <Illia.Silin@amd.com>

600fc000

10 Nov, 2023 2 commits

Support multi AB for grouped conv fwd xdl (#1027) · 49e52bb3

Bartłomiej Kocot authored Nov 10, 2023

* Support multi AB for grouped conv fwd xdl

* Add instances

* Add client example

* Add example

* Add interface test

* Minor fixes

Minor fixes

Minor fixes

* Comment fixes

* Fixes

* Reference fix

* Test xdl fixes

* Improve multi_ab interface test

49e52bb3

Backward of gamma and beta for layernorm and groupnorm (#1013) · 1db75603

rocking authored Nov 10, 2023

* Add layernorm backward reference code

* Add groupnorm backward reference code

* Add example

* clang format

* Fixc bug of reference layernorm and groupnorm

* Fix naming

* Refine naming

* Add device op for normalization bwd gamma and beta

* Refine template parameter

* Add bwd gamma & beta of kernel

* 1. Add groupnorm example
2. Refine layernorm naming

* Narrow down the static check for performance

* Refine variable name

1db75603

09 Nov, 2023 3 commits

add linker script to QA builds (#1030) · 68f2b5e7
Illia Silin authored Nov 08, 2023

68f2b5e7

Transpose 3d (#984) · 3af8c81a

arai713 authored Nov 08, 2023



* added working example for 5D input using 1D kernel

* example with 5D input tensor and 2d kernel - not working: issues with arguments

* added updated version of 3d device op - changed descriptors/dims

* added example file to check kernel

* fixed descriptor and isSupportedArgument stride problem

* added and modified kernel for 3d - updated tids/loop

* adding some more 5d example files

* fixed some issues

* changes made for testing

* working version: fixed error in stride for A, still a bit inefficient

* cleaned up formatting/comments

* updating formatting

* more formatting fixes

* fixing cmake, adding back gpu targets in cmake script

* adding client example

* added instances for client example

* fixed errors in client example

* implemented client ex with device_elementwise.hpp and device_elementwise_3d_impl.hpp

* removed extra files

* minor formatting and naming fixes

* adding test files and profiler

* fixing minor error

* minor fix

* removed unneccesary comments, renamed files

* updated instance list for client example, added different layout example

* removing instances

* fixed error in instance generation

* remove comments

* update profiler and client example tensor layouts

* fixed errors in test/profiler

* updated vector dim access to enable vector load

* updated test/profiler files

* updated example with 1d kernel

* updating profiler

* renamed files

---------
Co-authored-by: Jing Zhang <jizha@amd.com>

3af8c81a

Layernorm4d (#1022) · a3d9a2cd

rocking authored Nov 09, 2023



* Rename folder

* Add layernorm 4d fwd example

* Rename original layernorm example

* Add layernorm 4d f16  test

* Add layernorm4d_fwd client example

* Support layernorm4D in ckProfiler

* Rename groupnorm to groupnorm fwd in example

* Rename layernorm and group fwd in test

* Rename normalization to normalization_fwd (instances)

* Add fwd to DeviceNormalization

* Rename external api header

* Rename folder, because we can also add bwd in this folder

* Add fwd in layernorm and groupnorm (profiler

* Fix compile error

---------
Co-authored-by: Po Yen Chen <PoYen.Chen@amd.com>

a3d9a2cd

08 Nov, 2023 1 commit
- Support fp64 contraction on gfx94x. (#1029) · ce526211
  Illia Silin authored Nov 08, 2023
```
* enable contraction fp64 on gfx94*

* fix the logic
```
  ce526211
07 Nov, 2023 2 commits

Add Gemm instances for performance improvement (#1018) · 98fd41f5

zjing14 authored Nov 07, 2023



* improve kpad

* more tuning parameters

* f16_f8_fp16

* cut test time

* add f16_f8_fp16

* add f16_f8_f16

* testing instances for skinny cases

* format

* clean

* add fp16_f8_fp16

* clang-format

* add grouped gemm instalces

* fixed profile grouped_gemm

* clean

* clean

* clean

* clean

* clean

* add missing instance func

* fixed inferface

---------
Co-authored-by: Jing Zhang <jizha@amd.com>
Co-authored-by: root <root@sh5-1e707-rc06-38.mkm.dcgpu>

98fd41f5

Add compute type check for convolution instances (#1015) · aa0b9798

Daming Feng authored Nov 06, 2023



* add compute type check for fp16 in forward convolution instances

* Add compute type check for default compute types

---------
Co-authored-by: Bartlomiej Kocot <barkocot@amd.com>

aa0b9798

03 Nov, 2023 2 commits
- switch the hipTensor testing from mainline to develop branch (#1025) · b0568b72
  Illia Silin authored Nov 03, 2023
  
  b0568b72
- Add missing ComputeDatatype in contraction_multi_ABD_xdl_fp16 (#1024) · 16eb824c
  Bartlomiej Wroblewski authored Nov 03, 2023
  
  16eb824c