Commits · bc367a779cfef94ba243fe809a4fe21fea9bcd89 · gaoqiong / composable_kernel

04 Dec, 2023 3 commits
- Merge branch 'develop' into barkocot/lwpck-1063-dev · bc367a77
  Bartłomiej Kocot authored Dec 04, 2023
  
  bc367a77
- Fix comments · 0dc0af2e
  Bartlomiej Kocot authored Dec 04, 2023
  
  0dc0af2e
- Revert "Update cmake files" · 53fdf365
  Bartlomiej Kocot authored Dec 04, 2023
```
This reverts commit c27f88b5.
```
  53fdf365
03 Dec, 2023 1 commit

Add support for double buffering in direct load GEMM kernel (#1052) · bc4bf9bd

Bartlomiej Wroblewski authored Dec 03, 2023

This PR introduces support for double buffering in LDS into GEMM kernels that use direct load instructions.

Direct loads now use inline asm instead of intrinsics. Usage of intrinsics results in compiler adding additional waitcnt instructions what breaks possible load/compute overlap in case of double buffering.

Usage of inline asm results in the need to use sched_barrier in order to make sure that compiler cannot incorrectly reschedule instructions since it does not know the data dependencies between global->LDS and LDS->registers.

bc4bf9bd

01 Dec, 2023 2 commits
- Update cmake files · c27f88b5
  Bartlomiej Kocot authored Dec 01, 2023
  
  c27f88b5
- Merge branch 'develop' into barkocot/lwpck-1063-dev · f741895f
  Bartłomiej Kocot authored Dec 01, 2023
  
  f741895f
30 Nov, 2023 4 commits
- [CI] Update Jenkinsfile (#1073) · c7d5c772
  Jun Liu authored Nov 30, 2023
  
  c7d5c772
- Fixed GroupedGemmFixedNK with hipGraph (#1065) · 49df1dc5
  zjing14 authored Nov 30, 2023
```
* fixed examples; add async_mem_set

* add stream to all deviceOp using SetWorkspace

---------
Co-authored-by: Jing Zhang <jizha@amd.com>
```
  49df1dc5
- Introduce wrapper library · e5f05e71
  Bartlomiej Kocot authored Nov 30, 2023
  
  e5f05e71
- Introduce wrapper for layout (#1054) · 8ff845f2
  Bartłomiej Kocot authored Nov 30, 2023
```
* Introduce wrapper for layout

* Extend functionality

* Fix for getLength

* Comment fixes

* Add comments and remove not needed getters
```
  8ff845f2
29 Nov, 2023 1 commit

Disable transpose device op for MI300 (#1050) · a2969aa8

arai713 authored Nov 29, 2023



* added working example for 5D input using 1D kernel

* example with 5D input tensor and 2d kernel - not working: issues with arguments

* added updated version of 3d device op - changed descriptors/dims

* added example file to check kernel

* fixed descriptor and isSupportedArgument stride problem

* added and modified kernel for 3d - updated tids/loop

* adding some more 5d example files

* fixed some issues

* changes made for testing

* working version: fixed error in stride for A, still a bit inefficient

* cleaned up formatting/comments

* updating formatting

* more formatting fixes

* fixing cmake, adding back gpu targets in cmake script

* adding client example

* added instances for client example

* fixed errors in client example

* implemented client ex with device_elementwise.hpp and device_elementwise_3d_impl.hpp

* removed extra files

* minor formatting and naming fixes

* adding test files and profiler

* fixing minor error

* minor fix

* removed unneccesary comments, renamed files

* updated instance list for client example, added different layout example

* removing instances

* fixed error in instance generation

* remove comments

* update profiler and client example tensor layouts

* fixed errors in test/profiler

* updated vector dim access to enable vector load

* updated test/profiler files

* updated example with 1d kernel

* updating profiler

* renamed files

* disabled device op for MI300

* skip  elementwise_permute_2d on gfx94x

* Update CMakeLists.txt

* fixing CMake - disabling some GPU targets

---------
Co-authored-by: Jing Zhang <jizha@amd.com>
Co-authored-by: Jing Zhang <jizhan@amd.com>
Co-authored-by: zjing14 <zhangjing14@gmail.com>

a2969aa8

28 Nov, 2023 3 commits

recover default niter (#1064) · ae5e5181
zjing14 authored Nov 28, 2023

ae5e5181

Split the static library into several files. (#1044) · 7965d66a

Illia Silin authored Nov 28, 2023

* spolit the static library into several

* update lib paths and fix client example

* do not use device_mha_operarions for client examples

* use appropriate libs to link to client examples

* remove the gpu/transpose path from the list

* try fixing clinet examples 3,4,9

* add necessary libs for client examples

* fix the layernorm client example

* fix the client examples 23 and 24

* fix typo

* add interface library and refresh clang format

7965d66a

Switch default f8 conversion to stochastic rounding (#1048) · 6ef034f6

Rostyslav Geyyer authored Nov 27, 2023

* Switch default f8 conversion to stochastic rounding

* Refactor f8-related type_converts

* Add an element-wise op

6ef034f6

27 Nov, 2023 2 commits
- Add missing check for K padding in XDL GEMM (#1056) · 60ecfd73
  Bartlomiej Wroblewski authored Nov 27, 2023
  
  60ecfd73
- Fix cluster length arrange order in fp16 GEMM example (#1055) · bfecc193
  Bartlomiej Wroblewski authored Nov 27, 2023
  
  bfecc193
25 Nov, 2023 1 commit

Add basic support for direct loads from global to LDS (#999) · 627054b9

Bartlomiej Wroblewski authored Nov 25, 2023

* Add basic support for direct loads from global to LDS

* Clean the code and comments

* Add support for fp16

* Add comments

* Add check for thread cluster lengths

* Align non-direct-load fp16 example

* Small fixes

* Extend IsSupported to check for supported GPU gens

* Build examples only on the supported HW

* Do not throw when instance not supported in 04 example

* Review: Apply review suggestions

* Review: small fix

* Review: small fix

627054b9

17 Nov, 2023 1 commit

Improve 4k gemm perf (#1047) · e8cddfdc

zjing14 authored Nov 17, 2023



* improve 4k gemm perf

* add f8 instances

* format

---------
Co-authored-by: Jing Zhang <jizha@amd.com>

e8cddfdc

16 Nov, 2023 2 commits

[Hotfix] Remove unsed profile_transpose.cpp (#1046) · e1fa0091
Chao Liu authored Nov 16, 2023

e1fa0091

Bump rocm-docs-core from 0.26.0 to 0.27.0 in /docs/sphinx (#1023) · 61cce232

dependabot[bot] authored Nov 15, 2023

Bumps [rocm-docs-core](https://github.com/RadeonOpenCompute/rocm-docs-core) from 0.26.0 to 0.27.0.
- [Release notes](https://github.com/RadeonOpenCompute/rocm-docs-core/releases)
- [Changelog](https://github.com/RadeonOpenCompute/rocm-docs-core/blob/develop/CHANGELOG.md)
- [Commits](https://github.com/RadeonOpenCompute/rocm-docs-core/compare/v0.26.0...v0.27.0

)

---
updated-dependencies:
- dependency-name: rocm-docs-core
  dependency-type: direct:production
  update-type: version-update:semver-minor
...
Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>

61cce232

15 Nov, 2023 2 commits

Log CDEBlockTransferScalarPerVector_NPerBlock in conv fwd multiD xdl (#1042) · 1fefd82e

Bartłomiej Kocot authored Nov 15, 2023

* Log CDEBlockTransferScalarPerVector_NPerBlock in conv_fwd_multi_d_xdl implementation

* Log CDEBlockTransferScalarPerVector_NPerBlock in conv fwd multiD xdl

1fefd82e

Fix check for conv Fwd Filter1x1Pad0 (#1040) · 3ef3102f
Bartłomiej Kocot authored Nov 15, 2023
```
* Fix check for conv Fwd Filter1x1Pad0

* Fix check for conv Fwd Filter1x1Pad0
```
3ef3102f

14 Nov, 2023 1 commit

Introduce multiABD api and deprecate multiD (#1035) · f2398f61

Bartłomiej Kocot authored Nov 14, 2023

* Introduce multiABD api and deprecate multiD

* Replace multiD with multiABD

* Mark structures as deprecated

* Change doxygen deprecated to note to avoid warnings

f2398f61

13 Nov, 2023 2 commits

Add conv bwd weight client example (#1005) · 5356c4a9

Rostyslav Geyyer authored Nov 13, 2023

* Add conv bwd weight client example

* Update instance selector

* Fake the conversion

* Bring the conversion back

5356c4a9

Hip tensor permute (#1002) · 454cf7bd

arai713 authored Nov 13, 2023

* adding files for F32 example

* adding functioning implementation with scalar multiplication and unary operator support

* added fp 16 type check in unary square

* updating scalar multiplication as an operator

* functioning version with scalar operator

* changing strides for col major

* updated column major implementation

* working column major implementation

* cleaned up comments, rearranged/renamed files

454cf7bd

11 Nov, 2023 1 commit

add more instances for bfp16 gemm (#1036) · 600fc000

zjing14 authored Nov 11, 2023



* add more instances for bfp16

* reduce the gemm input values to prevent round-off errors

---------
Co-authored-by: Jing Zhang <jizha@amd.com>
Co-authored-by: illsilin <Illia.Silin@amd.com>

600fc000

10 Nov, 2023 2 commits

Support multi AB for grouped conv fwd xdl (#1027) · 49e52bb3

Bartłomiej Kocot authored Nov 10, 2023

* Support multi AB for grouped conv fwd xdl

* Add instances

* Add client example

* Add example

* Add interface test

* Minor fixes

Minor fixes

Minor fixes

* Comment fixes

* Fixes

* Reference fix

* Test xdl fixes

* Improve multi_ab interface test

49e52bb3

Backward of gamma and beta for layernorm and groupnorm (#1013) · 1db75603

rocking authored Nov 10, 2023

* Add layernorm backward reference code

* Add groupnorm backward reference code

* Add example

* clang format

* Fixc bug of reference layernorm and groupnorm

* Fix naming

* Refine naming

* Add device op for normalization bwd gamma and beta

* Refine template parameter

* Add bwd gamma & beta of kernel

* 1. Add groupnorm example
2. Refine layernorm naming

* Narrow down the static check for performance

* Refine variable name

1db75603

09 Nov, 2023 3 commits

add linker script to QA builds (#1030) · 68f2b5e7
Illia Silin authored Nov 08, 2023

68f2b5e7

Transpose 3d (#984) · 3af8c81a

arai713 authored Nov 08, 2023



* added working example for 5D input using 1D kernel

* example with 5D input tensor and 2d kernel - not working: issues with arguments

* added updated version of 3d device op - changed descriptors/dims

* added example file to check kernel

* fixed descriptor and isSupportedArgument stride problem

* added and modified kernel for 3d - updated tids/loop

* adding some more 5d example files

* fixed some issues

* changes made for testing

* working version: fixed error in stride for A, still a bit inefficient

* cleaned up formatting/comments

* updating formatting

* more formatting fixes

* fixing cmake, adding back gpu targets in cmake script

* adding client example

* added instances for client example

* fixed errors in client example

* implemented client ex with device_elementwise.hpp and device_elementwise_3d_impl.hpp

* removed extra files

* minor formatting and naming fixes

* adding test files and profiler

* fixing minor error

* minor fix

* removed unneccesary comments, renamed files

* updated instance list for client example, added different layout example

* removing instances

* fixed error in instance generation

* remove comments

* update profiler and client example tensor layouts

* fixed errors in test/profiler

* updated vector dim access to enable vector load

* updated test/profiler files

* updated example with 1d kernel

* updating profiler

* renamed files

---------
Co-authored-by: Jing Zhang <jizha@amd.com>

3af8c81a

Layernorm4d (#1022) · a3d9a2cd

rocking authored Nov 09, 2023



* Rename folder

* Add layernorm 4d fwd example

* Rename original layernorm example

* Add layernorm 4d f16  test

* Add layernorm4d_fwd client example

* Support layernorm4D in ckProfiler

* Rename groupnorm to groupnorm fwd in example

* Rename layernorm and group fwd in test

* Rename normalization to normalization_fwd (instances)

* Add fwd to DeviceNormalization

* Rename external api header

* Rename folder, because we can also add bwd in this folder

* Add fwd in layernorm and groupnorm (profiler

* Fix compile error

---------
Co-authored-by: Po Yen Chen <PoYen.Chen@amd.com>

a3d9a2cd

08 Nov, 2023 1 commit
- Support fp64 contraction on gfx94x. (#1029) · ce526211
  Illia Silin authored Nov 08, 2023
```
* enable contraction fp64 on gfx94*

* fix the logic
```
  ce526211
07 Nov, 2023 2 commits

Add Gemm instances for performance improvement (#1018) · 98fd41f5

zjing14 authored Nov 07, 2023



* improve kpad

* more tuning parameters

* f16_f8_fp16

* cut test time

* add f16_f8_fp16

* add f16_f8_f16

* testing instances for skinny cases

* format

* clean

* add fp16_f8_fp16

* clang-format

* add grouped gemm instalces

* fixed profile grouped_gemm

* clean

* clean

* clean

* clean

* clean

* add missing instance func

* fixed inferface

---------
Co-authored-by: Jing Zhang <jizha@amd.com>
Co-authored-by: root <root@sh5-1e707-rc06-38.mkm.dcgpu>

98fd41f5

Add compute type check for convolution instances (#1015) · aa0b9798

Daming Feng authored Nov 06, 2023



* add compute type check for fp16 in forward convolution instances

* Add compute type check for default compute types

---------
Co-authored-by: Bartlomiej Kocot <barkocot@amd.com>

aa0b9798

03 Nov, 2023 2 commits
- switch the hipTensor testing from mainline to develop branch (#1025) · b0568b72
  Illia Silin authored Nov 03, 2023
  
  b0568b72
- Add missing ComputeDatatype in contraction_multi_ABD_xdl_fp16 (#1024) · 16eb824c
  Bartlomiej Wroblewski authored Nov 03, 2023
  
  16eb824c
02 Nov, 2023 2 commits

Add support for mixed precision in contraction scale and bilinear (#973) · 4ef704d8

Bartlomiej Wroblewski authored Nov 02, 2023



* Add support for mixed precision in contraction scale and bilinear (#936)

* Extract common functionality to separate files

* Reference contraction: Remove incorrect consts from type_converts

* Reference contraction: Add missing type_convert for dst value

* Reference contraction: Fix incorrect order of B matrix dimensions

* Add support for mixed precision in contraction scale and bilinear

* Move using statements from instances to a common file

* Move using statements from examples to a common file

* Fix the order of B matrix dimensions across examples and profiler

* Fix the computation of error threshold

* Make ComputeDataType an optional argument

* Include possible DataType -> ComputeDataType casting error in the threshold

* Remove commented code

* Make the ComputeDataType an optional argument in instance

---------
Co-authored-by: Illia Silin <98187287+illsilin@users.noreply.github.com>

4ef704d8

Bump rocm-docs-core from 0.24.0 to 0.26.0 in /docs/sphinx (#987) · 73743aa0

dependabot[bot] authored Nov 02, 2023

Bumps [rocm-docs-core](https://github.com/RadeonOpenCompute/rocm-docs-core) from 0.24.0 to 0.26.0.
- [Release notes](https://github.com/RadeonOpenCompute/rocm-docs-core/releases)
- [Changelog](https://github.com/RadeonOpenCompute/rocm-docs-core/blob/develop/CHANGELOG.md)
- [Commits](https://github.com/RadeonOpenCompute/rocm-docs-core/compare/v0.24.0...v0.26.0

)

---
updated-dependencies:
- dependency-name: rocm-docs-core
  dependency-type: direct:production
  update-type: version-update:semver-minor
...
Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>

73743aa0

01 Nov, 2023 2 commits
- Add ScaleAddScaleAddRelu post op for conv fwd (#1006) · f27ea94e
  Bartłomiej Kocot authored Nov 02, 2023
```
* Add ScaleAddScaleAddRelu post op for conv fwd

* Fixes

* Fix instance file name

* Minor fix
```
  f27ea94e
- handle the exception when cannot connect to redis server (#1019) · 306fd506
  Illia Silin authored Nov 01, 2023
  
  306fd506