Commits · 522b7aeedef4e5fa578e4505f793e5528022efff · gaoqiong / composable_kernel_ROCM

29 Jan, 2024 2 commits
- Clean up debug code and reuse new neighbour count func. · 1b462ab5
  Adam Osewski authored Jan 29, 2024
  
  1b462ab5
- Clean up and change how neighbours are counted. · e954c206
  Adam Osewski authored Jan 29, 2024
  
  e954c206
24 Jan, 2024 1 commit

Fixing most of the cppcheck errors. (#1142) · 180e5720

Illia Silin authored Jan 24, 2024

* fix cppcheck errors, first pass

* fix format

* fix returned value in examples

* add macro definitions for cppcheck

* fix the profile_gemm logic

* update the gemm profiler logic

* add more difinitions to cppcheck, fix couple more errors

* replace runtime error with message in device function

* fix a couple of int4 issues

* no return for fill function

* fix errors in data_types.hpp

* fix format

* fix few remaining errors

* fix errors in data_types.hpp

* fix last couple of errors in datat_types.hpp

180e5720

23 Jan, 2024 1 commit
- Commit debug WIP for sharing. · 7e71ea99
  Adam Osewski authored Jan 23, 2024
  
  7e71ea99
19 Jan, 2024 2 commits

[GEMM] Optimization for MI200/300. (#1135) · bb63b973

Haocong WANG authored Jan 19, 2024

* Optimize GEMM on MI200/300:
1. Add new blockwise gemm pipeline
2. Add irregular splitk intances

* clang format + typo fix

* Fix a bug

bb63b973

Add optimized copy to ck wrapper (#1126) · 7e4eb4b8

Bartłomiej Kocot authored Jan 19, 2024



* Add optimized copy to ck wrapper

* Example optimizations

* Fixes

* Move img2col test to client example

* Refactor example

* Fix docs

* Fixes

* Fix

* Fixes

* Fixes

* Fixes

* Fixes

* Fixes

---------
Co-authored-by: zjing14 <zhangjing14@gmail.com>

7e4eb4b8

15 Jan, 2024 1 commit

Add cppcheck to CK CI. (#1125) · e6d099c8

Illia Silin authored Jan 15, 2024

* add cppcheck to the CK CI

* fix the path to CK source for cppcheck

* fix the path to CK source for cppcheck one more time

* fix the path to CK source for cppcheck third time

* change the path to ck_cppcheck.log

* install latest cppcheck from source

* fix bug in ck.hpp and use 20 threads for cppcheck

* create a switch to turn cppckeck on and off in CI

e6d099c8

11 Jan, 2024 2 commits
- Add MK-KN FP16 instances. · 734df790
  root authored Jan 11, 2024
  
  734df790
- Do not clear cthread buffer if needed. · f9f2cdf9
  root authored Jan 11, 2024
```
- Add output stream operators for LoopSched and PiplineVer
```
  f9f2cdf9
10 Jan, 2024 1 commit
- Hide unused tparams from device op and copy kernel args directly when setting pointer · 66c70dfe
  Adam Osewski authored Jan 10, 2024
  
  66c70dfe
09 Jan, 2024 2 commits

Add an option to change the number of warm-up cycles and iterations. (#1124) · 886d9eeb
Illia Silin authored Jan 09, 2024
```
* allow setting the number of warmup cycles and iterations for profiler

* fix the gemm_splitk and grouped_gemm examples
```
886d9eeb

SWDEV-439954 - Use hard coded filename rather than using the macro __FILE__... · e699dbd8

raramakr authored Jan 09, 2024


SWDEV-439954 - Use hard coded filename rather than using the macro __FILE__ for debug prints. (#1123)

* SWDEV-439954 - Use hard coded filename rather than using the macro __FILE__ for debug prints.

Hiptensor library is using the header files from CK. Hard coded ROCm path was getting embedded into the hiptensor library, since the header file was having the macro __FILE__. Replace the macro with filename.

* fix syntax

---------
Co-authored-by: illsilin <Illia.Silin@amd.com>

e699dbd8

03 Jan, 2024 1 commit

Add tensor partition and generic copy for ck wrapper (#1108) · 4234b3a6

Bartłomiej Kocot authored Jan 03, 2024

* Add tensor partition and generic copy for ck wrapper

* Update changelog

* Stylistic fixes

* Change shape/strides logic to descriptor transforms

* Fixes

* Fix client example

* Fix comments

4234b3a6

20 Dec, 2023 2 commits

enable compilation of INSTANCES_ONLY for Windows (#1082) · fb5bd51b

Artur Wojcik authored Dec 20, 2023



* enable compilation of INSTANCES_ONLY for Windows

* suppress ROCMChecks warnings on GoogleTests

* suppress -Wfloat-equal warning on GoogleTests

---------
Co-authored-by: Illia Silin <98187287+illsilin@users.noreply.github.com>

fb5bd51b

Fix compiler errors. · 1a1fd0b3
Adam Osewski authored Dec 20, 2023

1a1fd0b3

19 Dec, 2023 1 commit
- Fix StorePartials. · 134fc2e7
  Adam Osewski authored Dec 19, 2023
```
Pass pointer to whole workspace not the shifted one.
```
  134fc2e7
18 Dec, 2023 1 commit

layernorm and groupnorm backward data (#1083) · a69aa2a1

rocking authored Dec 19, 2023

* rename folder

* Add type string

* Remove typo

* Add deviceOp to backward x

* Add comment to describe the behavior of backward normalization

* Add kernel function, prepare to implement

* implement generic kernel

* Check vector size

* Add sweep once pipeline for small reduce size

* Fix bug of KRaw_ error

* Fix bug of dx stride

* sanity check for mean and rstd

* backward x for groupnorm

* Add bwd x instance

* add layernorm 2d bwd gamma beta instances

* Change save mean var type from f32 to f16 in f16 mode

* Change the example to f16

* Add groupnorm bwd gamma beta instance

* Add groupnorm bwd x instance

* Fix naming

* Add layernorm bwd x ckprofiler

* Add groupnorm bwd x profiler

* clang format

* Rename bwd x to bwd data

* Fix bug of verification in profiler

* Add test of layernorm and groupnorm bwd data

* Add missing cmake

* Add layernorm2d bwd data

* rename fwd example

* Add groupnorm client example

* Fix typo. replace Invarient with Invariant

* Add checking before running the best instance

a69aa2a1

15 Dec, 2023 1 commit
- Add tensor structure to wrapper (#1098) · 07092d68
  Bartłomiej Kocot authored Dec 15, 2023
```
* Add tensor structure to wrapper

* update changelog

* Fix names

* Comment fixes
```
  07092d68
13 Dec, 2023 2 commits
- [Doc][Werror] Fix security alerts and sync with MIOpen (#1085) · 3a3b98ef
  Jun Liu authored Dec 13, 2023
```
* fix Werror unused-parameter

* sync doc requirements

* fix blank space format

* fix dependency issue
```
  3a3b98ef
- Fix the bugs (#1099) · 6891e4d1
  Rostyslav Geyyer authored Dec 13, 2023
  
  6891e4d1
11 Dec, 2023 1 commit

Fix IsSupported check in the contraction op (#1066) · 89ee4746

Bartlomiej Wroblewski authored Dec 11, 2023

Current implementation of IsSupported method in contraction ops does not cover a lot of possible cases in which ScalarPerVector cannot really be used to read A, B or D, or write E.

This PR extends both the regular and multiABD contraction ops with improved checks and also adds new instances with smaller values of ScalarPerVector to support instances that are not supported by other instances.

89ee4746

08 Dec, 2023 1 commit
- Support broadcast for bias in grouped conv fwd (#1081) · f8369848
  Bartłomiej Kocot authored Dec 08, 2023
```
* Support broadcast for bias in grouped conv fwd

* Fix comment

* Comment fixes

* Remove GK layout
```
  f8369848
07 Dec, 2023 1 commit

Switch from ROCmSoftwarePlatform to ROCm org (#1091) · d939411d

Illia Silin authored Dec 07, 2023

* switch from ROCmSoftwarePlatform to ROCm org

* replace ROCmSoftwarePlatform with ROCm in few more places

d939411d

06 Dec, 2023 2 commits

Few fixes. · 88a4fbfb

Adam Osewski authored Dec 06, 2023

* prevent clearing c_thread_buffer between consecutive k-dim data tiles GEMM.
* Limit the number of launched thread blocks.

88a4fbfb

Introduce wrapper library (#1071) · 836b7e55

Bartłomiej Kocot authored Dec 06, 2023

* Introduce wrapper library

* Update cmake files

* Revert "Update cmake files"

This reverts commit c27f88b56590c11a88e26d5d0df7aca51a08133d.

* Fix comments

836b7e55

03 Dec, 2023 1 commit

Add support for double buffering in direct load GEMM kernel (#1052) · bc4bf9bd

Bartlomiej Wroblewski authored Dec 03, 2023

This PR introduces support for double buffering in LDS into GEMM kernels that use direct load instructions.

Direct loads now use inline asm instead of intrinsics. Usage of intrinsics results in compiler adding additional waitcnt instructions what breaks possible load/compute overlap in case of double buffering.

Usage of inline asm results in the need to use sched_barrier in order to make sure that compiler cannot incorrectly reschedule instructions since it does not know the data dependencies between global->LDS and LDS->registers.

bc4bf9bd

01 Dec, 2023 3 commits

Multiple changes to gridwise gemm. · 4841d991
Adam Osewski authored Dec 01, 2023

4841d991

Multiple changes to global kernel function. · ad0e4083

Adam Osewski authored Dec 01, 2023

* StorePartials work on offseted pointer.
* Read flags as uint32_t value
* Accumulate partials only if there is more than one cooperating workgroup
* Add condition for waiting on reduction end, only when there is still work to do.
* Fix creation od a/b grid desc in CheckArgument.
* LaunchKernel will use preprocess lambda to set flags value to zero.
* Add condition in IsSupportedArgument to check if xdl is supported.

ad0e4083

Change return type from inxed_t to uint32_t for GetFlagValue. · f8cbbd1b
Adam Osewski authored Dec 01, 2023
```
Update doc.
```
f8cbbd1b

30 Nov, 2023 2 commits

Fixed GroupedGemmFixedNK with hipGraph (#1065) · 49df1dc5

zjing14 authored Nov 30, 2023



* fixed examples; add async_mem_set

* add stream to all deviceOp using SetWorkspace

---------
Co-authored-by: Jing Zhang <jizha@amd.com>

49df1dc5

Introduce wrapper for layout (#1054) · 8ff845f2

Bartłomiej Kocot authored Nov 30, 2023

* Introduce wrapper for layout

* Extend functionality

* Fix for getLength

* Comment fixes

* Add comments and remove not needed getters

8ff845f2

29 Nov, 2023 1 commit

Disable transpose device op for MI300 (#1050) · a2969aa8

arai713 authored Nov 29, 2023



* added working example for 5D input using 1D kernel

* example with 5D input tensor and 2d kernel - not working: issues with arguments

* added updated version of 3d device op - changed descriptors/dims

* added example file to check kernel

* fixed descriptor and isSupportedArgument stride problem

* added and modified kernel for 3d - updated tids/loop

* adding some more 5d example files

* fixed some issues

* changes made for testing

* working version: fixed error in stride for A, still a bit inefficient

* cleaned up formatting/comments

* updating formatting

* more formatting fixes

* fixing cmake, adding back gpu targets in cmake script

* adding client example

* added instances for client example

* fixed errors in client example

* implemented client ex with device_elementwise.hpp and device_elementwise_3d_impl.hpp

* removed extra files

* minor formatting and naming fixes

* adding test files and profiler

* fixing minor error

* minor fix

* removed unneccesary comments, renamed files

* updated instance list for client example, added different layout example

* removing instances

* fixed error in instance generation

* remove comments

* update profiler and client example tensor layouts

* fixed errors in test/profiler

* updated vector dim access to enable vector load

* updated test/profiler files

* updated example with 1d kernel

* updating profiler

* renamed files

* disabled device op for MI300

* skip  elementwise_permute_2d on gfx94x

* Update CMakeLists.txt

* fixing CMake - disabling some GPU targets

---------
Co-authored-by: Jing Zhang <jizha@amd.com>
Co-authored-by: Jing Zhang <jizhan@amd.com>
Co-authored-by: zjing14 <zhangjing14@gmail.com>

a2969aa8

28 Nov, 2023 2 commits
- recover default niter (#1064) · ae5e5181
  zjing14 authored Nov 28, 2023
  
  ae5e5181
- Switch default f8 conversion to stochastic rounding (#1048) · 6ef034f6
  Rostyslav Geyyer authored Nov 27, 2023
```
* Switch default f8 conversion to stochastic rounding

* Refactor f8-related type_converts

* Add an element-wise op
```
  6ef034f6
27 Nov, 2023 1 commit
- Add missing check for K padding in XDL GEMM (#1056) · 60ecfd73
  Bartlomiej Wroblewski authored Nov 27, 2023
  
  60ecfd73
25 Nov, 2023 1 commit

Add basic support for direct loads from global to LDS (#999) · 627054b9

Bartlomiej Wroblewski authored Nov 25, 2023

* Add basic support for direct loads from global to LDS

* Clean the code and comments

* Add support for fp16

* Add comments

* Add check for thread cluster lengths

* Align non-direct-load fp16 example

* Small fixes

* Extend IsSupported to check for supported GPU gens

* Build examples only on the supported HW

* Do not throw when instance not supported in 04 example

* Review: Apply review suggestions

* Review: small fix

* Review: small fix

627054b9

17 Nov, 2023 1 commit

Improve 4k gemm perf (#1047) · e8cddfdc

zjing14 authored Nov 17, 2023



* improve 4k gemm perf

* add f8 instances

* format

---------
Co-authored-by: Jing Zhang <jizha@amd.com>

e8cddfdc

15 Nov, 2023 2 commits

Log CDEBlockTransferScalarPerVector_NPerBlock in conv fwd multiD xdl (#1042) · 1fefd82e

Bartłomiej Kocot authored Nov 15, 2023

* Log CDEBlockTransferScalarPerVector_NPerBlock in conv_fwd_multi_d_xdl implementation

* Log CDEBlockTransferScalarPerVector_NPerBlock in conv fwd multiD xdl

1fefd82e

Fix check for conv Fwd Filter1x1Pad0 (#1040) · 3ef3102f
Bartłomiej Kocot authored Nov 15, 2023
```
* Fix check for conv Fwd Filter1x1Pad0

* Fix check for conv Fwd Filter1x1Pad0
```
3ef3102f

14 Nov, 2023 1 commit

Introduce multiABD api and deprecate multiD (#1035) · f2398f61

Bartłomiej Kocot authored Nov 14, 2023

* Introduce multiABD api and deprecate multiD

* Replace multiD with multiABD

* Mark structures as deprecated

* Change doxygen deprecated to note to avoid warnings

f2398f61