Commits · b7c6350f145a82cfc64814f13a6595ee8be90502 · gaoqiong / composable_kernel

20 Oct, 2023 5 commits
- fix the condition syntax · b7c6350f
  illsilin authored Oct 20, 2023
  
  b7c6350f
- Merge branch 'develop' into lwpck-976 · 1b5af83d
  illsilin authored Oct 20, 2023
  
  1b5af83d
- do not use sccache with staging compiler · aac26d32
  illsilin authored Oct 20, 2023
  
  aac26d32
- use the new sccache redis server · 16402e96
  illsilin authored Oct 20, 2023
  
  16402e96
- Fix bf8 conversion issues (#1003) · 1fd27d52
  Rostyslav Geyyer authored Oct 20, 2023
```
* Fix the conversion

* Add bf8 functionality

* Enable example on MI200 as well
```
  1fd27d52
19 Oct, 2023 5 commits

Fix the DL kernel issues on Navi3x. (#998) · f7331c60
Illia Silin authored Oct 19, 2023
```
* apply the patch for dl kernels on gfx11

* build DL kernels on navi32 CI
```
f7331c60

Qianfeng authored Oct 20, 2023

* reinterpret_cast to const char* in dumpBufferToFile to be compatible with both const and non-const input pointers

* Add seed input to GeneratorTensor_4 for normal_distribution generator

* Add GetTypeString() for DeviceElementwiseImpl

* Add HIP_CHECK_ERROR macro

b4fc4d0b

Extend available elementwise operations with conv examples (#995) · 82f3a835

Bartłomiej Kocot authored Oct 19, 2023

* Extend available elementwise operations with conv examples

* Fixes

* Remove not needed convert

* Update CMakeFile and dir name

82f3a835

Avoid force setting ENABLE_PIPELINE_V2_OPT to OFF (#961) · deef92d5
Po Yen Chen authored Oct 19, 2023
```
* Avoid force setting ENABLE_PIPELINE_V2_OPT to OFF

* Remove compilation option variable MAX_ILP_OPTS
```
deef92d5
Change 1d,2d,... to 1D,2D,... (#997) · 0abc0f87
Bartlomiej Wroblewski authored Oct 19, 2023

0abc0f87

18 Oct, 2023 4 commits

Layernorm and groupnorm support to save mean and inverse std in forward (#929) · 3696fe1c

rocking authored Oct 19, 2023

* save mean and inverse std in normalization

* Save mean and inverse std in splitK

* Vector save mean and inv std

* Modify instance for save mean and std

* simplify the layernorm example

* Save mean and std in groupnorm example

* Save mean and inv std in ckProfiler and test

* Remove compute data type from base class

* Save mean and inv std in client example

* Add changelog

* clang format

* Fix compile error

* Refine naming

* Avoid error in bf16

* revert changelog

3696fe1c

fixed math-ci error; suspend a warning (#996) · 58338bb2
zjing14 authored Oct 18, 2023
```
Co-authored-by: Jing Zhang <jizha@amd.com>
```
58338bb2

Clean DTYPES conditions in CMake (#974) · bf435140

zjing14 authored Oct 18, 2023



* Add a condition to build fp8 instances

* simplified buffer_load/store

* add bfp8/fp8

* fixed

* remove all f8/bf8 condition include folder

* fixed cmake conditions

* fixed DTYPES=fp16/bfp16

* fix

* fixed buffer_load

* fixed buffer_store

* fix

* clean example cmake files

* fixed ci

* fixed cit

---------
Co-authored-by: Rostyslav Geyyer <rosty.geyyer@amd.com>
Co-authored-by: Jing Zhang <jizha@amd.com>

bf435140

Add contraction_multi_abd (#972) · 1cc36ba5

zjing14 authored Oct 17, 2023



* add gridwise_multi_abd

* move element_op into RunRead

* merge element_wise op with data read

* add multiABD example

* allow packed elementwise_op

* changed example

* clean

* clean

* add is_detected

* fix

* minor fix

* add scaleAdd_vec4 example

* init commit for contraction_multi_ABD

* add examples

* add examples of multiA and broadcast

* update example

* fixed comments

* Update cmake-ck-dev.sh

* Update cmake-ck-dev.sh

* Add comments into the example

* Update CMakeLists.txt

---------
Co-authored-by: Jing Zhang <jizha@amd.com>

1cc36ba5

17 Oct, 2023 9 commits
- add redis and ping packages, and redis port · bf52b430
  illsilin authored Oct 17, 2023
  
  bf52b430
- define SCCACHE_REDIS · cacc7f70
  illsilin authored Oct 17, 2023
  
  cacc7f70
- try using redis server with sccache · f6ef6163
  illsilin authored Oct 17, 2023
  
  f6ef6163
- added ab_elementwise_op support into splitK Gemm (#956) · bf0addb5
  zjing14 authored Oct 17, 2023
```
* add ab_elementwise

* fixed ci

* fixed a merge issue

* fixed pr comments

* fixed a conflict

* remove 61_example

---------
Co-authored-by: Jing Zhang <jizha@amd.com>
```
  bf0addb5
- Add grouped conv bwd weight wmma (#985) · 16d7c4d2
  Bartłomiej Kocot authored Oct 17, 2023
```
* Add grouped conv bwd weight wmma

* Update README, changelog, profiler

* Minor fixes

* Fix grouped conv bwd wei dl kernel

* Minor fixes

* Minor stylistic fixes
```
  16d7c4d2
- try using sccache --start-server instead of wrapper · c9124e3d
  illsilin authored Oct 16, 2023
  
  c9124e3d
- use /tmp instead of /usr/local for cache · 41cf1699
  illsilin authored Oct 16, 2023
  
  41cf1699
- set the paths before calling the sccache_wrapper · 05b0b90c
  illsilin authored Oct 16, 2023
  
  05b0b90c
- run sccache_wrapper before build if ccache server found · 6ecef55f
  illsilin authored Oct 16, 2023
  
  6ecef55f
16 Oct, 2023 8 commits
- workaround with float (#992) · 39430bfd
  zjing14 authored Oct 16, 2023
```
Co-authored-by: Jing Zhang <jizha@amd.com>
```
  39430bfd
- fix the pymysql package issue · 35cf9b18
  illsilin authored Oct 16, 2023
  
  35cf9b18
- fix the package version syntax · 5fe19b4a
  illsilin authored Oct 16, 2023
  
  5fe19b4a
- add sccashe_wrapper.sh script · 2e7ea550
  illsilin authored Oct 16, 2023
  
  2e7ea550
- put ccache back temporarily to avoid breaking other CI jobs · bc7d5d5e
  illsilin authored Oct 16, 2023
  
  bc7d5d5e
- Merge branch 'develop' into lwpck-976 · 8ff32881
  illsilin authored Oct 16, 2023
  
  8ff32881
- replace ccache with sccache, pin package versions · 8ebc0d8c
  illsilin authored Oct 16, 2023
  
  8ebc0d8c
- Add hipTensor build and test to CK CI. (#990) · 707ad002
  Illia Silin authored Oct 16, 2023
```
* add a hipTensor test to CI

* use jenkins git plugin

* change hipTensor folder location in CI

* change the git method for hipTensor

* run tests usign ctest

* check the hipTensor contents

* only build hipTensor on MI100/200

* pull hipTensor as zip archive

* fix jenkins syntax

* add path to the CK installation

* combine build commands into one shell

* change jenkins syntax for CK installer path

* try different syntax

* allow unzip overwrite

* fix jenkins file syntax

* remove any old versions of hipTensor before building

* add option to select hipTensor branch for testing
```
  707ad002
13 Oct, 2023 2 commits

Add splitk gemm fp16 @ fp16 with fp8 compute instances (#983) · fa753f27
Rostyslav Geyyer authored Oct 13, 2023
```
* Add ComputeType

* Update for compatibility

* Add instances

* Update profiler api
```
fa753f27

add vector_type support into thread_copy_v3r1 (#969) · 2ce9b56c

zjing14 authored Oct 13, 2023



* add vector_type support into thread_copy_v3r1

* remove unncessary type_convert

* fixed datatype

* fixed dataType

* changed API with is_packx_invocable

* changed example

* add missing cmake file

* fixed ci

* fixed cmake

---------
Co-authored-by: Jing Zhang <jizha@amd.com>

2ce9b56c

12 Oct, 2023 2 commits

Bump gitpython from 3.1.31 to 3.1.35 in /docs/sphinx (#898) · a3c80265

dependabot[bot] authored Oct 12, 2023

Bumps [gitpython](https://github.com/gitpython-developers/GitPython) from 3.1.31 to 3.1.35.
- [Release notes](https://github.com/gitpython-developers/GitPython/releases)
- [Changelog](https://github.com/gitpython-developers/GitPython/blob/main/CHANGES)
- [Commits](https://github.com/gitpython-developers/GitPython/compare/3.1.31...3.1.35

)

---
updated-dependencies:
- dependency-name: gitpython
  dependency-type: indirect
...
Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>

a3c80265

simplified buffer_load/store (#971) · f3b02ecf

zjing14 authored Oct 11, 2023



* simplified buffer_load/store

* add bfp8/fp8

* fixed

* fixed buffer_load

* fixed buffer_store

---------
Co-authored-by: Jing Zhang <jizha@amd.com>

f3b02ecf

11 Oct, 2023 2 commits

Revert "Grouped Gemm with looping over the tiles. (#788)" (#982) · c99323be
zjing14 authored Oct 11, 2023
```
This reverts commit a4f72a31.
```
c99323be

Grouped Gemm with looping over the tiles. (#788) · a4f72a31

Adam Osewski authored Oct 11, 2023



* Introduce LocalBlockToCTileMap.

* Change the signature of CalculateBottomIndex() function which now does
not accept any argument. The B2C map which is already passed as an
argument to the kernel Run function is calculating block's local id
already outside at kernel entry point __global__ function.
The LocalB2C map stores as members local block ID.

* Use LocalBlockToCTile map in device ops.

* First draft of tile loop work distribution.

* Fix typo.

* Simplify kernel arguments.

Calculate descriptors & B2C maps on the device.

* Use looping kernel.

* Fix B2C constructor.

* Fix Navi21 errors.

* Calculate tile start/end in device kernel.

* Change Run API to accept user provided workspace buffer.

* Add new line at EOF.

* Move Gemm KernelArguments to device op interface.

* Remove unused code.

* Update API.

* Launch grid size which is min of occupancy vs tile count

* Get back to use constant memory for gemm descriptors.

* Remove unused code.

* Add default virtual method implementation.

* Update comments to conform with doxygen style.

* Fix doc style and unused parameters.

* Add thread cluster lengths to kernel name.

* Remove old splitk impl and replace it with tile looping one.

* Modify instances.

* set KPerBlock to 64
* maximize wherever possible vector load size.

* Fix instances cluster lengths.

* Change comment style.

* Use 128b store where possible in instances.

* Update test cases, since KPerBlock has doubled.

* Update output stream operator for Sequence.

* Add pipeline version to GroupedGEMM device op type string.

* Fix pipeline version type logging.

* Fix input tensors type after merge.

* Fix compiler error.

* Fix output stream operator for Pipeline version.

* Store using 128b.

* Set of instances with kpb 32/64

* Limit number of instances

* Remove commented out instances.

* Fix function name.

* Limit the number of instances.

Add pipline version to the regular instances

* Change thr cluster layout for reading B tensor.

* disabled failed instances

---------
Co-authored-by: Adam Osewski <aosewski@amd.com>
Co-authored-by: zjing14 <zhangjing14@gmail.com>
Co-authored-by: Jing Zhang <jizha@amd.com>

a4f72a31

10 Oct, 2023 2 commits

Fix MNKPadding in gridwise_gemm_xdlops_v2r3 (#981) · 98c80714
Bartłomiej Kocot authored Oct 10, 2023

98c80714

Fixed f8_gemm NaN (#975) · ac9595a9

zjing14 authored Oct 10, 2023



* workaround nan problem by changing output to fp16

* enable f8/bf8 gemm tests on MI200

* workaround f16 to f8 conversion

---------
Co-authored-by: Jing Zhang <jizha@amd.com>

ac9595a9

05 Oct, 2023 1 commit
- Replace CMake `return` from later CMake (#970) · 59136091
  Lauren Wrubleski authored Oct 05, 2023
  
  59136091