Commits · f3bbfe3efe7a4b1065b5abea5af7b0ce6e2aeb40 · gaoqiong / composable_kernel_ROCM

15 Nov, 2024 2 commits
- re-enable fp8 gemms in ckProfiler (#1667) · b4a79045
  Illia Silin authored Nov 14, 2024
  
  b4a79045
- re-enable coerce-illegal-types flag for rocm6.3 (#1668) · 3b6a481e
  Illia Silin authored Nov 14, 2024
  
  3b6a481e
05 Nov, 2024 1 commit
- remove gfx940;gfx941 from default target lists (#1640) · 54440cf5
  Illia Silin authored Nov 05, 2024
  
  54440cf5
01 Nov, 2024 1 commit

Illia Silin authored Oct 31, 2024

* disable fp8 gemm_universal on gfx90a and gfx908 by default

* fix cmake syntax

* fix clang format

* add ifdefs in amd_xdlops

* disable fp8 gemm instances on gfx90a by default

* update readme

03c6448b

23 Oct, 2024 1 commit
- fix the logic of enabling XDL and WMMA instances (#1595) · 8e22e1ae
  Illia Silin authored Oct 23, 2024
  
  8e22e1ae
18 Oct, 2024 1 commit
- add the lsr-drop-solution=1 compiler flag (#1582) · 88e6fa7f
  Illia Silin authored Oct 18, 2024
  
  88e6fa7f
10 Oct, 2024 1 commit
- only build tests and examples if user sets GPU_TARGETS (#1565) · f46a9eee
  Illia Silin authored Oct 10, 2024
  
  f46a9eee
09 Oct, 2024 1 commit
- fix the target selection logic (#1561) · 2e1165c1
  Illia Silin authored Oct 09, 2024
  
  2e1165c1
07 Oct, 2024 2 commits
- add a CK_USE_CODEGEN build argument to enable codegen (#1552) · 7733ae16
  Illia Silin authored Oct 07, 2024
```
* add a CK_USE_CODEGEN build argument to enable codegen

* fix cmake codegen logic
```
  7733ae16
- Fix build logic using GRU_ARCHS. (#1536) · 7d8ea5f0
  Illia Silin authored Oct 07, 2024
```
* update build logic with GPU_ARCHS

* fix the GPU_ARCHS build for codegen

* unset GPU_TARGETS when GPU_ARCHS are set
```
  7d8ea5f0
04 Oct, 2024 1 commit

Codegen build (#1526) · b545de17

arai713 authored Oct 04, 2024

* updating codegen build for MIOpen access: adding .cmake for codegen component

(cherry picked from commit 652a7c04)

* updating CMake

(cherry picked from commit a685822e)

b545de17

13 Sep, 2024 1 commit

Customize filesystem in CK for legacy systems (#1509) · 81bc1496

Jun Liu authored Sep 13, 2024



* Legacy support: customized filesystem

* Update cmakefile for python alternative path

* fix build issues

* CK has no boost dependency

* More fixes to issues found on legay systems

* fix clang format issue

* Check if blob is correctly generated in cmake

* fix the python issues

* add a compiler flag for codegen when using alternative python

* use target_link_options instead of target_compile_options

---------
Co-authored-by: illsilin <Illia.Silin@amd.com>

81bc1496

04 Sep, 2024 1 commit
- Add an option to select an alternative python version during build. (#1496) · 841009c5
  Illia Silin authored Sep 04, 2024
```
* locate a newwer version of python when -DRHEL=ON flag is set

* allow setting python version on cmake command line
```
  841009c5
23 Aug, 2024 1 commit
- fix codegen rtc lib build issue (#1485) · 25935b57
  Illia Silin authored Aug 23, 2024
  
  25935b57
22 Aug, 2024 1 commit

Codegen INSTANCES_ONLY build (#1468) · 967b1f0f

arai713 authored Aug 22, 2024



* initial push - altering codegen build

* fix the codegen cmake

* enable codegen build for gfx908 and gfx90a

* enable building codegen with INSTANCES_ONLY=ON

* updating ck_rtc

* remove gpu targets for codegen and rename tests

* make codegen tests dependencies of tests and check targets

---------
Co-authored-by: illsilin <Illia.Silin@amd.com>
Co-authored-by: Illia Silin <98187287+illsilin@users.noreply.github.com>

967b1f0f

16 Aug, 2024 1 commit

Re-enable fp8 types for all architectures. (#1470) · c8b6b642

Illia Silin authored Aug 16, 2024

* re-enable fp8 and bf8 for all targets

* restore the fp8 gemm instances

* re-enable conv_3d fp8 on all architectures

* diasble several fp8 gemm instances on all architectures except gfx94

* clang format fix

c8b6b642

15 Aug, 2024 1 commit

Check compiler flags before using (#1403) · 49769ec8

trixirt authored Aug 14, 2024



* Check compiler flags before using

The user's compiler may not support these flags, so check.
Resolves failures on Fedora.
Signed-off-by: Tom Rix <trix@redhat.com>

* fix syntax CMakeLists.txt

Fix syntax in the check_cxx_compiler_flag.

---------
Signed-off-by: Tom Rix <trix@redhat.com>
Co-authored-by: Tom Rix <trix@redhat.com>
Co-authored-by: Illia Silin <98187287+illsilin@users.noreply.github.com>

49769ec8

14 Aug, 2024 1 commit

[GEMM] gemm_universal related optimization (#1453) · 3049b546

Haocong WANG authored Aug 14, 2024



* replace buffer_atomic with global_atomic

* fixed global_atomic_add

* added bf16 atomic_add

* format

* clang-format-12

* clean

* clean

* add guards

* Update gtest.cmake

* enabled splitk_gemm_multi_d

* format

* add ckProfiler

* format

* fixed naming

* format

* clean

* clean

* add guards

* fix clang format

* format

* add kbatch printout

* clean

* Add rocm6.2 related gemm optimization

* Limit bf16 atomic usage

* remove redundant RCR gemm_universal instance

* Add RRR fp8 gemm universal instance

* Bug fix

* Add GPU_TARGET guard to FP8/BF8 target

* bug fix

* update cmake

* remove all fp8/bf8 example if arch not support

* Enable fp8 RRR support in ckProfiler

* limit greedy-reverse flag to gemm_universal in ckProfiler

---------
Co-authored-by: Jing Zhang <jizhan@fb.com>
Co-authored-by: Jing Zhang <jizhan@meta.com>
Co-authored-by: zjing14 <zhangjing14@gmail.com>
Co-authored-by: Illia Silin <98187287+illsilin@users.noreply.github.com>
Co-authored-by: illsilin <Illia.Silin@amd.com>

3049b546

09 Aug, 2024 1 commit

Codegen build w/CK (#1428) · da214a5a

arai713 authored Aug 09, 2024



* initial push

* cleaned up compiler errors

* removed commented code

* build codegen folder only for gfx9 targets

* remove separate stage for codegen tests from CI

* removed commented code from CMake

---------
Co-authored-by: Illia Silin <98187287+illsilin@users.noreply.github.com>
Co-authored-by: illsilin <Illia.Silin@amd.com>

da214a5a

08 Aug, 2024 1 commit
- check if the coerce-illegal-types flag is supported (#1451) · ae3b8ff8
  Illia Silin authored Aug 08, 2024
  
  ae3b8ff8
06 Aug, 2024 1 commit
- Fix ROCm 6.2 compiler not fully supporting gfx12 when building CK with INSTANCES_ONLY (#1446) · afbf6350
  Jun Liu authored Aug 06, 2024
  
  afbf6350
01 Aug, 2024 1 commit

Add compiler flags for ROCm versions 6.2+ (#1429) · d311c953

Illia Silin authored Aug 01, 2024

* add compiler flags to fix compiler issues

* fix typo.

* disable test_smfmac_op on all devices except gfx942

* specify full path to compiler in CI

d311c953

26 Jul, 2024 1 commit

Introduce cmake USE_GLIBCXX_ASSERTIONS option (#1404) · 733f33af

trixirt authored Jul 25, 2024



A standard option in Fedora packaging that is used to check
the correctness of c++ use of the standard c++ library.
Signed-off-by: Tom Rix <trix@redhat.com>
Co-authored-by: Illia Silin <98187287+illsilin@users.noreply.github.com>

733f33af

16 Jul, 2024 2 commits
- An option whether to colorize output during build (#1390) · 9cac2827
  Mateusz Ozga authored Jul 16, 2024
  
  9cac2827
- [ASAN builds] Modify the list of default targets for ASAN builds. (#1389) · 4c3107fd
  Illia Silin authored Jul 16, 2024
```
* add a build parameter to build only XNACK targets

* use ENABLE_ASAN_PACKAGING flag to set targets for ASAN builds

---------
Co-authored-by: Bartłomiej Kocot <barkocot@amd.com>
```
  4c3107fd
10 Jul, 2024 1 commit
- [gfx12] add gfx12 to the default target list (#1379) · a8eb8720
  Illia Silin authored Jul 10, 2024
  
  a8eb8720
27 Jun, 2024 1 commit
- Merging the gfx12 code into public repo. (#1362) · 941d1f7c
  Illia Silin authored Jun 27, 2024
  
  941d1f7c
19 Jun, 2024 1 commit
- Remove gfx900 and gfx906 from default target device to reduce package size (#1351) · 8db331a5
  zjing14 authored Jun 19, 2024
  
  8db331a5
22 May, 2024 1 commit

Select appropriate GPU targets for instances, tests, and examples. (#1304) · 7b027d56

Illia Silin authored May 22, 2024

* set individual gpu targets for instances, examples, tests

* fix path to hip compiler

* fix path to hip compiler once more

* aggregate device macros in ck_tile config header

* fix the cmake logic for instances

* fix clang format

* add gfx900 and gfx906 to default set of targets

7b027d56

10 May, 2024 1 commit
- Code clean-up (#1285) · 566b6480
  Illia Silin authored May 10, 2024
```
* code clean-up

* remove the profiling output samples
```
  566b6480
01 May, 2024 1 commit
- Downgrade minimum required python version to 3.6 (#1274) · 7797f7c7
  Illia Silin authored May 01, 2024
  
  7797f7c7
18 Apr, 2024 1 commit
- Upgrade to ROCm6.1 and turn on the -enable-post-misched=0 compiler flag. (#1250) · caae537d
  Illia Silin authored Apr 18, 2024
```
* add rocm6.1 docker and make it default for CI

* fix typo

* move the rocm6.1 image into public dockerhub repo
```
  caae537d
16 Apr, 2024 1 commit

introducing ck_tile! (#1216) · db376dd8

carlushuang authored Apr 16, 2024

* enable gfx940

* switch between intrinsic mfma routines on mi100/200 and mi300

* fix mfma_int8 on MI300

* disable 2 int8 examples on MI300

* Update cmake-ck-dev.sh

* restore gitignore file

* modify Jenkinsfile to the internal repo

* Bump rocm-docs-core from 0.24.0 to 0.29.0 in /docs/sphinx

Bumps [rocm-docs-core](https://github.com/RadeonOpenCompute/rocm-docs-core) from 0.24.0 to 0.29.0.
- [Release notes](https://github.com/RadeonOpenCompute/rocm-docs-core/releases)
- [Changelog](https://github.com/RadeonOpenCompute/rocm-docs-core/blob/develop/CHANGELOG.md)
- [Commits](https://github.com/RadeonOpenCompute/rocm-docs-core/compare/v0.24.0...v0.29.0

)

---
updated-dependencies:
- dependency-name: rocm-docs-core
  dependency-type: direct:production
  update-type: version-update:semver-minor
...
Signed-off-by: dependabot[bot] <support@github.com>

* initial enablement of gfx950

* fix clang format

* disable examples 31 and 41 int8 on gfx950

* add code

* fix build wip

* fix xx

* now can build

* naming

* minor fix

* wip fix

* fix macro for exp2; fix warpgemm a/b in transposedC

* unify as tuple_array

* Update the required Python version to 3.9

* Update executable name in test scripts

* re-structure tuple/array to avoid spill

* Merge function templates

* Fix format

* Add constraint to array<> ctor

* Re-use function

* Some minor changes

* remove wrong code in store_raw()

* fix compile issue in transpose

* Rename enum
Rename 'cood_transform_enum' to 'coord_transform_enum'

* let more integral_constant->constant, and formating

* make sure thread_buffer can be tuple/array

* temp fix buffer_store spill

* not using custom data type by default, now we can have ISA-level same code as opt_padding

* fix compile error, fp8 not ready now

* fix fp8 duplicated move/shift/and/or problem

* Default use CK_TILE_FLOAT_TO_FP8_STOCHASTIC rounding mode

* fix scratch in fp8 kernel

* update some readme

* fix merge from upstream

* sync with upstream

* sync upstream again

* sync 22

* remove unused

* fix clang-format

* update README of ck_tile example

* fix several issue

* let python version to be 3.8 as minimal

* remove ck_tile example from default cmake target like all/install/check

* remove mistake

* 1).support receipe in generate.py 2).use simplified mask type 3).change left/right to pass into karg

* fix some bug in group-mode masking and codegen. update README

* F8 quantization for FMHA forward (#1224)

* Add SAccElementFunction, PComputeElementFunction, OAccElementFunction in pipeline

* Add element function to fmha api

* Adjust P elementwise function

* Fix bug of elementwise op, our elementwise op is not inout

* Add some elementwise op, prepare to quantization

* Let generate.py can generate different elementwise function

* To prevent compiler issue, remove the elementwise function we have not used.

* Remove f8 pipeline, we should share the same pipeline even in f8

* Remove remove_cvref_t

* Avoid warning

* Fix wrong fp8 QK/KV block gemm setting

* Check fp8 rounding error in check_err()

* Set fp8 rounding error for check_err()

* Use CK_TILE_FLOAT_TO_FP8_STANDARD as default fp8 rounding mode

* 1. codgen the f8 api and kernel
2. f8 host code

* prevent warning in filter mode

* Remove not-in-use elementwise function kargs

* Remove more not-in-use elementwise function kargs

* Small refinements in C++ source files

* Use conditional_t<> to simplify code

* Support heterogeneous argument for binary function types

* Re-use already-existing scales<> functor template

* Fix wrong value produced by saturating

* Generalize the composes<> template

* Unify saturates<> implementation

* Fix type errors in composes<>

* Extend less_equal<>

* Reuse the existing template less_equal<> in check_err()

* Add equal<float> & equal<double>

* Rename check_err() parameter

* Rename check_err() parameter

* Add FIXME comment for adding new macro in future

* Remove unnecessary cast to void

* Eliminate duplicated code

* Avoid dividing api pool into more than 2 groups

* Use more clear variable names

* Use affirmative condition in if stmt

* Remove blank lines

* Donot perfect forwarding in composes<>

* To fix compile error, revert generate.py back to 4439cc107dd90302d68a6494bdd33113318709f8

* Fix bug of p element function

* Add compute element op to host softmax

* Remove element function in api interface

* Extract user parameter

* Rename pscale and oscale variable

* rename f8 to fp8

* rename more f8 to fp8

* Add pipeline::operator() without element_functor

* 1. Remove deprecated pipeline enum
2. Refine host code parameter

* Use quantization range as input

* 1. Rename max_dtype to dtype_max.
2. Rename scale to scale_s
3.Add init description

* Refine description

* prevent early return

* unify _squant kernel name in cpp, update README

* Adjust the default range.

* Refine error message and bias range

* Add fp8 benchmark and smoke test

* fix fp8 swizzle_factor=4 case

---------
Co-authored-by: Po Yen Chen <PoYen.Chen@amd.com>
Co-authored-by: carlushuang <carlus.huang@amd.com>

---------
Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: illsilin <Illia.Silin@amd.com>
Co-authored-by: Illia Silin <98187287+illsilin@users.noreply.github.com>
Co-authored-by: Jing Zhang <jizha@amd.com>
Co-authored-by: zjing14 <zhangjing14@gmail.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Po-Yen, Chen <PoYen.Chen@amd.com>
Co-authored-by: rocking <ChunYu.Lai@amd.com>

db376dd8

12 Apr, 2024 1 commit
- Update the config.h after the CK_USE_XDL/WMMA are set. (#1236) · 7cdf5a96
  Illia Silin authored Apr 12, 2024
```
* pass XDL and WMMA macros to libs that use CK

* update config.h after XDL and WMMA macros get set
```
  7cdf5a96
02 Apr, 2024 1 commit

Split the instances by architecture. (#1223) · ae57e593

Illia Silin authored Apr 02, 2024

* parse examples inside the add_example_executable function

* fix the example 64 cmake file

* add xdl flag to the gemm_bias_softmax_gemm_permute example

* add filtering of tests based on architecture type

* enable test_grouped_gemm for gfx9 only

* enable test_transpose only for gfx9

* only linnk test_transpose if it gets built

* split the gemm instances by architectures

* split gemm_bilinear,grouped_conv_bwd_weight instances by targets

* split instances by architecture

* split grouped_conv instances by architecture

* fix clang format

* fix the if-else logic in group_conv headers

* small fix for grouped convolution instances

* fix the grouped conv bwd weight dl instances

* fix client examples

* only enable client examples 3 and 4 on gfx9

* set the gfx9 macro

* make sure the architecture macros are set by cmake

* use separate set of xdl/wmma flags for host code

* sinmplify the main cmake file

* add conv_fwd_bf8 instance declaration

ae57e593

03 Jan, 2024 1 commit
- fix the cmake option syntax (#1117) · fbf31a2e
  Illia Silin authored Jan 03, 2024
  
  fbf31a2e
02 Jan, 2024 1 commit
- adding -Wno-switch-default compiler flag (#1115) · b268f273
  Illia Silin authored Jan 02, 2024
  
  b268f273
20 Dec, 2023 1 commit

enable compilation of INSTANCES_ONLY for Windows (#1082) · fb5bd51b

Artur Wojcik authored Dec 20, 2023



* enable compilation of INSTANCES_ONLY for Windows

* suppress ROCMChecks warnings on GoogleTests

* suppress -Wfloat-equal warning on GoogleTests

---------
Co-authored-by: Illia Silin <98187287+illsilin@users.noreply.github.com>

fb5bd51b

19 Dec, 2023 2 commits
- ROCm 6.0 replaces all __HIP_PLATFORM_HCC__ with __HIP_PLATFORM_AMD__ (#1106) · 3ab1838f
  Jun Liu authored Dec 19, 2023
```
* ROCm 6.0 replaces all __HIP_PLATFORM_HCC__ with __HIP_PLATFORM_AMD__

* make it backward compatible

* Update .clang-tidy

* Update ClangTidy.cmake
```
  3ab1838f
- add -Wno-pass-failed compiler flag (#1105) · 3726a173
  Illia Silin authored Dec 19, 2023
  
  3726a173