Commits · f4855e8c8516de81a3bb66dc086aee6e0db1ec22 · gaoqiong / composable_kernel_ROCM

26 Dec, 2023 1 commit
- Merge branch 'develop' into amd-develop · f4855e8c
  Jun Liu authored Dec 25, 2023
  
  f4855e8c
23 Dec, 2023 1 commit
- Fix results verify in test_tensor (#1109) · 20b1ae7c
  Bartłomiej Kocot authored Dec 23, 2023
  
  20b1ae7c
20 Dec, 2023 2 commits

Bump rocm-docs-core from 0.30.2 to 0.30.3 in /docs/sphinx (#1107) · 78eb3f0b

dependabot[bot] authored Dec 20, 2023

Bumps [rocm-docs-core](https://github.com/RadeonOpenCompute/rocm-docs-core) from 0.30.2 to 0.30.3.
- [Release notes](https://github.com/RadeonOpenCompute/rocm-docs-core/releases)
- [Changelog](https://github.com/RadeonOpenCompute/rocm-docs-core/blob/develop/CHANGELOG.md)
- [Commits](https://github.com/RadeonOpenCompute/rocm-docs-core/compare/v0.30.2...v0.30.3

)

---
updated-dependencies:
- dependency-name: rocm-docs-core
  dependency-type: direct:production
  update-type: version-update:semver-patch
...
Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>

78eb3f0b

enable compilation of INSTANCES_ONLY for Windows (#1082) · fb5bd51b

Artur Wojcik authored Dec 20, 2023



* enable compilation of INSTANCES_ONLY for Windows

* suppress ROCMChecks warnings on GoogleTests

* suppress -Wfloat-equal warning on GoogleTests

---------
Co-authored-by: Illia Silin <98187287+illsilin@users.noreply.github.com>

fb5bd51b

19 Dec, 2023 6 commits

Merge branch 'develop' into amd-develop · d0f355a3
Jun Liu authored Dec 19, 2023

d0f355a3

Remove index tensor in avgpool (#1093) · b305a29e

rocking authored Dec 19, 2023



* Remove index tensor

* fix syntax

---------
Co-authored-by: Illia Silin <98187287+illsilin@users.noreply.github.com>
Co-authored-by: illsilin <Illia.Silin@amd.com>

b305a29e

Bump rocm-docs-core from 0.30.1 to 0.30.2 in /docs/sphinx (#1104) · a167e3c7

dependabot[bot] authored Dec 19, 2023

Bumps [rocm-docs-core](https://github.com/RadeonOpenCompute/rocm-docs-core) from 0.30.1 to 0.30.2.
- [Release notes](https://github.com/RadeonOpenCompute/rocm-docs-core/releases)
- [Changelog](https://github.com/RadeonOpenCompute/rocm-docs-core/blob/develop/CHANGELOG.md)
- [Commits](https://github.com/RadeonOpenCompute/rocm-docs-core/compare/v0.30.1...v0.30.2

)

---
updated-dependencies:
- dependency-name: rocm-docs-core
  dependency-type: direct:production
  update-type: version-update:semver-patch
...
Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>

a167e3c7

ROCm 6.0 replaces all __HIP_PLATFORM_HCC__ with __HIP_PLATFORM_AMD__ (#1106) · 3ab1838f

Jun Liu authored Dec 19, 2023

* ROCm 6.0 replaces all __HIP_PLATFORM_HCC__ with __HIP_PLATFORM_AMD__

* make it backward compatible

* Update .clang-tidy

* Update ClangTidy.cmake

3ab1838f

add -Wno-pass-failed compiler flag (#1105) · 3726a173
Illia Silin authored Dec 19, 2023

3726a173

Hip tensor permute unit test (#1068) · 12a8883c

arai713 authored Dec 18, 2023

* adding files for F32 example

* adding functioning implementation with scalar multiplication and unary operator support

* added fp 16 type check in unary square

* updating scalar multiplication as an operator

* functioning version with scalar operator

* changing strides for col major

* updated column major implementation

* working column major implementation

* cleaned up comments, rearranged/renamed files

* small edits to 3d transpose profiler

* adding test/profiler/instance files for hipTensor permute unit test

* added more test instances

* cleaned up errors, randomized input tensor, added more instances

* turned off time printouts

* removed conflicting transpose profiler

* rearranged some files

12a8883c

18 Dec, 2023 2 commits

layernorm and groupnorm backward data (#1083) · a69aa2a1

rocking authored Dec 19, 2023

* rename folder

* Add type string

* Remove typo

* Add deviceOp to backward x

* Add comment to describe the behavior of backward normalization

* Add kernel function, prepare to implement

* implement generic kernel

* Check vector size

* Add sweep once pipeline for small reduce size

* Fix bug of KRaw_ error

* Fix bug of dx stride

* sanity check for mean and rstd

* backward x for groupnorm

* Add bwd x instance

* add layernorm 2d bwd gamma beta instances

* Change save mean var type from f32 to f16 in f16 mode

* Change the example to f16

* Add groupnorm bwd gamma beta instance

* Add groupnorm bwd x instance

* Fix naming

* Add layernorm bwd x ckprofiler

* Add groupnorm bwd x profiler

* clang format

* Rename bwd x to bwd data

* Fix bug of verification in profiler

* Add test of layernorm and groupnorm bwd data

* Add missing cmake

* Add layernorm2d bwd data

* rename fwd example

* Add groupnorm client example

* Fix typo. replace Invarient with Invariant

* Add checking before running the best instance

a69aa2a1

Optimize fp16 direct load GEMM instances (#1086) · ad0a8e4c

Bartlomiej Wroblewski authored Dec 18, 2023

This PR optimizes fp16 instances of direct load GEMM kernel introduced in #999 and #1052.

Measured the performance of new instances on CDNA2 GPU and compared it against the performance of the best non-direct-load GEMM instances. Used 76 different GEMM problems.
On average, this change improves the performance of the tested problems by 47%. For cases known as latency-bound, the speedup is around 126%.

ad0a8e4c

17 Dec, 2023 1 commit
- Merge branch 'develop' into amd-develop · 55a89c74
  Jun Liu authored Dec 16, 2023
  
  55a89c74
16 Dec, 2023 1 commit

Upgrade the default compiler to ROCm6.0 release. (#1103) · dcedf363

Illia Silin authored Dec 16, 2023

* upgrade to rocm6.0 compiler

* move rocm6.0 from private to public repo

* switch to testing hipTensor mainline in CI

dcedf363

15 Dec, 2023 3 commits

Adding Issue Template (#1094) · 3246d1f6

abhimeda authored Dec 15, 2023



* Add files via upload

* fixed extra space typo

* add mi300 GPU architectures and rocm versions 5.6.1 and 6.0.0

---------
Co-authored-by: illsilin <Illia.Silin@amd.com>
Co-authored-by: Illia Silin <98187287+illsilin@users.noreply.github.com>

3246d1f6

Add tensor structure to wrapper (#1098) · 07092d68
Bartłomiej Kocot authored Dec 15, 2023
```
* Add tensor structure to wrapper

* update changelog

* Fix names

* Comment fixes
```
07092d68

cmake: Add CK_PARALLEL_LINK_JOBS and CK_PARALLEL_COMPILE_JOBS options (#1063) · efaf3106

trixirt authored Dec 14, 2023



Copied from the llvm-project LLVM_PARALLEL_*_JOBS

Concurrent linking can break the build as well as having too many
compile jobs for the avaiable memory.  These options allow the user
to fine tune the build to fit within their machines memory
constraints.

An example use on linux is
COMPILE_JOBS=`cat /proc/cpuinfo | grep -m 1 'cpu cores' | awk '{ print $4 }'`
if [ ${COMPILE_JOBS}x = x ]; then
  COMPILE_JOBS=1
fi
BUILD_MEM=4
MEM_KB=0
MEM_KB=`cat /proc/meminfo | grep MemTotal | awk '{ print $2 }'`
MEM_MB=`eval "expr ${MEM_KB} / 1024"`
MEM_GB=`eval "expr ${MEM_MB} / 1024"`
COMPILE_JOBS_MEM=`eval "expr 1 + ${MEM_GB} / ${BUILD_MEM}"`
if [ "$COMPILE_JOBS_MEM" -lt "$COMPILE_JOBS" ]; then
  COMPILE_JOBS=$COMPILE_JOBS_MEM
fi
LINK_MEM=32
LINK_JOBS=`eval "expr 1 + ${MEM_GB} / ${LINK_MEM}"`

cmake -G Ninja -DCK_PARALLEL_LINK_JOBS=$LINK_JOBS
               -DCK_PARALLEL_COMPILE_JOBS=$COMPILE_JOBS
Signed-off-by: Tom Rix <trix@redhat.com>

efaf3106

14 Dec, 2023 1 commit

fix typo (#1067) · 281f8369

Lisa authored Dec 14, 2023


Co-authored-by: Illia Silin <98187287+illsilin@users.noreply.github.com>

281f8369

13 Dec, 2023 3 commits
- [Doc][Werror] Fix security alerts and sync with MIOpen (#1085) · 3a3b98ef
  Jun Liu authored Dec 13, 2023
```
* fix Werror unused-parameter

* sync doc requirements

* fix blank space format

* fix dependency issue
```
  3a3b98ef
- Fix the bugs (#1099) · 0dacd895
  Rostyslav Geyyer authored Dec 13, 2023
  
  0dacd895
- Fix the bugs (#1099) · 6891e4d1
  Rostyslav Geyyer authored Dec 13, 2023
  
  6891e4d1
12 Dec, 2023 1 commit

disabling some fp8 gemm instances to reduce build time (#1084) · c004e0d9

Illia Silin authored Dec 11, 2023

* disabling some fp8 gemm instances to reduce build time

* disable fp8 gemm instances to reduce build time

* remove the unused variable

* build fp8 gemm default and padded instances separately

* fix include pathsc

c004e0d9

11 Dec, 2023 1 commit

Fix IsSupported check in the contraction op (#1066) · 89ee4746

Bartlomiej Wroblewski authored Dec 11, 2023

Current implementation of IsSupported method in contraction ops does not cover a lot of possible cases in which ScalarPerVector cannot really be used to read A, B or D, or write E.

This PR extends both the regular and multiABD contraction ops with improved checks and also adds new instances with smaller values of ScalarPerVector to support instances that are not supported by other instances.

89ee4746

08 Dec, 2023 3 commits
- fix clang format (#1095) · f199035b
  Illia Silin authored Dec 08, 2023
  
  f199035b
- Add F8 dtype definition in f16_f8_f16 gemm instances (#1092) · b4dcd580
  Nicolas Macchioni authored Dec 08, 2023
  
  b4dcd580
- Support broadcast for bias in grouped conv fwd (#1081) · f8369848
  Bartłomiej Kocot authored Dec 08, 2023
```
* Support broadcast for bias in grouped conv fwd

* Fix comment

* Comment fixes

* Remove GK layout
```
  f8369848
07 Dec, 2023 3 commits

Switch from ROCmSoftwarePlatform to ROCm org (#1091) · d939411d

Illia Silin authored Dec 07, 2023

* switch from ROCmSoftwarePlatform to ROCm org

* replace ROCmSoftwarePlatform with ROCm in few more places

d939411d

remove imcomplete transpose profiler (#1088) · 33600202

zjing14 authored Dec 07, 2023


Co-authored-by: Jing Zhang <jizha@amd.com>
Co-authored-by: Illia Silin <98187287+illsilin@users.noreply.github.com>

33600202

Bump rocm-docs-core from 0.29.0 to 0.30.1 in /docs/sphinx (#1090) · 957281ce

dependabot[bot] authored Dec 07, 2023

Bumps [rocm-docs-core](https://github.com/RadeonOpenCompute/rocm-docs-core) from 0.29.0 to 0.30.1.
- [Release notes](https://github.com/RadeonOpenCompute/rocm-docs-core/releases)
- [Changelog](https://github.com/RadeonOpenCompute/rocm-docs-core/blob/develop/CHANGELOG.md)
- [Commits](https://github.com/RadeonOpenCompute/rocm-docs-core/compare/v0.29.0...v0.30.1

)

---
updated-dependencies:
- dependency-name: rocm-docs-core
  dependency-type: direct:production
  update-type: version-update:semver-minor
...
Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>

957281ce

06 Dec, 2023 2 commits

Fix the CI builds using clang++ directly. (#1087) · 6896c3b0

Illia Silin authored Dec 06, 2023

* turn on -O3 compiler flag explicitly

* change cmake syntax for CI

* modify cmake line breaks in jenkinsfile

6896c3b0

Introduce wrapper library (#1071) · 836b7e55

Bartłomiej Kocot authored Dec 06, 2023

* Introduce wrapper library

* Update cmake files

* Revert "Update cmake files"

This reverts commit c27f88b56590c11a88e26d5d0df7aca51a08133d.

* Fix comments

836b7e55

05 Dec, 2023 4 commits
- Standardize documentation for ReadtheDocs (#1057) · f60cd9d7
  Sam Wu authored Dec 05, 2023
```
Relates to https://github.com/RadeonOpenCompute/rocm-docs-core/issues/330
```
  f60cd9d7
- Merge branch 'develop' into amd-develop · df467969
  Jun Liu authored Dec 04, 2023
  
  df467969
- [SWDEV-435347] disable instances failed with mainlien compiler (#1077) · ff24b537
  Jun Liu authored Dec 04, 2023
  
  ff24b537
- Add daily run with mainline compiler. (#1075) · afe46220
  Illia Silin authored Dec 04, 2023
```
* add daily build with mainline compiler

* fix the compiler paths for ci

* remove the -flto flag

* build with clang by default
```
  afe46220
03 Dec, 2023 1 commit

Add support for double buffering in direct load GEMM kernel (#1052) · bc4bf9bd

Bartlomiej Wroblewski authored Dec 03, 2023

This PR introduces support for double buffering in LDS into GEMM kernels that use direct load instructions.

Direct loads now use inline asm instead of intrinsics. Usage of intrinsics results in compiler adding additional waitcnt instructions what breaks possible load/compute overlap in case of double buffering.

Usage of inline asm results in the need to use sched_barrier in order to make sure that compiler cannot incorrectly reschedule instructions since it does not know the data dependencies between global->LDS and LDS->registers.

bc4bf9bd

30 Nov, 2023 3 commits

[CI] Update Jenkinsfile (#1073) · c7d5c772
Jun Liu authored Nov 30, 2023

c7d5c772

Fixed GroupedGemmFixedNK with hipGraph (#1065) · 49df1dc5

zjing14 authored Nov 30, 2023



* fixed examples; add async_mem_set

* add stream to all deviceOp using SetWorkspace

---------
Co-authored-by: Jing Zhang <jizha@amd.com>

49df1dc5

Introduce wrapper for layout (#1054) · 8ff845f2

Bartłomiej Kocot authored Nov 30, 2023

* Introduce wrapper for layout

* Extend functionality

* Fix for getLength

* Comment fixes

* Add comments and remove not needed getters

8ff845f2

29 Nov, 2023 1 commit
- Merge branch 'develop' into amd-develop · e9047ab9
  Jun Liu authored Nov 29, 2023
  
  e9047ab9