Commits · dcb013fcf22a5375770330416c2bca35141c4f84 · gaoqiong / composable_kernel_ROCM

08 Nov, 2023 1 commit
- Avoid force setting ENABLE_PIPELINE_V2_OPT to OFF (#961) · dcb013fc
  Po Yen Chen authored Oct 19, 2023
```
* Avoid force setting ENABLE_PIPELINE_V2_OPT to OFF

* Remove compilation option variable MAX_ILP_OPTS
```
  dcb013fc
11 Oct, 2023 2 commits
- Merge branch 'amd-develop' into amd-master · 50320413
  Jun Liu authored Oct 11, 2023
  
  50320413
- Merge commit 'ac9595a9 ' into amd-develop · 91b414cd
  Jun Liu authored Oct 11, 2023
  
  91b414cd
10 Oct, 2023 1 commit

zjing14 authored Oct 10, 2023



* workaround nan problem by changing output to fp16

* enable f8/bf8 gemm tests on MI200

* workaround f16 to f8 conversion

---------
Co-authored-by: Jing Zhang <jizha@amd.com>

ac9595a9

05 Oct, 2023 5 commits
- Merge branch 'amd-develop' into amd-master · 0b70e1cd
  Jun Liu authored Oct 05, 2023
  
  0b70e1cd
- Merge branch 'develop' into amd-develop · 082cf643
  Jun Liu authored Oct 05, 2023
  
  082cf643
- Replace CMake `return` from later CMake (#970) · 59136091
  Lauren Wrubleski authored Oct 05, 2023
  
  59136091
- Revert "Add support for mixed precision in contraction scale and bilinear" (#967) · 4daedf8c
  Illia Silin authored Oct 05, 2023
```
* Revert "Add support for mixed precision in contraction scale and bilinear (#936)"

This reverts commit f0748506.

* revert commits #957 and #960
```
  4daedf8c
- remove example 60 (#963) · 570ff3dd
  zjing14 authored Oct 05, 2023
```
Co-authored-by: Jing Zhang <jizha@amd.com>
```
  570ff3dd
04 Oct, 2023 3 commits

Grouped conv bwd data with fp16 input and bf8fp8 comp (#962) · 04f93aad

zjing14 authored Oct 04, 2023



* Add f8 bf8 gemm example

* Add element-wise ops

* Add intrinsics

* Update reference calculation

* Add an additional type option for xdlops gemm

* Fix build process

* Add bf8 to buffer addressing

* Update blockwise op, split typeA and typeB

* Update for compatibility

* Uppdate naming to f8->fp8

* Update naming

* Format

* Update naming (#937)

* Add a client example

* Add computetypes to device and gridwise ops

* Add instances, update instance factory

* Format

* Fix a flag

* Add ckProfiler mode

* Fix typos

* Add an example

* Add bf8 generator

* add bf8 mfma; fixed type_convert for bf8

* move verfication ahead of timing

* Update reference calculation

* Fix reference

* Narrow down float init range

* Fix bf8 bf8 mfma

* Add bf8 @ fp8 mfma

* Update example

* Update instances

* Update profiler api

* Update for compatibility

* Format

* Remove extra example

* Clean up

* workaround convert

* added instance of f16_bf8f8, and client example

* fixed mfma selector

* format

---------
Co-authored-by: Rostyslav Geyyer <rosty.geyyer@amd.com>
Co-authored-by: Rostyslav Geyyer <46627076+geyyer@users.noreply.github.com>
Co-authored-by: Jing Zhang <jizha@amd.com>

04f93aad

Add conv bwd weight fp16 comp bf8 fp8 op, instances and example (#945) · 42facfc6

Rostyslav Geyyer authored Oct 04, 2023



* Add f8 bf8 gemm example

* Add element-wise ops

* Add intrinsics

* Update reference calculation

* Add an additional type option for xdlops gemm

* Fix build process

* Add bf8 to buffer addressing

* Update blockwise op, split typeA and typeB

* Update for compatibility

* Uppdate naming to f8->fp8

* Update naming

* Format

* Update naming (#937)

* Add a client example

* Add computetypes to device and gridwise ops

* Add instances, update instance factory

* Format

* Fix a flag

* Add ckProfiler mode

* Fix typos

* Add an example

* Add bf8 generator

* add bf8 mfma; fixed type_convert for bf8

* move verfication ahead of timing

* Update reference calculation

* Fix reference

* Narrow down float init range

* Fix bf8 bf8 mfma

* Add bf8 @ fp8 mfma

* Update example

* Update instances

* Update profiler api

* Update for compatibility

* Format

* Remove extra example

* Clean up

* workaround convert

---------
Co-authored-by: Jing Zhang <jizha@amd.com>

42facfc6

3d grouped conv fwd with input/output fp16 and comp fp8 (#931) · e921e1f0

zjing14 authored Oct 03, 2023



* add f8 comp instance

* fixed

* fixed comments

* rename

* fixed dtype

* format

* fixed CI

* fixed ci

* add missing ComputeType

* fixed cit

* fixed

* Update cmake-ck-dev.sh

---------
Co-authored-by: Jing Zhang <jizha@amd.com>

e921e1f0

03 Oct, 2023 5 commits
- changed test for grouped_gemm to be random (#959) · 5311d1b3
  zjing14 authored Oct 03, 2023
```
Co-authored-by: Jing Zhang <jizha@amd.com>
```
  5311d1b3
- Fixed contraction issues (#960) · aa46039f
  zjing14 authored Oct 03, 2023
```
* add missing ComputeType

* fixed

* Update cmake-ck-dev.sh

---------
Co-authored-by: Jing Zhang <jizha@amd.com>
```
  aa46039f
- add generic instances (#947) · f477fca4
  zjing14 authored Oct 03, 2023
```
Co-authored-by: Jing Zhang <jizha@amd.com>
```
  f477fca4
- Merge branch 'amd-develop' into amd-master · 7b7a3978
  Jun Liu authored Oct 02, 2023
  
  7b7a3978
- Merge branch 'develop' into amd-develop · 7e8230da
  Jun Liu authored Oct 02, 2023
  
  7e8230da
02 Oct, 2023 3 commits

Add fp8 @ bf8 gemm support and example (#933) · bd09b5c5

Rostyslav Geyyer authored Oct 02, 2023

* Add f8 bf8 gemm example

* Add element-wise ops

* Add intrinsics

* Update reference calculation

* Add an additional type option for xdlops gemm

* Fix build process

* Add bf8 to buffer addressing

* Update blockwise op, split typeA and typeB

* Update for compatibility

* Uppdate naming to f8->fp8

* Update naming

* Format

bd09b5c5

get rid of gfx900/906, set rocm5.7 as default (#958) · 59dbb01f
Illia Silin authored Oct 02, 2023

59dbb01f

Contraction multi abd (#957) · 9d58c421

zjing14 authored Oct 02, 2023



* add gridwise_multi_abd

* move element_op into RunRead

* merge element_wise op with data read

* add multiABD example

* allow packed elementwise_op

* changed example

* clean

* clean

* add is_detected

* fix

* minor fix

* add scaleAdd_vec4 example

* init commit for contraction_multi_ABD

* add examples

* add examples of multiA and broadcast

* update example

* fixed comments

* Update cmake-ck-dev.sh

* Update cmake-ck-dev.sh

* Add comments into the example

---------
Co-authored-by: Jing Zhang <jizha@amd.com>

9d58c421

29 Sep, 2023 2 commits

add gfx942 target to the daily ckprofiler package (#955) · 6b5f6473
Illia Silin authored Sep 29, 2023

6b5f6473

Add support for mixed precision in contraction scale and bilinear (#936) · f0748506

Bartlomiej Wroblewski authored Sep 29, 2023

* Extract common functionality to separate files

* Reference contraction: Remove incorrect consts from type_converts

* Reference contraction: Add missing type_convert for dst value

* Reference contraction: Fix incorrect order of B matrix dimensions

* Add support for mixed precision in contraction scale and bilinear

* Move using statements from instances to a common file

* Move using statements from examples to a common file

* Fix the order of B matrix dimensions across examples and profiler

* Fix the computation of error threshold

* Make ComputeDataType an optional argument

* Include possible DataType -> ComputeDataType casting error in the threshold

* Remove commented code

f0748506

28 Sep, 2023 4 commits

Add grouped conv bwd data wmma (#950) · cb538740

Bartłomiej Kocot authored Sep 28, 2023

* Add grouped conv bwd data wmma

* Fix copyrights

* Add instances with smaller NPerBlock

* Update interface test

* Minor stylistic fixes

* Minor stylistic fixes

cb538740

Add grouped convolution changes to changelog (#952) · 271ef645

Bartłomiej Kocot authored Sep 28, 2023



* Add grouped convolution changes to changelog

* Fix 0.2.0 ck release rocm version

* Suggested CHANGELOG.md edits

* Update CHANGELOG.md

* Update CHANGELOG.md

* Update CHANGELOG.md

* Update CHANGELOG.md

* Update CHANGELOG.md

* Update CHANGELOG.md

---------
Co-authored-by: Lisa <lisajdelaney@gmail.com>

271ef645

Merge branch 'amd-develop' into amd-master · b24d93a1
Jun Liu authored Sep 28, 2023

b24d93a1
Merge branch 'develop' into amd-develop · 56c72035
Jun Liu authored Sep 28, 2023

56c72035

27 Sep, 2023 5 commits

Fix gemm_splitk test, add hip_check_error after kernel calls in kernel_launch. (#951) · bc1108bb

Illia Silin authored Sep 27, 2023



* Added error check after kernel launch (#919)
Co-authored-by: Xiaodong Wang <xdwang@meta.com>
Co-authored-by: Xiaodong Wang <xw285@cornell.edu>

* remove M=0 test cases for test_gemm_splitk

---------
Co-authored-by: Xiaodong Wang <xdwang@meta.com>
Co-authored-by: Xiaodong Wang <xw285@cornell.edu>

bc1108bb

Handle type conversions to a const datatype (#944) · f4af5aed

Bartlomiej Wroblewski authored Sep 27, 2023

* Handle type conversions to a const datatype

* Review: Handle X being const data type as well

* Review: Remove typo

f4af5aed

Add column to image kernel (#930) · e2243a4d

Bartłomiej Kocot authored Sep 27, 2023

* Add column to image kernel

* Minor fixes for dtypes and client examples

* Disable tests for disabled dtypes

* Disable add instances functions for disabled data types

* Minor stylistic fixes

* Revert "Disable add instances functions for disabled data types"

This reverts commit 728b86956378dcd9415fd0f2557833a068fe1c10.

* Instances reduction

* Add comments in device_column_to_image_impl

* Update changelog and Copyrights

* Improve changelog

e2243a4d

Add multiple A/B support (#906) · 11676c7e

zjing14 authored Sep 26, 2023



* add gridwise_multi_abd

* move element_op into RunRead

* merge element_wise op with data read

* add multiABD example

* allow packed elementwise_op

* changed example

* clean

* clean

* add is_detected

* fix

* minor fix

* add scaleAdd_vec4 example

---------
Co-authored-by: Jing Zhang <jizha@amd.com>

11676c7e

Use lower case for ckprofiler package. (#948) · 420b5a03
Illia Silin authored Sep 26, 2023
```
* split ckProfiler gfx9 package into gfx90 and gfx94

* use lower case for package names
```
420b5a03

26 Sep, 2023 6 commits
- Fixed Gemmv2r3 kpad (#938) · 48ba6e8a
  zjing14 authored Sep 26, 2023
```
* added kpad support into v2r3

* add generic instances

* fixed comments

* fixed mnk padding

* Update device_batched_gemm_xdl.hpp

* fixed kpad

---------
Co-authored-by: Jing Zhang <jizha@amd.com>
```
  48ba6e8a
- Add fp8 gemm instances (#920) · 94bfa502
  Rostyslav Geyyer authored Sep 26, 2023
```
* Add fp8 gemm instances

* Update instance naming
```
  94bfa502
- Merge branch 'amd-develop' into amd-master · 742dd3aa
  Jun Liu authored Sep 26, 2023
  
  742dd3aa
- Merge branch 'develop' into amd-develop · 1f02eaef
  Jun Liu authored Sep 26, 2023
  
  1f02eaef
- split ckProfiler gfx9 package into gfx90 and gfx94 (#946) · 0b296a27
  Illia Silin authored Sep 26, 2023
  
  0b296a27
- Resolve some data type issues and cmake policy. (#940) · 2ea75bd6
  Illia Silin authored Sep 26, 2023
```
* split the types in gemm_bilinear instances, add condition to cmake policy

* fix syntax

* split the data types in batchnorm examples

* fix the batchnorm_bwd test

* fix types in the batchnorm_bwd test
```
  2ea75bd6
25 Sep, 2023 1 commit
- Merge branch 'amd-develop' into amd-master · c9013009
  Jun Liu authored Sep 25, 2023
  
  c9013009
24 Sep, 2023 1 commit
- Merge branch 'develop' into amd-develop · 84dcf5d0
  Jun Liu authored Sep 23, 2023
  
  84dcf5d0
23 Sep, 2023 1 commit

Add 3d grouped conv fwd wmma instances (#935) · c9553832

Bartłomiej Kocot authored Sep 23, 2023

* Add 3d grouped conv fwd wmma instances

* Refactor fwd conv tests

* Split wmma instances for each specialization

* Minor stylistic fixes

c9553832