Commits · fcc81cf227dbf49eb81d5f31675d78893f953797 · gaoqiong / composable_kernel

12 Dec, 2022 6 commits
- Remove no-longer used commands · fcc81cf2
  Po-Yen, Chen authored Dec 13, 2022
  
  fcc81cf2
- Update 16-MFMA scheduling pattern · e39a4794
  Po-Yen, Chen authored Dec 12, 2022
  
  e39a4794
- Remove unused opt-strategies · 686a0c20
  Po-Yen, Chen authored Dec 12, 2022
  
  686a0c20
- Use v2 pipeline for all instances · 52cff1b8
  Po-Yen, Chen authored Dec 12, 2022
  
  52cff1b8
- Update sheduling pattern for IGLP_OPT_STRATEGY=1 · fcef1550
  Po-Yen, Chen authored Dec 12, 2022
  
  fcef1550
- Comment out unused instance definition · 08d12d4f
  Po-Yen, Chen authored Dec 12, 2022
  
  08d12d4f
09 Dec, 2022 5 commits
- Allow using macro to switch between opt strategies · 3a457e27
  Po-Yen, Chen authored Dec 09, 2022
  
  3a457e27
- Use macro to decide mfma cluster size · a9255313
  Po-Yen, Chen authored Dec 09, 2022
  
  a9255313
- Use pipeline v2 for all layout=NN instances · 3fea2678
  Po-Yen, Chen authored Dec 09, 2022
  
  3fea2678
- Make CMake commands more extensible · 937e917b
  Po-Yen, Chen authored Dec 09, 2022
  
  937e917b
- Use macro to toggle pipeline v2 optimization · 4bf3b87e
  Po-Yen, Chen authored Dec 09, 2022
  
  4bf3b87e
08 Dec, 2022 6 commits
- Use better scheduling pattern · 2a6d05f4
  Po-Yen, Chen authored Dec 09, 2022
  
  2a6d05f4
- Default enable intrinsic function calls · ea724623
  Po-Yen, Chen authored Dec 09, 2022
  
  ea724623
- Remove dependencies on other layouts · 1e1e8fd9
  Po-Yen, Chen authored Dec 09, 2022
  
  1e1e8fd9
- Disable all instances except NN one · 2b7e1b40
  Po-Yen, Chen authored Dec 09, 2022
  
  2b7e1b40
- Use pipeline v2 instances · a095f4b1
  Po-Yen, Chen authored Dec 09, 2022
  
  a095f4b1
- Use intrinsic functions to improve pipeline v2 · 43df1a8b
  Po-Yen, Chen authored Dec 09, 2022
  
  43df1a8b
06 Dec, 2022 1 commit

Illia Silin authored Dec 06, 2022

* ignore .git folder when doing clang-format

* fix syntax

* add backslashes before quotes

* add path filter for several extensions

d072790f

02 Dec, 2022 3 commits

Fix bug where scaling may not be applied in some code path (#526) · d1567094
Anthony Chang authored Dec 03, 2022
```
* fix bug where scaling may not be applied in some code path

* more test

* revert accidental example code changes
```
d1567094

Add multiple d gridwise gemm on Navi21 for ResNet50 (#517) · 23ecf0fa

ltqin authored Dec 03, 2022



* start add example

* add multiple d fp16 example

* device transfer elementwiseop to gridwise

* gridwise add multiple d

* change example for multiple d

* fix spill registers

* fix for passthrough element op

* fix int8 overflow

* change example file name

* add instance for dl multiple d

* example add DsDataType

* remove grouped_convolution_forward_dl.hpp

* add head file(was deleted before)

* fix not support device issue

* format

* remove passthrough check
Co-authored-by: letaoqin <letaoqin@amd.com>

23ecf0fa

[Navi3x-LWPCK-449] wmma_op + unit test (#484) · abf9cc6c

Haocong WANG authored Dec 03, 2022



* wmma_op + unit test

* add arch limitation to wmma test

* change arch limitation

* Refactor + Add all type unit test(int4 compile failed)

* Add f32_16x16x16_bf16 unit test

* Remote int4 related

* delete deprecated test
Co-authored-by: Po Yen Chen <PoYen.Chen@amd.com>
Co-authored-by: Chao Liu <chao.liu2@amd.com>

abf9cc6c

01 Dec, 2022 1 commit

Modularize ckProfiler operations (#514) · 8784a72e

Po Yen Chen authored Dec 02, 2022



* Re-structure ckProfiler source files

* Rename profiler.cpp to main.cpp

* Modularize ckProfiler operations

* Add description for profiler operations

* Use longer name to avoid name collision

* Use macro to delay expansion

* Use std::move() to avoid object copying

* Prohibit users from calling dtor

* Use macro to eliminate redundant code

* Make friend function hidden

* Add missing include directive <iostream>

* Fix wrong include directives

* Remove int8 from batchnorm-forward instances since it is not needed for forward training and could fail test
Co-authored-by: Qianfeng Zhang <Qianfeng.Zhang@amd.com>

8784a72e

30 Nov, 2022 2 commits

gemm, conv perchannel quantization (#503) · ad541ad6

rocking5566 authored Dec 01, 2022

* Use gemm_multiple_D instead

* Add gemm bias relu quantization example

* Add pure gemm quantization example

* Add quantization of perchannel conv + bias + relu example

* Refine the code

* Rename multiplier to requant_scale

* Rename the folder

* Remove redundant comment

* Rename the file. Prepare to add perchannel

* Add conv perchannel instance

* Move to quantization folder

* Add conv perchannel client example

* Apply Rangify constructor of HostTensorDescriptor & Tensor<>

* Fix merge error

ad541ad6

BatchNorm backward instance/external API/profiler/tests (#519) · 63af525c

Qianfeng authored Dec 01, 2022

* Refine the device batchnorm-backward base API templates and data type assignments

* Remove duplicated kernel file

* Add batchnorm backward instances and external API

* Add batchnorm-backward profiler and tests

* Add client example which uses batchnorm backward external API

* Merge test/batchnorm_fwd and test/batchnorm_bwd into one directory

* Loose the threshold for batchnorm-backward check_err()

63af525c

29 Nov, 2022 3 commits

Fix split-k gemm test (#231) · 236bd148

Anthony Chang authored Nov 30, 2022



* properly return error flag; reveals bug in split-k gemm

* fix bug in split k

* update split-k test case
Co-authored-by: Chao Liu <chao.liu2@amd.com>

236bd148

fix GetTypeString · 0e9c88ce
fsx950223 authored Nov 16, 2022

0e9c88ce

BatchNorm backward implementation (#461) · 44789d99

Qianfeng authored Nov 29, 2022

* Implemented batchnorm-backward Blockwise and Multiblock kernels

* Add batchnorm-backward device op

* Add batchnorm-backward host-reference op

* Add batchnorm-backward example

* Parameters renaming in batchnorm backward kernels and device op

* Change in the example to loose the threshold for ScaleDiff checking

* Add comments to explain the implementation of batchnorm-backward

* Parameters renaming again in batchnorm backward kernels

* Improve the expression calculation for performance

* Add batchnorm backward to README

* Add comments to explain inv-variance in batchnorm forward and backward

* Renaming the batchnorm forward training and inferring examples

* Add/update the comments for batchnorm-backward kernels

* Renaming again

* Add block_sync_lds between two consecutive blockwise reductions

* Move common expression 1/N out of the static_for loops

* Add dy_elementwise_op

* Renaming in backward example again

* Add checking for reduceDims in reference_batchnorm_backward

* Update to comments and codes format

* Rename in the comments

* Remove common expression out of the loop in reference_batchnorm_backward_nhwc_c

* Add block_sync_lds() between blockwise reduction again

* Fix comments again

* Remove int8 from batchnorm-forward instances since it is not needed for forward training and could fail test

44789d99

28 Nov, 2022 1 commit
- Remove int8 from batchnorm-forward instances since it is not needed for... · 5bf0475a
  Qianfeng authored Nov 29, 2022
```
Remove int8 from batchnorm-forward instances since it is not needed for forward training and could fail test (#516)
```
  5bf0475a
25 Nov, 2022 1 commit

BatchNorm forward instance/external api/profiler/tests/client example (#511) · 4e6a5575

Qianfeng authored Nov 25, 2022



* Update to device_batchnorm_forward base class to include all template parameters for problem description

* Add batchnorm forward instances and external api

* Add batchnorm forward profiler module which uses the external api

* Add some comments in batchnorm_forward example to explain the dimensions in lengths[]

* Replace the reference_batchnorm_forward_nhwc_c by generic reference_batchnorm_forward

* Improvement to the batchnorm infer base API

* Add batchnorm forward client example which shows using the batchnorm forward external API

* Add test for batchnorm forward

* Tuning the batchnorm profiler initialized values and error threshold

* Add support for bhalf_t in instances/external api/tests

* Add support for int8_t in instances/external api/tests

* Add support for double in instances/external api/tests

* Let ScaleDataType and BiasDataType be same as XDataType and YDataType when creating instances

* Checking before running best instance in batchnorm_fwd_nhwc client example

* Add checking for YElementwiseOp in batchnorm_forward external API

* Add more types in batchnorm forward profiler

* Add more test lengths
Co-authored-by: rocking5566 <ChunYu.Lai@amd.com>

4e6a5575

20 Nov, 2022 1 commit

Client examples AddFastGelu and FastGelu + instances. (#509) · 43a889b7

Adam Osewski authored Nov 20, 2022



* FastGelu support for more data types.

* AddFastGelu & FastGelu instances.

* Client example.

* clang-format

* Remove unused stride variable.

* Add new line at EOF.
Co-authored-by: Adam Osewski <aosewski@amd.com>

43a889b7

17 Nov, 2022 1 commit
- Work around develop validation failure (#513) · 892a8d76
  Anthony Chang authored Nov 18, 2022
```
* workaround bf16 atten fwd issue on gfx908

* typo
```
  892a8d76
15 Nov, 2022 4 commits

Add BF16 tests for batched_gemm_softmax_gemm_permute (#504) · 4c4c7328

guangzlu authored Nov 16, 2022



* fixed bug in softmax reference & add bf16 examples for batched_gemm_scale_softmax_gemm

* added bf16 tests for batched_gemm_softmax_gemm_permute

* changed format of device_batched_gemm_softmax_gemm_permute_xdl_cshuffle_bf16_bf16_bf16_bf16_gmk_gnk_gno_gmo_instance.cpp

* changed format device_batched_gemm_softmax_gemm_permute_xdl_cshuffle_bf16_bf16_bf16_bf16_gmk_gnk_gno_gmo_instance.cpp

* aligned annotations

* modified CMakeLists for examples

* add common example code of fp16/bf16 version for batched_gemm_scale_softmax_gemm_xdl

* use macro to control the instances

* added macro control into instances

* clang-format some files

* changed error tolerance for bf16

* changed index for 10_elementwise_normalization

* fixed xdlops code bug in amd_xdlops.hpp
Co-authored-by: Po Yen Chen <PoYen.Chen@amd.com>

4c4c7328

Add Conv Backward Data on Navi21 for ResNet50 (#499) · db0eb1ea

ltqin authored Nov 16, 2022



* start add example

* add device dl

* change launch kernel

* change init data method

* change example config

* add config valid check

* add instance for dl bwd

* add instance to ckProfiler

* reserver to profiler and cmakelist

* add instance to ckProfiler2

* change instance f32 config

* fix example return value
Co-authored-by: letaoqin <letaoqin@amd.com>
Co-authored-by: Po Yen Chen <PoYen.Chen@amd.com>

db0eb1ea

Avoid reporting unused member function error (#507) · 7038723a
Po Yen Chen authored Nov 15, 2022

7038723a

Introduce ck::accumulate_n() (#439) · 730204ee

Po Yen Chen authored Nov 15, 2022

We can use this template to eliminate duplicated iterator computing
logics. By providing return type to ck::accumulate_n(), we can avoid
type conversion operations.

730204ee

14 Nov, 2022 1 commit

Rangify STL algorithms (#438) · dc663fae

Po Yen Chen authored Nov 15, 2022

* Rangify STL algorithms

This commit adapts rangified std::copy(), std::fill() & std::transform()

* Re-write more std::copy() calls

* Re-write std::copy() calls in profiler

dc663fae

11 Nov, 2022 3 commits

Rangify check_err() (#444) · b79bbbc2

Po Yen Chen authored Nov 12, 2022

* Rangify check_err()

By rangifying check_err(), we can not only compare values between
std::vector<>s, but also compare any ranges which have same value
type.

* Re-format example code

b79bbbc2

Fix build errors on CI server (#506) · 4382b414
Po Yen Chen authored Nov 12, 2022
```
* Add missing ignore expression

* Add missing include directive
```
4382b414

Rangify constructor of HostTensorDescriptor & Tensor<> (#445) · 4a2a56c2

Po Yen Chen authored Nov 12, 2022

* Rangify STL algorithms

This commit adapts rangified std::copy(), std::fill() & std::transform()

* Rangify check_err()

By rangifying check_err(), we can not only compare values between
std::vector<>s, but also compare any ranges which have same value
type.

* Allow constructing Tensor<> like a HostTensorDescriptor

* Simplify Tensor<> object construction logics

* Remove more unnecessary 'HostTensorDescriptor' objects

* Re-format example code

* Re-write more HostTensorDescriptor ctor call

4a2a56c2

10 Nov, 2022 1 commit
- Add packages for examples and profiler (#502) · 37f2e918
  Lauren Wrubleski authored Nov 10, 2022
```
* Add packages for example and profiler

* correct TEST_NAME -> EXAMPLE_NAME
```
  37f2e918