Commits · 105bd708c799f625e5d6b750fcfe615b511db86f · gaoqiong / composable_kernel_ROCM

25 Jul, 2024 1 commit

Add rotating buff for gemm_multi_d (#1411) · 105bd708

zjing14 authored Jul 25, 2024



* add rotating_buff for gemm_multi_d

* format

* Update flush_cache.hpp

* Update gtest.cmake

---------
Co-authored-by: Jing Zhang <jizhan@fb.com>
Co-authored-by: Haocong WANG <haocwang@amd.com>

105bd708

24 Jul, 2024 2 commits

Adding more instances of grouped convolution 3d forward for FP8 with... · 4a8a1bef

Andriy Roshchenko authored Jul 24, 2024

Adding more instances of grouped convolution 3d forward for FP8 with ConvScale+Bias element-wise operation. (#1412)

* Add CMakePresets configurations.

* Add binary elementwise ConvScaleAdd and an example.

* Numerical verification of results.

Observed significant irregularities in F8 to F32 type conversions:
```log
ConvScaleAdd: float=145.000000   f8_t=160.000000    e=144.000000
ConvScaleAdd: float=97.000000   f8_t=96.000000    e=104.000000
ConvScaleAdd: float=65.000000   f8_t=64.000000    e=72.000000
```

* Implemented ConvScaleAdd + Example.

* Add ConvScale+Bias Instances

* Add Client Example for ConvScale+Bias

* Fix number of bytes in an example..

* Cleanup.

4a8a1bef

Add support for half_t and bfloat to reduction operations (#1395) · ffabd70a
Bartłomiej Kocot authored Jul 24, 2024
```
* Add support for half_t and bfloat to reduction operations

* Fix bhalf convert

* Next fix bf16
```
ffabd70a

22 Jul, 2024 1 commit
- Revert Support access per groups and filter2x3 in grouped conv fwd (#1382) (#1406) · 5d8c3d81
  Bartłomiej Kocot authored Jul 22, 2024
  
  5d8c3d81
19 Jul, 2024 3 commits

[GEMM] F8 GEMM, performance optimized. (#1384) · 8c90f25b

Haocong WANG authored Jul 19, 2024



* add ab_scale init support

* enabled interwave

* add scale type; update isSupport

* adjust example

* clean

* enable f8 pure gemm rcr ckprofiler

* Add gemm_multiply_multiply instances

* clang format

* Optimize for ScaleBlockMNK=128

* enable abscale f8 gemm ck profiler

* Add pure f8 gemm test suite

* Reverting to the state of project at f60fd77

* update copyright

* clang format

* update copyright

---------
Co-authored-by: root <jizhan@amd.com>

8c90f25b

Universal gemm splitk using reduce (with multi-d) (#1341) · c544eb4d

ltqin authored Jul 19, 2024



* init for reduce_threadwise multi_d

* add reduce_threadwise_multi_d

* add reduce_multi_d

* clean

* start add an other splitk device op

* add reduce template parameter to SplitKBatchOffset

* add reduce c matrix

* clean up code

* change example data type to bf16

* add bf16Ai8B example

* remove reduce template parameter

* add splitk atomic status to v4

* example add multi d parameters

* device op add multi-d parameters

* add multi-d to reduce

* fix kbach=1 bug

* change B layout to col in  bf16Ai8B example

* remove float adding struct

* change  multi-d interface

* change file and class name

* remove multi-d of bf16Ai8B example

* change IsReduce function to IsReduceAdd

* change example layout to RRR from RCR

* according layout to set ds stride

* reset parameter layout

* add gemm universal reduce instance

* add reduce factory

* add profile_gemm_universal_reduce

* add reduce to profiler

* fix reduce instance

* fix profiler reduce compiling bug

* format

* format library instance code

* add mem instance for reduce library

* fix call instance names

* add workspace for reduce in ckProfiler

* format

* add mnpading to reduce library instance

* add fp16 instance to reduce of profiler

* change copyright time

* restore profiler cmake file

* add reduce text to instances

* add DsLayout and DsDataType to instances template parameter

* fixed gemm_reduce_multi_d

* add an example without multi_d

* Update common.hpp

* Update gtest.cmake

* Update gemm_xdl_splitk_reduce_bf16.cpp

* clean

* Update gtest.cmake

* format

* fixe api

* format

* default parameter change to RRR

* add vector_len for multi_d

* format

* Update gtest.cmake

* fix bf16A iBB elementwiseop

* add ReduceDataType

* move ReduceDataType to end position

* format

* remove googletest git method  address

* fix copyright time

* update init data

---------
Co-authored-by: root <jizhan@amd.com>
Co-authored-by: letaoqin <letaoqin@amd.com>
Co-authored-by: Jing Zhang <jizhan@meta.com>
Co-authored-by: zjing14 <zhangjing14@gmail.com>

c544eb4d

Refactor transform conv to gemm fwd (#1391) · 70a814f1
Bartłomiej Kocot authored Jul 19, 2024
```
* Refactor transform conv to gemm fwd

* fixes codegen

* wmma fixes

* fix wmma

* Fix copyright
```
70a814f1

17 Jul, 2024 1 commit
- Replace the using of __expf by __ocml_exp_f32 to work-around the test_softmax_rank4 failure (#1394) · ee768148
  Qianfeng authored Jul 18, 2024
  
  ee768148
16 Jul, 2024 1 commit

Adding more instances of grouped convolution 3d forward for FP8 with ConvScale... · 802a8a1d

Andriy Roshchenko authored Jul 16, 2024

Adding more instances of grouped convolution 3d forward for FP8 with ConvScale element-wise operation and ReLU activation. (#1386)

* Add CMakePresets configurations.

* Add ConvScale+ReLU Functor and an Example

* Account for ReLU FLOPs.

* Add instances of 3D convolutions with ConvscaleRelu operation.

* Implement Client Example

* Cleanup

802a8a1d

12 Jul, 2024 1 commit
- Support access per groups and filter3x3 in grouped conv fwd (#1382) · 82e8a78a
  Bartłomiej Kocot authored Jul 12, 2024
```
* Support access per groups and filter3x3 in grouped conv fwd

* Fixes for large cases

* Fixes for large tensors
```
  82e8a78a
06 Jul, 2024 1 commit

Universal streamk with atomics (#1360) · 75e622f0

Harisankar Sadasivan authored Jul 05, 2024

* universal streamk with atomics with ckprofiler support. grid_size and streamk strategy are tunable. grid_size of -1 leads to #WGs = maximum occupancy X num_CUs. implementation supports many different streamk policies: 1-tile, 2-tile, 3-tile and 4-tile. streamk strategy of -1 leads to default streamk policy (4-tile). 

* Update README.md

* fixing clang-format issues

* removed conflicts in struct members between streamk and universal streamk

* corrected arg parsing for streamk and universal streamk

* added stream-k policies for 3 tile and 4 tile

* fixed argument type issue with parsing cmd args

* changes suggested in PR review are made- removing comments and correcting copyright

* file permissions updated

* added default value support for grid_size and streamk-policy selection set to -1

* print messages for arguments

* print messages for arguments

* print messages for arguments1

75e622f0

04 Jul, 2024 2 commits
- Add structural sparsity xdlops (#1363) · eaa870a1
  jakpiase authored Jul 04, 2024
```
* Implemented smfmac xdlops

* add reviewer comments
```
  eaa870a1
- Fix issue with multiple targets and remove smfmac tests from unsupported test targets (#1372) · 95907384
  Jun Liu authored Jul 03, 2024
  
  95907384
27 Jun, 2024 2 commits
- Add structural sparsity gemm instruction tests (#1309) · ed21948b
  jakpiase authored Jun 27, 2024
```
* first version of smfmac test

* add reviewer comments

* add reviewer suggestions
```
  ed21948b
- Merging the gfx12 code into public repo. (#1362) · 941d1f7c
  Illia Silin authored Jun 27, 2024
  
  941d1f7c
25 Jun, 2024 1 commit

CK Instance Gen (#1145) · 3e9711f0

arai713 authored Jun 25, 2024



* Format

* Format

* Format

* Remove const

* Use the right template

* Format

* Format

* add row/col instances

* Add missing file

* fixed

* fixing block to etile error

* Format

* Updates

* Format

* fixed rrr layout

* generating a sample JSON file: currently contains includes, prologue/epilogue and instances

* version where the json is passed into the instances to generate a key

* updated run function to just launch kernel

* updated run function: only contains kernel object, json file is updated but still needs to be cleaned up, added front-end API to parse JSON into character buffer

* adding in testing files

* cleaned up comments, still need to work on including header files

* removed unneeded files

* removed/commented out JSON implementation

* added fusion(prologue/epilogue) into instance generation

* working on instance selection

* added instance selection, need to fix instance validation

* removed block2etile map validity check for testing purposes

* test running: failing due to incorrect files/input

* all grid descs/ptrs completed, but device file not found

* Update test and embed modules

* Restore older version

* added convolution operation, written test, debugging generated code for compilation

* attempting to include CK in host directory: _Float16 error

* CK header file issues

* slight fix

* don't crash when hip can't report total memory

* dump generated code to a file

* changing sizes

* creating tensor descriptors using CK methods: set up grid desc manually, also trying to set up an argument pointer - this needs to be fixed

* some fixes to call the device code

* separating test files for conv and gemm

* completed arg ptr, now have linking errors

* clang format fix

* resolved linker issues in conv test

* remove dependency on libutility from ck

* resolved num dim error

* properly passing arg ptr, errors with passing typenames: redefinition/redeclaration

* undo the commenting of device function

* hand created kernel code to find rtc issues

* dump the full src to file

* resolved redeclaration errors, cleaned up errors for Amber's kernel code

* debugging purposes: redeclaration error

* config files

* resolved errors for NumTensor and redeclaration, formatted version.h

* resolved most errors in manually added kernel and my own. error with calling kernel object: overloaded function type

* WIP: close to getting kernel compiled

* WIP: fixing rtc errors

* fixed sequence errors, formatting, still one error with run fcn

* yay: kernel compiles and runs

* updated templated/generated version to run and compile

* minor fixes

* working generated example, resolved memory access error due to padding

* adding in reference kernel, validation failing against reference

* debugging: printing kernel argsz

* reduced error in results

* debugged reference kernel and output errors, added to generated version, currently debugging prologue function issues

* working validation (using reference convolution) with prologue function for both hard-coded and generated version

* WIP: create an alt version that creates Argument on the device

* wip: added new duplicate files, fixed fusion templating errors from working example, setting up kernel arguments

* wip: making necessary methods device code

* added grid descs, working on grid pointers, errors with stl numerics

* wip: updating kernel args - issue, replacing some std functions

* replaced std::accumulate call with temp hardcoded version

* wip: args causing memory issue

* Construct Argument object inside the kernel and use it to call convolution device function. Code runs and verification passes

* adding object file dump

* temporary hardcoding of grid size, can remove device op inst + arg ptr

* minor fix for grid size

* added modified example where arg ptr is created on the device for generated version as well

* removed device op instance and arg ptr from modified examples

* moving device op file for testing purposes and to properly build CK

* commenting out print-outs

* adjust compiler args to produce a valid ELF file

* temporary removal of validation

* reverting compiler args back for working example

* retrieve necessary arguments from generated template parameters in correct format

* calculating grid size on host-side, still need to clean up process, pass parameters to host functions properly

* scaled up factory functions/wrapper structs to implement host-side launch parameter calculations using CK host side functions - in hard-coded example

* temporary change to generate ELF format binary object file

* removed unecessary code, added comments

* formatting fix

* cleaned up code, added new tests, restructured library: move helper into CK

* refactored launch parameter calculation to be more concise

* renamed files and variables for more clarity/uniformity

* more code cleaning, removed debug statements

* moved majority of my files into codegen directory, running properly

* updated Embed.cmake(string_view) in codegen directory

* updated host directory to match Embed.cmake as well

* added old tests in

* updated instance generation methods to be more concise

* removed layout from launch parameter calculation

* working test

* fixed issue with verification, all instances working

* updated verification in other tests

* removed duplicate matrix padder file, removed code dumps

* removed old hard-coded tests

* removed old host directory, all files in codegen directory now

* fixed copyright in files

* commenting out validation

* renamed files

* made changes for review: fixed copyright, renamed files for clarity, removed comments, refactored code

* updated headers

* removing duplicate file for fwd conv to gemm, merging with original file

* fix building codegen with clang++ directly

* resolving build error from conv_fwd_to_gemm

* fix for previous error

* renaming tests

* created common test file

* cleaned up code, added comments

* renamed device op

* fixed typos in comments

* removed extra space

* code cleanup: resolving Amber's comments

* removed wrapper struct for matrix padder, fixed template

* cleaned up if statements for better readability

---------
Co-authored-by: Paul <pfultz2@yahoo.com>
Co-authored-by: Jing Zhang <jizha@amd.com>
Co-authored-by: M. Amber Hassaan <amber_474@yahoo.com>
Co-authored-by: illsilin <Illia.Silin@amd.com>
Co-authored-by: Illia Silin <98187287+illsilin@users.noreply.github.com>

3e9711f0

21 Jun, 2024 2 commits

WA for rocm-6.2+ s constrait for buffer resource (#1346) · fa129c1a
carlushuang authored Jun 22, 2024
```
* WA for rocm-6.2+ s constrait for buffer resource

* add missing memory clobber
```
fa129c1a

Fix cmake warnings (#1342) · 510325a4

Bartłomiej Kocot authored Jun 21, 2024

* Cmake add -Wno-nvcc-compt

* Remove template without initialization list

* dpp remove template without init list

* Fixes

510325a4

20 Jun, 2024 1 commit

Adding Missed Activation Functions for Grouped 2D/3D Convolutions (#1348) · 0162a5f6

ThruptiRajLakshmanaGowda authored Jun 20, 2024

* Initial Push

* First Push

* Fixed Clang format

* Resolve merge conflict

* Addressed review comments

* Addressed review comments

* Addressed review comments

0162a5f6

18 Jun, 2024 3 commits
- Add read_first_lane function for int64 (#1347) · 8faec23c
  Bartłomiej Kocot authored Jun 18, 2024
  
  8faec23c
- Switch to universal gemm in grouped gemm tile loop (#1335) · e2d13920
  jakpiase authored Jun 18, 2024
```
* switch to universal gemm in grouped gemm tile loop

* minor fixes

* add reviewers comments

---------
Co-authored-by: Adam Osewski <19374865+aosewski@users.noreply.github.com>
```
  e2d13920
- Fix continous dim selection in contraction (#1336) · 933951ed
  Bartłomiej Kocot authored Jun 18, 2024
```
* Fix continous dim selection in contraction

* Fixes
```
  933951ed
17 Jun, 2024 1 commit
- disabled lds direct load inline asm (#1331) · e0210316
  zjing14 authored Jun 16, 2024
  
  e0210316
14 Jun, 2024 1 commit
- Support large tensors in grouped conv fwd (#1332) · dc1e9c5d
  Bartłomiej Kocot authored Jun 14, 2024
```
* Support large tensors in grouped conv fwd

* Multi ABD fixes

* Fix calculate element space size
```
  dc1e9c5d
10 Jun, 2024 1 commit

Add a convinvscale op, related instances and examples (#1307) · ce66277a

Rostyslav Geyyer authored Jun 10, 2024



* Update the element op

* Add an example

* Add instances

* Add a client example

* make sure new instances only build on gfx9

* Update element op and its handling

* Format

* Update instances to take element op as an argument

* Update examples to use random scale values

* Format

* Update client example with random scales

* Format

---------
Co-authored-by: illsilin <Illia.Silin@amd.com>

ce66277a

05 Jun, 2024 2 commits

Integrate universal gemm with conv forward (#1320) · ac58cc5d

Bartłomiej Kocot authored Jun 05, 2024

* Integrate universal gemm with conv fwd

* Fix conv fwd wmma test

* Fix instances

* Remove direct load check

ac58cc5d

Add a scale op, related instances and examples (#1242) · cb0645be

Rostyslav Geyyer authored Jun 04, 2024



* Add a scale op

* Update the element op

* Add instances

* Add an example

* Add a client example

* Add a flag check

* Revert flag check addition

* Fix flag check

* Update d strides in example

* Update d strides in client example

* Apply suggestions from code review

Update copyright header
Co-authored-by: Bartłomiej Kocot <barkocot@amd.com>

* Move the example

* Move the client example

* Update element op

* Update example with the new element op

* Add scalar layout

* Update example

* Update kernel for scalar Ds

* Revert kernel changes

* Update element op

* Update example to use scales' pointers

* Format

* Update instances

* Update client example

* Move element op to unary elements

* Update element op to work with values instead of pointers

* Update instances to take element op as an argument

* Update examples to use random scale values

---------
Co-authored-by: Bartłomiej Kocot <barkocot@amd.com>

cb0645be

01 Jun, 2024 1 commit

Post-merge fix of PR 1300 (#1313) · 6fb1f4e0

zjing14 authored Jun 01, 2024

* add f8 gemm with multiD for both row/col wise

* change compute_type to fp8

* changed tuning parameters in the example

* add rcr example

* post-merge fix

* fix

* reduce init range

6fb1f4e0

28 May, 2024 1 commit

add f8 gemm multiD with both row/col wise scale (#1300) · 80db62f0

zjing14 authored May 28, 2024

* add f8 gemm with multiD for both row/col wise

* change compute_type to fp8

* changed tuning parameters in the example

* add rcr example

80db62f0

22 May, 2024 1 commit
- Optimize grouped conv bwd weight for small M and N (#1303) · fd72380a
  Bartłomiej Kocot authored May 22, 2024
```
* Optimize grouped conv bwd weight for small M and N

* Fixes
```
  fd72380a
20 May, 2024 1 commit
- aggregate device macros in ck_tile config header (#1297) · 06b891c5
  Illia Silin authored May 20, 2024
  
  06b891c5
17 May, 2024 1 commit
- replace the ENV macro with CK_ENV (#1296) · 1274861a
  Illia Silin authored May 17, 2024
  
  1274861a
15 May, 2024 2 commits
- remove wrong use of nonexistent class members (#1290) · c4413783
  Illia Silin authored May 15, 2024
  
  c4413783
- Add unit tests for grouped gemm two stage (#1256) · 3e3471d5
  jakpiase authored May 15, 2024
```
* add unit tests for grouped gemm two stage

* add reviewers suggestions

---------
Co-authored-by: Adam Osewski <19374865+aosewski@users.noreply.github.com>
```
  3e3471d5
10 May, 2024 2 commits
- Code clean-up (#1285) · 566b6480
  Illia Silin authored May 10, 2024
```
* code clean-up

* remove the profiling output samples
```
  566b6480
- Change output gemm type to AccDataType in two stage conv bwd wei (#1283) · 8346af9c
  Bartłomiej Kocot authored May 10, 2024
  
  8346af9c
09 May, 2024 2 commits
- Fix MakeArgument (#1284) · a0ae1c61
  Adam Osewski authored May 09, 2024
  
  a0ae1c61
- Add vector instruction coherency bits for gfx94 targets. (#1268) · 3c043cd1
  Adam Osewski authored May 09, 2024
  
  3c043cd1
08 May, 2024 2 commits
- fix the output formatting (#1282) · fdbf8ccb
  Illia Silin authored May 08, 2024
  
  fdbf8ccb
- Add two stage grouped conv bwd weight kernel (#1280) · 0b6b5d17
  Bartłomiej Kocot authored May 08, 2024
  
  0b6b5d17