Commits · bfefc6b8864fa9df3c21177a2f1c1790bcd01aa9 · gaoqiong / composable_kernel_ROCM

10 Mar, 2024 1 commit
- enabled gemm · bfefc6b8
  Jing Zhang authored Mar 10, 2024
  
  bfefc6b8
01 Mar, 2024 1 commit
- Update clipping for fp8/bf8 conversion (#1182) · acfb3392
  Rostyslav Geyyer authored Mar 01, 2024
```
* Update clipping for fp8 conversion

* Add clipping for bf8 conversion

* Format
```
  acfb3392
27 Feb, 2024 1 commit

Clip fp8 to +/-240 on all targets. (#1172) · d0c7b451

Illia Silin authored Feb 27, 2024

* clip fp8 to +/-240 on all targets

* if inputs to fp8 conversion are +/-inf, they remain unaltered

* increase tolerance for test_elementwise_layernorm to prevent false errors

* change the input values for gemm examples to floats

* reduce gemm example float input values to prevent errors

* increase the tolerance for gemm examples

d0c7b451

26 Feb, 2024 1 commit
- Todo: fix gemm_bilinear_wmma instances compilation bug · 18d5297b
  aska-0096 authored Feb 26, 2024
  
  18d5297b
24 Feb, 2024 1 commit
- add wmma · 760b0c75
  Jing Zhang authored Feb 23, 2024
  
  760b0c75
17 Feb, 2024 1 commit
- fixed block_sync_lds · 8831b0d8
  Jing Zhang authored Feb 16, 2024
  
  8831b0d8
15 Feb, 2024 1 commit
- fix clang format · e60bf36c
  illsilin authored Feb 14, 2024
  
  e60bf36c
14 Feb, 2024 1 commit
- initial enablement of gfx950 · d66da6be
  illsilin authored Feb 14, 2024
  
  d66da6be
07 Feb, 2024 1 commit

Implement direct loads split-K GEMM kernel (#1137) · 69518582

Bartlomiej Wroblewski authored Feb 07, 2024



* WIP: Implement direct loads split-K GEMM kernel

* Clean the review

---------
Co-authored-by: Adam Osewski <19374865+aosewski@users.noreply.github.com>
Co-authored-by: Bartłomiej Kocot <barkocot@amd.com>

69518582

02 Feb, 2024 1 commit

Add support for more Navi2x and Navi3x models. (#1152) · 180f16f9

Illia Silin authored Feb 02, 2024

* add support for navi2x and navi3x models

* fix syntax

* use common macro for different mi300 architectures

180f16f9

24 Jan, 2024 1 commit

Fixing most of the cppcheck errors. (#1142) · 180e5720

Illia Silin authored Jan 24, 2024

* fix cppcheck errors, first pass

* fix format

* fix returned value in examples

* add macro definitions for cppcheck

* fix the profile_gemm logic

* update the gemm profiler logic

* add more difinitions to cppcheck, fix couple more errors

* replace runtime error with message in device function

* fix a couple of int4 issues

* no return for fill function

* fix errors in data_types.hpp

* fix format

* fix few remaining errors

* fix errors in data_types.hpp

* fix last couple of errors in datat_types.hpp

180e5720

19 Jan, 2024 1 commit

Add optimized copy to ck wrapper (#1126) · 7e4eb4b8

Bartłomiej Kocot authored Jan 19, 2024



* Add optimized copy to ck wrapper

* Example optimizations

* Fixes

* Move img2col test to client example

* Refactor example

* Fix docs

* Fixes

* Fix

* Fixes

* Fixes

* Fixes

* Fixes

* Fixes

---------
Co-authored-by: zjing14 <zhangjing14@gmail.com>

7e4eb4b8

03 Jan, 2024 1 commit

Add tensor partition and generic copy for ck wrapper (#1108) · 4234b3a6

Bartłomiej Kocot authored Jan 03, 2024

* Add tensor partition and generic copy for ck wrapper

* Update changelog

* Stylistic fixes

* Change shape/strides logic to descriptor transforms

* Fixes

* Fix client example

* Fix comments

4234b3a6

13 Dec, 2023 1 commit
- Fix the bugs (#1099) · 6891e4d1
  Rostyslav Geyyer authored Dec 13, 2023
  
  6891e4d1
06 Dec, 2023 1 commit

Introduce wrapper library (#1071) · 836b7e55

Bartłomiej Kocot authored Dec 06, 2023

* Introduce wrapper library

* Update cmake files

* Revert "Update cmake files"

This reverts commit c27f88b56590c11a88e26d5d0df7aca51a08133d.

* Fix comments

836b7e55

03 Dec, 2023 1 commit

Add support for double buffering in direct load GEMM kernel (#1052) · bc4bf9bd

Bartlomiej Wroblewski authored Dec 03, 2023

This PR introduces support for double buffering in LDS into GEMM kernels that use direct load instructions.

Direct loads now use inline asm instead of intrinsics. Usage of intrinsics results in compiler adding additional waitcnt instructions what breaks possible load/compute overlap in case of double buffering.

Usage of inline asm results in the need to use sched_barrier in order to make sure that compiler cannot incorrectly reschedule instructions since it does not know the data dependencies between global->LDS and LDS->registers.

bc4bf9bd

30 Nov, 2023 1 commit

Introduce wrapper for layout (#1054) · 8ff845f2

Bartłomiej Kocot authored Nov 30, 2023

* Introduce wrapper for layout

* Extend functionality

* Fix for getLength

* Comment fixes

* Add comments and remove not needed getters

8ff845f2

28 Nov, 2023 1 commit

Switch default f8 conversion to stochastic rounding (#1048) · 6ef034f6

Rostyslav Geyyer authored Nov 27, 2023

* Switch default f8 conversion to stochastic rounding

* Refactor f8-related type_converts

* Add an element-wise op

6ef034f6

25 Nov, 2023 1 commit

Add basic support for direct loads from global to LDS (#999) · 627054b9

Bartlomiej Wroblewski authored Nov 25, 2023

* Add basic support for direct loads from global to LDS

* Clean the code and comments

* Add support for fp16

* Add comments

* Add check for thread cluster lengths

* Align non-direct-load fp16 example

* Small fixes

* Extend IsSupported to check for supported GPU gens

* Build examples only on the supported HW

* Do not throw when instance not supported in 04 example

* Review: Apply review suggestions

* Review: small fix

* Review: small fix

627054b9

07 Nov, 2023 1 commit

Add Gemm instances for performance improvement (#1018) · 98fd41f5

zjing14 authored Nov 07, 2023



* improve kpad

* more tuning parameters

* f16_f8_fp16

* cut test time

* add f16_f8_fp16

* add f16_f8_f16

* testing instances for skinny cases

* format

* clean

* add fp16_f8_fp16

* clang-format

* add grouped gemm instalces

* fixed profile grouped_gemm

* clean

* clean

* clean

* clean

* clean

* add missing instance func

* fixed inferface

---------
Co-authored-by: Jing Zhang <jizha@amd.com>
Co-authored-by: root <root@sh5-1e707-rc06-38.mkm.dcgpu>

98fd41f5

28 Oct, 2023 1 commit

Fix the fp8 gemm for large tensors on MI300. (#1011) · f46a6ffa

Illia Silin authored Oct 27, 2023



* Fix the fp8 conversion

* Try clipping value before conversion

* Fix return

* Simplify with a const

* reduce the gemm input tensor values to reduce round-off error

* replace if-else with lambda

* fix syntax

---------
Co-authored-by: Rostyslav Geyyer <rosty.geyyer@amd.com>

f46a6ffa

20 Oct, 2023 1 commit
- Fix bf8 conversion issues (#1003) · 1fd27d52
  Rostyslav Geyyer authored Oct 20, 2023
```
* Fix the conversion

* Add bf8 functionality

* Enable example on MI200 as well
```
  1fd27d52
19 Oct, 2023 2 commits
- Fix the DL kernel issues on Navi3x. (#998) · f7331c60
  Illia Silin authored Oct 19, 2023
```
* apply the patch for dl kernels on gfx11

* build DL kernels on navi32 CI
```
  f7331c60
- Extend available elementwise operations with conv examples (#995) · 82f3a835
  Bartłomiej Kocot authored Oct 19, 2023
```
* Extend available elementwise operations with conv examples

* Fixes

* Remove not needed convert

* Update CMakeFile and dir name
```
  82f3a835
18 Oct, 2023 1 commit

Clean DTYPES conditions in CMake (#974) · bf435140

zjing14 authored Oct 18, 2023



* Add a condition to build fp8 instances

* simplified buffer_load/store

* add bfp8/fp8

* fixed

* remove all f8/bf8 condition include folder

* fixed cmake conditions

* fixed DTYPES=fp16/bfp16

* fix

* fixed buffer_load

* fixed buffer_store

* fix

* clean example cmake files

* fixed ci

* fixed cit

---------
Co-authored-by: Rostyslav Geyyer <rosty.geyyer@amd.com>
Co-authored-by: Jing Zhang <jizha@amd.com>

bf435140

16 Oct, 2023 1 commit
- workaround with float (#992) · 39430bfd
  zjing14 authored Oct 16, 2023
```
Co-authored-by: Jing Zhang <jizha@amd.com>
```
  39430bfd
13 Oct, 2023 1 commit

add vector_type support into thread_copy_v3r1 (#969) · 2ce9b56c

zjing14 authored Oct 13, 2023



* add vector_type support into thread_copy_v3r1

* remove unncessary type_convert

* fixed datatype

* fixed dataType

* changed API with is_packx_invocable

* changed example

* add missing cmake file

* fixed ci

* fixed cmake

---------
Co-authored-by: Jing Zhang <jizha@amd.com>

2ce9b56c

12 Oct, 2023 1 commit

simplified buffer_load/store (#971) · f3b02ecf

zjing14 authored Oct 11, 2023



* simplified buffer_load/store

* add bfp8/fp8

* fixed

* fixed buffer_load

* fixed buffer_store

---------
Co-authored-by: Jing Zhang <jizha@amd.com>

f3b02ecf

11 Oct, 2023 2 commits

Revert "Grouped Gemm with looping over the tiles. (#788)" (#982) · c99323be
zjing14 authored Oct 11, 2023
```
This reverts commit a4f72a31.
```
c99323be

Grouped Gemm with looping over the tiles. (#788) · a4f72a31

Adam Osewski authored Oct 11, 2023



* Introduce LocalBlockToCTileMap.

* Change the signature of CalculateBottomIndex() function which now does
not accept any argument. The B2C map which is already passed as an
argument to the kernel Run function is calculating block's local id
already outside at kernel entry point __global__ function.
The LocalB2C map stores as members local block ID.

* Use LocalBlockToCTile map in device ops.

* First draft of tile loop work distribution.

* Fix typo.

* Simplify kernel arguments.

Calculate descriptors & B2C maps on the device.

* Use looping kernel.

* Fix B2C constructor.

* Fix Navi21 errors.

* Calculate tile start/end in device kernel.

* Change Run API to accept user provided workspace buffer.

* Add new line at EOF.

* Move Gemm KernelArguments to device op interface.

* Remove unused code.

* Update API.

* Launch grid size which is min of occupancy vs tile count

* Get back to use constant memory for gemm descriptors.

* Remove unused code.

* Add default virtual method implementation.

* Update comments to conform with doxygen style.

* Fix doc style and unused parameters.

* Add thread cluster lengths to kernel name.

* Remove old splitk impl and replace it with tile looping one.

* Modify instances.

* set KPerBlock to 64
* maximize wherever possible vector load size.

* Fix instances cluster lengths.

* Change comment style.

* Use 128b store where possible in instances.

* Update test cases, since KPerBlock has doubled.

* Update output stream operator for Sequence.

* Add pipeline version to GroupedGEMM device op type string.

* Fix pipeline version type logging.

* Fix input tensors type after merge.

* Fix compiler error.

* Fix output stream operator for Pipeline version.

* Store using 128b.

* Set of instances with kpb 32/64

* Limit number of instances

* Remove commented out instances.

* Fix function name.

* Limit the number of instances.

Add pipline version to the regular instances

* Change thr cluster layout for reading B tensor.

* disabled failed instances

---------
Co-authored-by: Adam Osewski <aosewski@amd.com>
Co-authored-by: zjing14 <zhangjing14@gmail.com>
Co-authored-by: Jing Zhang <jizha@amd.com>

a4f72a31

10 Oct, 2023 1 commit

Fixed f8_gemm NaN (#975) · ac9595a9

zjing14 authored Oct 10, 2023



* workaround nan problem by changing output to fp16

* enable f8/bf8 gemm tests on MI200

* workaround f16 to f8 conversion

---------
Co-authored-by: Jing Zhang <jizha@amd.com>

ac9595a9

04 Oct, 2023 1 commit

Add conv bwd weight fp16 comp bf8 fp8 op, instances and example (#945) · 42facfc6

Rostyslav Geyyer authored Oct 04, 2023



* Add f8 bf8 gemm example

* Add element-wise ops

* Add intrinsics

* Update reference calculation

* Add an additional type option for xdlops gemm

* Fix build process

* Add bf8 to buffer addressing

* Update blockwise op, split typeA and typeB

* Update for compatibility

* Uppdate naming to f8->fp8

* Update naming

* Format

* Update naming (#937)

* Add a client example

* Add computetypes to device and gridwise ops

* Add instances, update instance factory

* Format

* Fix a flag

* Add ckProfiler mode

* Fix typos

* Add an example

* Add bf8 generator

* add bf8 mfma; fixed type_convert for bf8

* move verfication ahead of timing

* Update reference calculation

* Fix reference

* Narrow down float init range

* Fix bf8 bf8 mfma

* Add bf8 @ fp8 mfma

* Update example

* Update instances

* Update profiler api

* Update for compatibility

* Format

* Remove extra example

* Clean up

* workaround convert

---------
Co-authored-by: Jing Zhang <jizha@amd.com>

42facfc6

02 Oct, 2023 1 commit

Add fp8 @ bf8 gemm support and example (#933) · bd09b5c5

Rostyslav Geyyer authored Oct 02, 2023

* Add f8 bf8 gemm example

* Add element-wise ops

* Add intrinsics

* Update reference calculation

* Add an additional type option for xdlops gemm

* Fix build process

* Add bf8 to buffer addressing

* Update blockwise op, split typeA and typeB

* Update for compatibility

* Uppdate naming to f8->fp8

* Update naming

* Format

bd09b5c5

27 Sep, 2023 3 commits

Handle type conversions to a const datatype (#944) · f4af5aed

Bartlomiej Wroblewski authored Sep 27, 2023

* Handle type conversions to a const datatype

* Review: Handle X being const data type as well

* Review: Remove typo

f4af5aed

Add column to image kernel (#930) · e2243a4d

Bartłomiej Kocot authored Sep 27, 2023

* Add column to image kernel

* Minor fixes for dtypes and client examples

* Disable tests for disabled dtypes

* Disable add instances functions for disabled data types

* Minor stylistic fixes

* Revert "Disable add instances functions for disabled data types"

This reverts commit 728b86956378dcd9415fd0f2557833a068fe1c10.

* Instances reduction

* Add comments in device_column_to_image_impl

* Update changelog and Copyrights

* Improve changelog

e2243a4d

Add multiple A/B support (#906) · 11676c7e

zjing14 authored Sep 26, 2023



* add gridwise_multi_abd

* move element_op into RunRead

* merge element_wise op with data read

* add multiABD example

* allow packed elementwise_op

* changed example

* clean

* clean

* add is_detected

* fix

* minor fix

* add scaleAdd_vec4 example

---------
Co-authored-by: Jing Zhang <jizha@amd.com>

11676c7e

18 Sep, 2023 1 commit
- Add native conversions fp8<->fp32 (#908) · f17af2e9
  Rostyslav Geyyer authored Sep 17, 2023
```
* Add native conversions

* Add bf8 conversions
```
  f17af2e9
13 Sep, 2023 1 commit

Add grouped conv bwd weight dl instances and new layout (#897) · 475188ca

Bartłomiej Kocot authored Sep 13, 2023

* Add grouped conv bwd weight dl instances and new layout

* Add M and N padding

* Remove todo comment

* Enable grouped conv fwd dl k,c=1 generic instance

* Comment fixes

475188ca

12 Sep, 2023 1 commit

Refactor f8_t, add bf8_t (#792) · 62d4af74

Rostyslav Geyyer authored Sep 12, 2023

* Refactor f8_t to add bf8_t

* Add check_err impl for f8_t

* Update fp8 test

* Format

* Revert the fix

* Update vector_type implementation

* Add bf8 test

* Add bf8, use BitInt types

* Add bf8 conversion methods

* Update type_convert for fp8/bf8

* Add check_err fp8/bf8 support

* Add subnorm fp8 tests

* Add subnorm bf8 tests

* Fix conversion

* Add bf8 cmake bindings

* Add macros to enable build with disabled fp8/bf8

* Remove is_native method

* Update flag combination for mixed precision instances

* Add more flag checks

* Add another flag to a client example

* Add type traits, decouple f8/bf8 casting

* Clean up

* Decouple fp8 and bf8 flags

* Remove more redundant flags

* Remove leftover comments

62d4af74

06 Sep, 2023 1 commit

Redesign the DPP8 GEMM kernel to use warp-wise component (#863) · 37a8c1f7

Bartlomiej Wroblewski authored Sep 06, 2023

* Redesign the DPP8 GEMM kernel to use warp-wise component

* Review: Improve error messages

* Review: Remove unnecessary empty lines

* Review: Fix M, N per thread names

* Review: Rename mfma_input_type to dpp_input_type

* Review: Fix tensor adaptor; remove unnecessary element

* Review: Remove calls to dpp_gemm's MakeCDescriptor

* Review: Add blockwise doc, change function names to include dimension names

* Review: Remove duplicated code; Move Block2CtileMap alias to the top of the file

* Review: Add __restrict__ keywords

* Review: Use MatrixPadder for padding A, B, C matrices

* Review: Remove hardcoded datatypes

* Review: Change names from FloatX to XDataType

* Review: Introduce AK0 and BK0 instead of a single K0

* Review: Remove construction of dpp_datatypes object

* Review: Rename DppInstrRunner to DppLanegroupGemm

37a8c1f7