- 16 Mar, 2023 3 commits
- 15 Mar, 2023 13 commits
-
-
zjing14 authored
-
Rostyslav Geyyer authored
* Add layout check to IsSupportedArgument * Format --------- Co-authored-by:
Rosty Geyyer <rosty.geyyer@amd.com> Co-authored-by:
zjing14 <zhangjing14@gmail.com>
-
Illia Silin authored
* make conv_fwd_bias_activation kernel id unique * add more parameters to conv and gemm kernel names * update GetTypeString for conv and gemm kernels * fix two more kernel strings
-
Haocong WANG authored
-
rocking authored
Let acc and CShuffleDataType be the same in xdlops
-
rocking authored
Merge branch 'conv_dlops/quantization' of github.com:ROCmSoftwarePlatform/composable_kernel into conv_dlops/quantization
-
rocking5566 authored
-
rocking authored
-
rocking authored
-
rocking authored
-
rocking authored
-
rocking authored
-
- 13 Mar, 2023 1 commit
-
-
rocking authored
-
- 10 Mar, 2023 2 commits
-
-
Rostyslav Geyyer authored
Co-authored-by:Rosty Geyyer <rosty.geyyer@amd.com>
-
Haocong WANG authored
* Change gridwise gemm mD blockwise gemm to naive * RRR Gemm fix * Fix RCR gemm bug * Isolate wmma instructions * Update amd_inline_asm.hpp * Update amd_wmma.hpp * Update amd_wmma.hpp * fix syntax and update Jenkinsfile --------- Co-authored-by:
zjing14 <zhangjing14@gmail.com> Co-authored-by:
Illia Silin <98187287+illsilin@users.noreply.github.com> Co-authored-by:
illsilin <Illia.Silin@amd.com>
-
- 09 Mar, 2023 7 commits
-
-
rocking authored
-
carlushuang authored
Co-authored-by:zjing14 <zhangjing14@gmail.com>
-
Illia Silin authored
* enable building on Nav31 * fix syntax * replace GPU_TARGETS with offload-arch * add gfx1102 rachitecture * fix typo * update changelog
-
rocking authored
-
rocking authored
-
rocking authored
-
rocking authored
-
- 08 Mar, 2023 7 commits
-
-
rocking authored
-
rocking authored
-
rocking authored
-
rocking authored
-
rocking5566 authored
-
rocking authored
-
Adam Osewski authored
* Grouped gemm + Gelu instances. * Device Instance Factory for GroupedGemm+Gelu * Client example * Rangify fill helper functions. * Fix name clash. * Profiler for grouped_gemm+gelu * No need to use full namespace name. * Add check for MRaw divisible by vector load. * Ugly fix for big errors. * Add grouped_gemm+gelu to profiler CMakelists. * Store in argument additional info. * Information about Mraw, Nraw, Kraw values. * Use FastGelu instead of Gelu. * Change client ex to use FastGelu * Remove relaxed error precision. * Remove duplicate output elementwise-op --------- Co-authored-by:
Adam Osewski <aosewski@amd.com> Co-authored-by:
zjing14 <zhangjing14@gmail.com>
-
- 07 Mar, 2023 3 commits
- 06 Mar, 2023 3 commits
-
-
Rostyslav Geyyer authored
Co-authored-by:Rosty Geyyer <rosty.geyyer@amd.com>
-
pmaybank authored
* Modify Doxygen config to pick up include directories recursively * Add DeviceMem struct to API Reference guide * Add classes that are used in Flash Attention kernel * Add a reference and config for generating bibliography Co-authored-by:Philip Maybank <Philip.Maybank@amd.com>
-
rocking authored
-
- 02 Mar, 2023 1 commit
-
-
Illia Silin authored
* add new parallel stage on navi node * dont run performance tests on navi, get rid of 9110 compiler * only run navi build when not doing QA * fix syntax * use navi21 label * dont stash profiler on navi nodes, scp deb package to ginger * disable tests on navi nodes * test posting a binary to ginger * add sshpass and use it to copy deb package * fix the scp example * fix syntax * debug the scp issues * add jenkins user to docker * dont try whoami * change jenkins uid and add user with uid=1002 * try scp from the last stage on micimaster * rename and stash the package, scp from micimaster
-