Grouped Gemm with looping over the tiles. (#788)

* Introduce LocalBlockToCTileMap. * Change the signature of CalculateBottomIndex() function which now does not accept any argument. The B2C map which is already passed as an argument to the kernel Run function is calculating block's local id already outside at kernel entry point __global__ function. The LocalB2C map stores as members local block ID. * Use LocalBlockToCTile map in device ops. * First draft of tile loop work distribution. * Fix typo. * Simplify kernel arguments. Calculate descriptors & B2C maps on the device. * Use looping kernel. * Fix B2C constructor. * Fix Navi21 errors. * Calculate tile start/end in device kernel. * Change Run API to accept user provided workspace buffer. * Add new line at EOF. * Move Gemm KernelArguments to device op interface. * Remove unused code. * Update API. * Launch grid size which is min of occupancy vs tile count * Get back to use constant memory for gemm descriptors. * Remove unused code. * Add default virtual method implementation. * Update comments to conform with doxygen style. * Fix doc style and unused parameters. * Add thread cluster lengths to kernel name. * Remove old splitk impl and replace it with tile looping one. * Modify instances. * set KPerBlock to 64 * maximize wherever possible vector load size. * Fix instances cluster lengths. * Change comment style. * Use 128b store where possible in instances. * Update test cases, since KPerBlock has doubled. * Update output stream operator for Sequence. * Add pipeline version to GroupedGEMM device op type string. * Fix pipeline version type logging. * Fix input tensors type after merge. * Fix compiler error. * Fix output stream operator for Pipeline version. * Store using 128b. * Set of instances with kpb 32/64 * Limit number of instances * Remove commented out instances. * Fix function name. * Limit the number of instances. Add pipline version to the regular instances * Change thr cluster layout for reading B tensor. * disabled failed instances --------- Co-authored-by: Adam Osewski <aosewski@amd.com> Co-authored-by: zjing14 <zhangjing14@gmail.com> Co-authored-by: Jing Zhang <jizha@amd.com>

Grouped Gemm with looping over the tiles. (#788)
* Introduce LocalBlockToCTileMap. * Change the signature of CalculateBottomIndex() function which now does not accept any argument. The B2C map which is already passed as an argument to the kernel Run function is calculating block's local id already outside at kernel entry point __global__ function. The LocalB2C map stores as members local block ID. * Use LocalBlockToCTile map in device ops. * First draft of tile loop work distribution. * Fix typo. * Simplify kernel arguments. Calculate descriptors & B2C maps on the device. * Use looping kernel. * Fix B2C constructor. * Fix Navi21 errors. * Calculate tile start/end in device kernel. * Change Run API to accept user provided workspace buffer. * Add new line at EOF. * Move Gemm KernelArguments to device op interface. * Remove unused code. * Update API. * Launch grid size which is min of occupancy vs tile count * Get back to use constant memory for gemm descriptors. * Remove unused code. * Add default virtual method implementation. * Update comments to conform with doxygen style. * Fix doc style and unused parameters. * Add thread cluster lengths to kernel name. * Remove old splitk impl and replace it with tile looping one. * Modify instances. * set KPerBlock to 64 * maximize wherever possible vector load size. * Fix instances cluster lengths. * Change comment style. * Use 128b store where possible in instances. * Update test cases, since KPerBlock has doubled. * Update output stream operator for Sequence. * Add pipeline version to GroupedGEMM device op type string. * Fix pipeline version type logging. * Fix input tensors type after merge. * Fix compiler error. * Fix output stream operator for Pipeline version. * Store using 128b. * Set of instances with kpb 32/64 * Limit number of instances * Remove commented out instances. * Fix function name. * Limit the number of instances. Add pipline version to the regular instances * Change thr cluster layout for reading B tensor. * disabled failed instances --------- Co-authored-by: Adam Osewski <aosewski@amd.com> Co-authored-by: zjing14 <zhangjing14@gmail.com> Co-authored-by: Jing Zhang <jizha@amd.com>
a4f72a31 · Adam Osewski · GitHub · 98c80714 · a4f72a31
Unverified Commit a4f72a31 authored Oct 11, 2023 by Adam Osewski Committed by GitHub Oct 10, 2023
Hide whitespace changes
Inline Side-by-side

Showing with 11 additions and 11 deletions

test/grouped_gemm/test_grouped_gemm_ut_cases.inc test/grouped_gemm/test_grouped_gemm_ut_cases.inc +11 -11

No files found.
--- a/test/grouped_gemm/test_grouped_gemm_ut_cases.inc
+++ b/test/grouped_gemm/test_grouped_gemm_ut_cases.inc
@@ -4,7 +4,7 @@ TEST_P(RRR_F16_F16_F16, TinyCases)
 {
    const std::vector<int> Ms{0, 1};
    constexpr int N = 768;
-    constexpr int K = 544;
+    constexpr int K = 1088;

    const std::vector<int> Ns(Ms.size(), N);
    const std::vector<int> Ks(Ms.size(), K);
@@ -17,9 +17,9 @@ TEST_P(RRR_F16_F16_F16, TinyCases)

 TEST_P(RRR_F16_F16_F16, SmallCases)
 {
-    const std::vector<int> Ms{2, 1, 3, 4, 5, 0};
+    const std::vector<int> Ms{2, 3, 4, 5};
    constexpr int N = 768;
-    constexpr int K = 544;
+    constexpr int K = 1088;

    const std::vector<int> Ns(Ms.size(), N);
    const std::vector<int> Ks(Ms.size(), K);
@@ -34,7 +34,7 @@ TEST_P(RRR_F16_F16_F16, MidCases)
 {
    const std::vector<int> Ms{167, 183, 177, 153, 139, 204};
    constexpr int N = 768;
-    constexpr int K = 544;
+    constexpr int K = 1088;

    const std::vector<int> Ns(Ms.size(), N);
    const std::vector<int> Ks(Ms.size(), K);
@@ -49,7 +49,7 @@ TEST_P(RRR_F16_F16_F16, Regular)
 {
    const std::vector<int> Ms{64, 128, 256};
    constexpr int N = 768;
-    constexpr int K = 320;
+    constexpr int K = 640;

    const std::vector<int> Ns(Ms.size(), N);
    const std::vector<int> Ks(Ms.size(), K);
@@ -79,7 +79,7 @@ TEST_P(RCR_F16_F16_F16, TinyCases)
 {
    const std::vector<int> Ms{0, 1};
    constexpr int N = 768;
-    constexpr int K = 544;
+    constexpr int K = 1088;

    const std::vector<int> Ns(Ms.size(), N);
    const std::vector<int> Ks(Ms.size(), K);
@@ -91,9 +91,9 @@ TEST_P(RCR_F16_F16_F16, TinyCases)

 TEST_P(RCR_F16_F16_F16, SmallCases)
 {
-    const std::vector<int> Ms{2, 1, 3, 4, 5, 0};
+    const std::vector<int> Ms{2, 3, 4, 5};
    constexpr int N = 768;
-    constexpr int K = 544;
+    constexpr int K = 1088;

    const std::vector<int> Ns(Ms.size(), N);
    const std::vector<int> Ks(Ms.size(), K);
@@ -123,7 +123,7 @@ TEST_P(RCR_F16_F16_F16, Regular)
 {
    const std::vector<int> Ms{32, 64, 128, 256};
    constexpr int N = 768;
-    constexpr int K = 320;
+    constexpr int K = 640;

    const std::vector<int> Ns(Ms.size(), N);
    const std::vector<int> Ks(Ms.size(), K);
@@ -151,9 +151,9 @@ TEST_P(RCR_F16_F16_F16, MNKPadded)

 TEST_P(RRR_F16_F16_F16_LargeK, TestLargeKBatch)
 {
-    const std::vector<int> Ms{188, 210};
+    const std::vector<int> Ms{127, 150, 188, 210};
    constexpr int N = 768;
-    constexpr int K = 4096;
+    constexpr int K = 8192;

    const std::vector<int> Ns(Ms.size(), N);
    const std::vector<int> Ks(Ms.size(), K);