Merge from internal (#1857)

* enable batched_gemm_softmax_gemm_perm_wmma for gfx12 * disable instances with blocksize=256 in attention examples * debuggging * debug * fixed lds_enabled * debugging * Fix and add limit to skiplds feature * Enable skipLds feature and fix compilation bugs * add ck_tile definitions for gfx12 * fix clang format and test/wmma_op * updage instances cmake for gfx12 * disable the test_wmma_op on gfx12 * fix the builds for gfx950 * add gfx12 and gfx950 to default target list * clean-up cmake file * Initial introduction of OFP8 data types. * Renamed FP8 and BF8 tests into FP8_FNUZ and BF8_FNUZ. * Implementation of ConvertFP32Nearest in test_fp8_ocp. * Remove dependence on possibly undeclared alias. * Implement FP8OCP test for stochastic rounding mode. * Implement FP8OCP tests for half_t type conversions. * enable bf16 atomic add on gfx950 * Implement ConvertFP32Nearest test. * Implement ConvertFP32Stochastic test. * Implement ConvertFP16Nearest and ConvertFP16Stochastic tests. * Refactoring. Move FP8 definitions into a separate header file. * Enable easy switching between architectures. * Fix compilation error for gfx942 architecture. * Add fp4 type with constants * only builf gfx950 branch for gfx950 target by default * Enable OCP build of example_gemm_xdl_fp8. * Fix formatting. * fix the build logic for gfx950 * Improve GEMM example verbosity. * Add constexpr where applicable. * fix the logic of enabling XDL and WMMA instances * Improve GEMM example verbosity. * Enable build of example_gemm_xdl_fp8_bf8 test. * Fix tests for gfx1101 architecture. * Build DPP examples only on gfx103 and gfx11 architectures. * Optionaly run either CPU or GPU verifications with GEMM examples. * Extend GeneratorTensor_Sequential to produce values of prescribed data types. * Add missing constructor. * Add scale type and mxfp conversions * Update conversions * Add conversion tests * Fix typo * Improve infrastructure for OFP8 data type support. * BUGFIX. Should not use FP8 as Compute/Accum data type. * Add custom target for grouped_convnd_bwd_weight tests. * Can build `tests` target on gfx950. * Bugfixes on gfx1101 architecture. * Fix dependencies. * Add stochastic rounding tests * Provide single point of truth for FP8 INF and NAN checks * Prevent instantiation of operators that are not supported by FP8 data types * Add FP8 type selection into client_axample CMakeLists.txt * Prevent sccache server from shutting down during build * Fix test success reporting logic * Change default verification method to CPU. GPU verification takes too much time to complete on the emulator. * Add scale <-> float conversions * Add scaled conversions with tests * Add device conversions * Make sure all tests and examples are built for gfx950 * Facilitate testing of FP8 data types on the emulator * Introduce two new tensor generators * Enable instances built for gfx94 to be built on gfx950 * Verify 35_splitk_gemm on floating point numbers. splitk gemm appears to be losing precision VS reference implementation when FP numbers are involved. * Format * Verify 04_gemm_add_add_fastgelu on floating point numbers * Verify 20_grouped_conv_bwd_weight on floating point numbers * Verify 38_grouped_conv_bwd_data_multiple_d on floating point numbers * Verify more tests on floating point data * Fix data types and improve testing verbocity. * Add fp4 vectors * Add debug tests * Upgrade to NPI 573 build docker. * Skip on gemm_universal tests. The tests take too long to complete on the emulator. Need to see if it is possible to reduce the scope of the testing to just FP8 data types. * Add new mfma instructions and examples * Add preprocessor directives for gfx950 specific code * Fix gfx1101 build * Document test availability * Re-enable fp8 gemms for gfx94/95 * Cherry-pick GEMM Universal tests for FP8 data types * Cleanup * Add vector types and tests * Add check_err function * Add tensor generators * CK_USE_GFX94 has already been set on this branch * Fix * Address formatting issues and leftovers * Make fail/pass logic consistent within 01_gemm folder Removed multiple negations in fail/pass logic to propagate `true` as the success indicator. * Fix GPU verification reporting logic. * Update year in copyright notice. * Cleanup * Use `enum class` instead of `enum` * Remove set_property for FP8 tests * Add vector conversions * Fix * Fix linker errror * Clean up * Fix gfx950 conversions * Clean up * Fix more gfx950 conversions * Fix even more gfx950 conversions * Narrowing the scope of PR to OCP FP8 enablement only * Add tests for OCP FP8 vector_type storage * Fix client examples build * Fix typo * Update e8m0 casting * Rename E8M0 type * Update unpack method * Cleanup merge artifacts * Enable gemm kernel on all gfx9 architectures (#227) * clean-up * Implement `non_native_vector_base` with `ext_vector_type` array. (#232) * Enable support of 1, 2, 4, and 8-byte custom types in CK. * Fix pool tests for OCP FP8 data type * Fix build * Add ckProfiler gemm instances for new mfma instructions and fix ckProfiler build on MI350 * fix clang format * Add new mfma instructions and examples * Add preprocessor directives for gfx950 specific code * Add ckProfiler gemm instances for new mfma instructions and fix ckProfiler build on MI350 * fix clang format * Fix clang format for the newly merged files * Use the existing example instances for fp16 bf16 and int8 * Remove comment on new mfma instructions in MfmaInstr * Update include/ck/tensor_operation/gpu/grid/gridwise_batched_gemm_gemm_xdl_cshuffle_v1.hpp Co-authored-by: Andriy Roshchenko <107577548+andriy-ca@users.noreply.github.com> * merge from public repo * Fix ck build * Fix ck build * Use double for max_abs_in_val * Move scaled_type_convert functions to a separate header (#251) * re-enable building mha lib and gemm_universal_f8 instances for gfx950 * Update library/src/tensor_operation_instance/gpu/CMakeLists.txt Co-authored-by: Andriy Roshchenko <107577548+andriy-ca@users.noreply.github.com> * fix typo for CK_USE_OCP_FP8 * fix typo for CK_USE_OCP_FP8 * Add FP6 and BF6 types (#261) * Add a rounding flag * Add FP6 and BF6 * Add tests Co-authored-by: Andriy Roshchenko <107577548+andriy-ca@users.noreply.github.com> * Clean up --------- Co-authored-by: Andriy Roshchenko <107577548+andriy-ca@users.noreply.github.com> * fix one more typo * Refactor E8M0 scale implementation (#262) * Refactor E8M0 scale implementation * Add MXFP6 and MXBF6 conversion methods (#270) * Add conversions * Add tests * Add docstrings * Add scaled conversions * Add fp6/bf6 tests * Remove misleading fp4 test case * Add docstrings * Clean up * Address comments * Set stricter tolerances for RNE tests * Add missing tests * Add native conversions to float * Revert "Add native conversions to float" This reverts commit 09467111f73b753c8cc3d597533b187940353dab. * Update copyright years * replace the fp6 with bf6 convert calls in test_bf6 * fix test_bf6 * enable smfmac test * [MX FP8] Add Scaled Type Convert Functions for OCP FP8/BF8 data types (#271) * Move scaled_type_convert functions to a separate header * Introduce MX data tests * Build MX tests only on relevant architectures * Refactor E8M0 scale implementation * Fix `config.h` typo * Cleanup deprecated symbols * Refactor `amd_ck_fp8.hpp` * `scaled_type_convert` for `f8_ocp_t` * Implement test for MX FP8 scaled type convert * Implement test for MX BF8 scaled type convert * Scaled type convert for vectors of 2 FP8 elements * Scaled type convert for vectors of 16 FP8 elements * Implementation of scaled conversion from F32 to F8 * Add tests for scaled conversions from FP32 to FP8 * Add documentation to the test functions * Implementation of scaled conversion from F32x2 to F8x2 * Implementation of scaled conversion from F32x16 to F8x16 * Implementation of scaled conversion from F32x32 to F8x32 * Implementation of scaled conversion from F8x32 to F32x32 * Verified on the emulator * MX FP GEMM - Example Template (#277) Temporarily uses `DeviceGemmMultiD_ABScale_Xdl_CShuffle_V3` kernel and 128x128 scaling matrices. Must be modified to use MX-native GEMM kernell with 16 or 32 component vectors per scale. Verified on the emulator. * Add vector support * Add tests * Add missing type aliases * Fix test naming * only build mx example for gfx950 * disable CK_USE_AMD_MFMA_GFX950 by default * fic build for multiple archs * fix typo * fix typo * Update unpack signature * Fix merge * Add size checks in pack function * Add a flag * Add conversions * Fix build logic * Update pack/unpack methods * Remove unneeded AsType accessors * Add docstrings * Add a flag to config file * Test the functionality of V_MFMA_F32_16X16X128_F8F6F4 and V_MFMA_F32_32X32X64_F8F6F4 instructions. (#293) * Introduced MFMA tests * Verified f8f6f4 MFMA Instructions * Move flag logic to scaled_type_convert header * Use pointers instead of array indices * Fix a typo * Update tests and pack functions * Fix gemm gemm on gfx950 * Fix clang format * restore the default gput target lists * fix the jenkinsfile * add missing ifdef --------- Co-authored-by: Jing Zhang <jizhan@amd.com> Co-authored-by: aska-0096 <haocwang@amd.com> Co-authored-by: Jun Liu <Liu.Jun@amd.com> Co-authored-by: Andriy Roshchenko <andriy.roshchenko@amd.com> Co-authored-by: Rostyslav Geyyer <rosty.geyyer@amd.com> Co-authored-by: Rostyslav Geyyer <46627076+geyyer@users.noreply.github.com> Co-authored-by: root <root@banff-cyxtera-s83-2.ctr.dcgpu> Co-authored-by: Andriy Roshchenko <107577548+andriy-ca@users.noreply.github.com> Co-authored-by: jefyang1 <146495389+jefyang1@users.noreply.github.com> Co-authored-by: jefyang1 <Jeffreyj.Yang@amd.com>

Merge from internal (#1857)
* enable batched_gemm_softmax_gemm_perm_wmma for gfx12 * disable instances with blocksize=256 in attention examples * debuggging * debug * fixed lds_enabled * debugging * Fix and add limit to skiplds feature * Enable skipLds feature and fix compilation bugs * add ck_tile definitions for gfx12 * fix clang format and test/wmma_op * updage instances cmake for gfx12 * disable the test_wmma_op on gfx12 * fix the builds for gfx950 * add gfx12 and gfx950 to default target list * clean-up cmake file * Initial introduction of OFP8 data types. * Renamed FP8 and BF8 tests into FP8_FNUZ and BF8_FNUZ. * Implementation of ConvertFP32Nearest in test_fp8_ocp. * Remove dependence on possibly undeclared alias. * Implement FP8OCP test for stochastic rounding mode. * Implement FP8OCP tests for half_t type conversions. * enable bf16 atomic add on gfx950 * Implement ConvertFP32Nearest test. * Implement ConvertFP32Stochastic test. * Implement ConvertFP16Nearest and ConvertFP16Stochastic tests. * Refactoring. Move FP8 definitions into a separate header file. * Enable easy switching between architectures. * Fix compilation error for gfx942 architecture. * Add fp4 type with constants * only builf gfx950 branch for gfx950 target by default * Enable OCP build of example_gemm_xdl_fp8. * Fix formatting. * fix the build logic for gfx950 * Improve GEMM example verbosity. * Add constexpr where applicable. * fix the logic of enabling XDL and WMMA instances * Improve GEMM example verbosity. * Enable build of example_gemm_xdl_fp8_bf8 test. * Fix tests for gfx1101 architecture. * Build DPP examples only on gfx103 and gfx11 architectures. * Optionaly run either CPU or GPU verifications with GEMM examples. * Extend GeneratorTensor_Sequential to produce values of prescribed data types. * Add missing constructor. * Add scale type and mxfp conversions * Update conversions * Add conversion tests * Fix typo * Improve infrastructure for OFP8 data type support. * BUGFIX. Should not use FP8 as Compute/Accum data type. * Add custom target for grouped_convnd_bwd_weight tests. * Can build `tests` target on gfx950. * Bugfixes on gfx1101 architecture. * Fix dependencies. * Add stochastic rounding tests * Provide single point of truth for FP8 INF and NAN checks * Prevent instantiation of operators that are not supported by FP8 data types * Add FP8 type selection into client_axample CMakeLists.txt * Prevent sccache server from shutting down during build * Fix test success reporting logic * Change default verification method to CPU. GPU verification takes too much time to complete on the emulator. * Add scale <-> float conversions * Add scaled conversions with tests * Add device conversions * Make sure all tests and examples are built for gfx950 * Facilitate testing of FP8 data types on the emulator * Introduce two new tensor generators * Enable instances built for gfx94 to be built on gfx950 * Verify 35_splitk_gemm on floating point numbers. splitk gemm appears to be losing precision VS reference implementation when FP numbers are involved. * Format * Verify 04_gemm_add_add_fastgelu on floating point numbers * Verify 20_grouped_conv_bwd_weight on floating point numbers * Verify 38_grouped_conv_bwd_data_multiple_d on floating point numbers * Verify more tests on floating point data * Fix data types and improve testing verbocity. * Add fp4 vectors * Add debug tests * Upgrade to NPI 573 build docker. * Skip on gemm_universal tests. The tests take too long to complete on the emulator. Need to see if it is possible to reduce the scope of the testing to just FP8 data types. * Add new mfma instructions and examples * Add preprocessor directives for gfx950 specific code * Fix gfx1101 build * Document test availability * Re-enable fp8 gemms for gfx94/95 * Cherry-pick GEMM Universal tests for FP8 data types * Cleanup * Add vector types and tests * Add check_err function * Add tensor generators * CK_USE_GFX94 has already been set on this branch * Fix * Address formatting issues and leftovers * Make fail/pass logic consistent within 01_gemm folder Removed multiple negations in fail/pass logic to propagate `true` as the success indicator. * Fix GPU verification reporting logic. * Update year in copyright notice. * Cleanup * Use `enum class` instead of `enum` * Remove set_property for FP8 tests * Add vector conversions * Fix * Fix linker errror * Clean up * Fix gfx950 conversions * Clean up * Fix more gfx950 conversions * Fix even more gfx950 conversions * Narrowing the scope of PR to OCP FP8 enablement only * Add tests for OCP FP8 vector_type storage * Fix client examples build * Fix typo * Update e8m0 casting * Rename E8M0 type * Update unpack method * Cleanup merge artifacts * Enable gemm kernel on all gfx9 architectures (#227) * clean-up * Implement `non_native_vector_base` with `ext_vector_type` array. (#232) * Enable support of 1, 2, 4, and 8-byte custom types in CK. * Fix pool tests for OCP FP8 data type * Fix build * Add ckProfiler gemm instances for new mfma instructions and fix ckProfiler build on MI350 * fix clang format * Add new mfma instructions and examples * Add preprocessor directives for gfx950 specific code * Add ckProfiler gemm instances for new mfma instructions and fix ckProfiler build on MI350 * fix clang format * Fix clang format for the newly merged files * Use the existing example instances for fp16 bf16 and int8 * Remove comment on new mfma instructions in MfmaInstr * Update include/ck/tensor_operation/gpu/grid/gridwise_batched_gemm_gemm_xdl_cshuffle_v1.hpp Co-authored-by: Andriy Roshchenko <107577548+andriy-ca@users.noreply.github.com> * merge from public repo * Fix ck build * Fix ck build * Use double for max_abs_in_val * Move scaled_type_convert functions to a separate header (#251) * re-enable building mha lib and gemm_universal_f8 instances for gfx950 * Update library/src/tensor_operation_instance/gpu/CMakeLists.txt Co-authored-by: Andriy Roshchenko <107577548+andriy-ca@users.noreply.github.com> * fix typo for CK_USE_OCP_FP8 * fix typo for CK_USE_OCP_FP8 * Add FP6 and BF6 types (#261) * Add a rounding flag * Add FP6 and BF6 * Add tests Co-authored-by: Andriy Roshchenko <107577548+andriy-ca@users.noreply.github.com> * Clean up --------- Co-authored-by: Andriy Roshchenko <107577548+andriy-ca@users.noreply.github.com> * fix one more typo * Refactor E8M0 scale implementation (#262) * Refactor E8M0 scale implementation * Add MXFP6 and MXBF6 conversion methods (#270) * Add conversions * Add tests * Add docstrings * Add scaled conversions * Add fp6/bf6 tests * Remove misleading fp4 test case * Add docstrings * Clean up * Address comments * Set stricter tolerances for RNE tests * Add missing tests * Add native conversions to float * Revert "Add native conversions to float" This reverts commit 09467111f73b753c8cc3d597533b187940353dab. * Update copyright years * replace the fp6 with bf6 convert calls in test_bf6 * fix test_bf6 * enable smfmac test * [MX FP8] Add Scaled Type Convert Functions for OCP FP8/BF8 data types (#271) * Move scaled_type_convert functions to a separate header * Introduce MX data tests * Build MX tests only on relevant architectures * Refactor E8M0 scale implementation * Fix `config.h` typo * Cleanup deprecated symbols * Refactor `amd_ck_fp8.hpp` * `scaled_type_convert` for `f8_ocp_t` * Implement test for MX FP8 scaled type convert * Implement test for MX BF8 scaled type convert * Scaled type convert for vectors of 2 FP8 elements * Scaled type convert for vectors of 16 FP8 elements * Implementation of scaled conversion from F32 to F8 * Add tests for scaled conversions from FP32 to FP8 * Add documentation to the test functions * Implementation of scaled conversion from F32x2 to F8x2 * Implementation of scaled conversion from F32x16 to F8x16 * Implementation of scaled conversion from F32x32 to F8x32 * Implementation of scaled conversion from F8x32 to F32x32 * Verified on the emulator * MX FP GEMM - Example Template (#277) Temporarily uses `DeviceGemmMultiD_ABScale_Xdl_CShuffle_V3` kernel and 128x128 scaling matrices. Must be modified to use MX-native GEMM kernell with 16 or 32 component vectors per scale. Verified on the emulator. * Add vector support * Add tests * Add missing type aliases * Fix test naming * only build mx example for gfx950 * disable CK_USE_AMD_MFMA_GFX950 by default * fic build for multiple archs * fix typo * fix typo * Update unpack signature * Fix merge * Add size checks in pack function * Add a flag * Add conversions * Fix build logic * Update pack/unpack methods * Remove unneeded AsType accessors * Add docstrings * Add a flag to config file * Test the functionality of V_MFMA_F32_16X16X128_F8F6F4 and V_MFMA_F32_32X32X64_F8F6F4 instructions. (#293) * Introduced MFMA tests * Verified f8f6f4 MFMA Instructions * Move flag logic to scaled_type_convert header * Use pointers instead of array indices * Fix a typo * Update tests and pack functions * Fix gemm gemm on gfx950 * Fix clang format * restore the default gput target lists * fix the jenkinsfile * add missing ifdef --------- Co-authored-by: Jing Zhang <jizhan@amd.com> Co-authored-by: aska-0096 <haocwang@amd.com> Co-authored-by: Jun Liu <Liu.Jun@amd.com> Co-authored-by: Andriy Roshchenko <andriy.roshchenko@amd.com> Co-authored-by: Rostyslav Geyyer <rosty.geyyer@amd.com> Co-authored-by: Rostyslav Geyyer <46627076+geyyer@users.noreply.github.com> Co-authored-by: root <root@banff-cyxtera-s83-2.ctr.dcgpu> Co-authored-by: Andriy Roshchenko <107577548+andriy-ca@users.noreply.github.com> Co-authored-by: jefyang1 <146495389+jefyang1@users.noreply.github.com> Co-authored-by: jefyang1 <Jeffreyj.Yang@amd.com>
800cf897 · Illia Silin · GitHub · 85d6fcd3 · 800cf897 · 800cf897
Unverified Commit 800cf897 authored Feb 04, 2025 by Illia Silin Committed by GitHub Feb 04, 2025
17 changed files
--- a/library/src/tensor_operation_instance/gpu/gemm_universal_streamk/device_gemm_xdl_universal_streamk_f16_f16_f16/device_gemm_xdl_universal_streamk_f16_f16_f16_mk_kn_mn.hpp
+++ b/library/src/tensor_operation_instance/gpu/gemm_universal_streamk/device_gemm_xdl_universal_streamk_f16_f16_f16/device_gemm_xdl_universal_streamk_f16_f16_f16_mk_kn_mn.hpp
@@ -34,12 +34,14 @@ static constexpr auto Interwave = BlockGemmPipelineScheduler::Interwave;
 template <GemmSpecialization GemmSpec>
 using device_gemm_xdl_universal_streamk_f16_f16_f16_mk_kn_mn_comp_instances = std::tuple<
-    // clang-format off
+// clang-format off
        //#########################| ALayout| BLayout| CLayout|AData| BData| CData| AccData| Cshuffle|           A|           B|           C|          GEMM| Block|  MPer|  NPer|  KPer| AK1| BK1|MPer| NPer| MXdl| NXdl|  ABlockTransfer| ABlockTransfer| ABlockTransfer| ABlockTransfer| ABlockTransfer| ABlockTransfer| ABlockLds|  BBlockTransfer| BBlockTransfer| BBlockTransfer| BlockTransfer| BBlockTransfer| BBlockTransfer| BBlockLds|    CShuffle|    CShuffle|     CBlockTransferClusterLengths|  CBlockTransfer|                         Block-wiseGemm|               Block-wiseGemm|
        //#########################|        |        |        | Type|  Type|  Type|    Type|     Type| Elementwise| Elementwise| Elementwise|Specialization|  Size| Block| Block| Block|    |    | XDL|  XDL|  Per|  Per|   ThreadCluster|  ThreadCluster| SrcAccessOrder|   SrcVectorDim|      SrcScalar|      DstScalar| AddExtraM|   ThreadCluster|  ThreadCluster| SrcAccessOrder|  SrcVectorDim|      SrcScalar|      DstScalar| AddExtraN| MXdlPerWave| NXdlPerWave| _MBlock_MXdlPerWave_MWaveMPerXdl| ScalarPerVector|                               Pipeline|                     Pipeline|
        //#########################|        |        |        |     |      |      |        |         |   Operation|   Operation|   Operation|              |      |      |      |      |    |    |    |     | Wave| Wave| Lengths_K0_M_K1|   ArrangeOrder|               |               |      PerVector|   PerVector_K1|          | Lengths_K0_N_K1|   ArrangeOrder|               |              |      PerVector|   PerVector_K1|          |  PerShuffle|  PerShuffle| _NBlock_NXdlPerWave_NWaveNPerXdl|   _NWaveNPerXdl|                              Scheduler|                     Verision|
        //#########################|        |        |        |     |      |      |        |         |            |            |            |              |      |      |      |      |    |    |    |     |     |     |                |               |               |               |               |               |          |                |               |               |              |               |               |          |            |            |                                 |                |                                       |                             |
+#if defined(__gfx950__)
+        DeviceGemm_Xdl_CShuffle_Streamk_V3<  Row,     Row,     Row,     F16,   F16,  F16,   F32,     F16,      PassThrough, PassThrough, PassThrough,       GemmSpec,   256,   256,   32,   128,  16,   2,  32,   32,    2,    1,     S<8, 32, 1>,     S<1, 0, 2>,    S<1, 0, 2>,               2,              8,              8,          0,    S<32, 8, 1>,     S<0, 2, 1>,     S<0, 2, 1>,             1,              4,              2,          0,          1,           1,                   S<1, 32, 1, 8>,               4,  BlockGemmPipelineScheduler::Intrawave, BlockGemmPipelineVersion::v2>
+#else        
        DeviceGemm_Xdl_CShuffle_Streamk_V3<  Row,     Row,     Row,     F16,   F16,  F16,   F32,     F16,      PassThrough, PassThrough, PassThrough,       GemmSpec,   256,   256,   256,    32,   8,   4,  32,   32,    4,    4,     S<4, 64, 1>,     S<1, 0, 2>,    S<1, 0, 2>,               2,              8,              8,          0,    S<8, 32, 1>,     S<0, 2, 1>,     S<0, 2, 1>,             1,              8,              4,          0,          1,           1,                   S<1, 32, 1, 8>,               8,  BlockGemmPipelineScheduler::Intrawave, BlockGemmPipelineVersion::v4>,
        DeviceGemm_Xdl_CShuffle_Streamk_V3<  Row,     Row,     Row,     F16,   F16,  F16,   F32,     F16,      PassThrough, PassThrough, PassThrough,       GemmSpec,   256,   256,   256,    32,   4,   4,  32,   32,    4,    4,     S<8, 32, 1>,     S<1, 0, 2>,    S<1, 0, 2>,               2,              4,              4,          0,    S<8, 32, 1>,     S<0, 2, 1>,     S<0, 2, 1>,             1,              4,              4,          0,          1,           1,                   S<1, 32, 1, 8>,               8,  BlockGemmPipelineScheduler::Intrawave, BlockGemmPipelineVersion::v4>,
        DeviceGemm_Xdl_CShuffle_Streamk_V3<  Row,     Row,     Row,     F16,   F16,  F16,   F32,     F16,      PassThrough, PassThrough, PassThrough,       GemmSpec,   256,   256,   256,    32,   2,   2,  32,   32,    4,    4,     S<16,16, 1>,     S<1, 0, 2>,    S<1, 0, 2>,               2,              2,              2,          0,    S<16,16, 1>,     S<0, 2, 1>,     S<0, 2, 1>,             1,              2,              2,          0,          1,           1,                   S<1, 32, 1, 8>,               8,  BlockGemmPipelineScheduler::Intrawave, BlockGemmPipelineVersion::v4>,
@@ -52,20 +54,23 @@ using device_gemm_xdl_universal_streamk_f16_f16_f16_mk_kn_mn_comp_instances = st
        DeviceGemm_Xdl_CShuffle_Streamk_V3<  Row,     Row,     Row,     F16,   F16,  F16,   F32,     F16,      PassThrough, PassThrough, PassThrough,       GemmSpec,   256,   128,   256,    32,   8,   4,  32,   32,    2,    4,     S<4, 64, 1>,     S<1, 0, 2>,    S<1, 0, 2>,               2,              8,              8,          0,    S<8, 32, 1>,     S<0, 2, 1>,     S<0, 2, 1>,             1,              8,              4,          0,          1,           1,                   S<1, 32, 1, 8>,               8,  BlockGemmPipelineScheduler::Interwave, BlockGemmPipelineVersion::v1>,
        DeviceGemm_Xdl_CShuffle_Streamk_V3<  Row,     Row,     Row,     F16,   F16,  F16,   F32,     F16,      PassThrough, PassThrough, PassThrough,       GemmSpec,   256,   256,   128,    32,   8,   4,  32,   32,    4,    2,     S<4, 64, 1>,     S<1, 0, 2>,    S<1, 0, 2>,               2,              8,              8,          0,    S<8, 32, 1>,     S<0, 2, 1>,     S<0, 2, 1>,             1,              4,              4,          0,          1,           1,                   S<1, 32, 1, 8>,               8,  BlockGemmPipelineScheduler::Interwave, BlockGemmPipelineVersion::v1>,
        DeviceGemm_Xdl_CShuffle_Streamk_V3<  Row,     Row,     Row,     F16,   F16,  F16,   F32,     F16,      PassThrough, PassThrough, PassThrough,       GemmSpec,   256,   128,   128,    64,   8,   4,  32,   32,    2,    2,     S<8, 32, 1>,     S<1, 0, 2>,    S<1, 0, 2>,               2,              8,              8,          0,    S<16, 16, 1>,    S<0, 2, 1>,     S<0, 2, 1>,             1,              8,              4,          0,          1,           1,                   S<1, 32, 1, 8>,               8,  BlockGemmPipelineScheduler::Interwave, BlockGemmPipelineVersion::v1>,
+        DeviceGemm_Xdl_CShuffle_Streamk_V3<  Row,     Row,     Row,     F16,   F16,  F16,   F32,     F16,      PassThrough, PassThrough, PassThrough,       GemmSpec,   256,   128,   128,    64,   8,   4,  32,   32,    2,    2,     S<8, 32, 1>,     S<1, 0, 2>,    S<1, 0, 2>,               2,              8,              8,          0,    S<16, 16, 1>,    S<0, 2, 1>,     S<0, 2, 1>,             1,              8,              4,          0,          1,           1,                   S<1, 32, 1, 8>,               8,  BlockGemmPipelineScheduler::Interwave, BlockGemmPipelineVersion::v1>,
        DeviceGemm_Xdl_CShuffle_Streamk_V3<  Row,     Row,     Row,     F16,   F16,  F16,   F32,     F16,      PassThrough, PassThrough, PassThrough,       GemmSpec,   256,   128,   128,    64,   4,   4,  32,   32,    2,    2,     S<16,16, 1>,     S<1, 0, 2>,    S<1, 0, 2>,               2,              4,              4,          0,    S<16, 16, 1>,    S<0, 2, 1>,     S<0, 2, 1>,             1,              4,              4,          0,          1,           1,                   S<1, 32, 1, 8>,               8,  BlockGemmPipelineScheduler::Interwave, BlockGemmPipelineVersion::v1>,
        DeviceGemm_Xdl_CShuffle_Streamk_V3<  Row,     Row,     Row,     F16,   F16,  F16,   F32,     F16,      PassThrough, PassThrough, PassThrough,       GemmSpec,   256,   128,   128,    64,   2,   2,  32,   32,    2,    2,     S<32, 8, 1>,     S<1, 0, 2>,    S<1, 0, 2>,               2,              2,              2,          0,    S<32,  8, 1>,    S<0, 2, 1>,     S<0, 2, 1>,             1,              2,              2,          0,          1,           1,                   S<1, 32, 1, 8>,               8,  BlockGemmPipelineScheduler::Interwave, BlockGemmPipelineVersion::v1>
+#endif // defined(__gfx950__)
    // clang-format on
    >;
 template <BlockGemmPipelineScheduler BlkGemmPipeSched, GemmSpecialization GemmSpec>
 using device_gemm_xdl_universal_streamk_f16_f16_f16_mk_kn_mn_mem_instances = std::tuple<
-    // clang-format off
+// clang-format off
        //#########################| ALayout| BLayout| CLayout|AData| BData| CData| AccData| Cshuffle|           A|           B|           C|          GEMM| Block|  MPer|  NPer|  KPer| AK1| BK1|MPer| NPer| MXdl| NXdl|  ABlockTransfer| ABlockTransfer| ABlockTransfer| ABlockTransfer| ABlockTransfer| ABlockTransfer| ABlockLds|  BBlockTransfer| BBlockTransfer| BBlockTransfer| BlockTransfer| BBlockTransfer| BBlockTransfer| BBlockLds|    CShuffle|    CShuffle|     CBlockTransferClusterLengths|  CBlockTransfer|    Block-wiseGemm|               Block-wiseGemm|
        //#########################|        |        |        | Type|  Type|  Type|    Type|     Type| Elementwise| Elementwise| Elementwise|Specialization|  Size| Block| Block| Block|    |    | XDL|  XDL|  Per|  Per|   ThreadCluster|  ThreadCluster| SrcAccessOrder|   SrcVectorDim|      SrcScalar|      DstScalar| AddExtraM|   ThreadCluster|  ThreadCluster| SrcAccessOrder|  SrcVectorDim|      SrcScalar|      DstScalar| AddExtraN| MXdlPerWave| NXdlPerWave| _MBlock_MXdlPerWave_MWaveMPerXdl| ScalarPerVector|          Pipeline|                     Pipeline|
        //#########################|        |        |        |     |      |      |        |         |   Operation|   Operation|   Operation|              |      |      |      |      |    |    |    |     | Wave| Wave| Lengths_K0_M_K1|   ArrangeOrder|               |               |      PerVector|   PerVector_K1|          | Lengths_K0_N_K1|   ArrangeOrder|               |              |      PerVector|   PerVector_K1|          |  PerShuffle|  PerShuffle| _NBlock_NXdlPerWave_NWaveNPerXdl|   _NWaveNPerXdl|         Scheduler|                     Verision|
        //#########################|        |        |        |     |      |      |        |         |            |            |            |              |      |      |      |      |    |    |    |     |     |     |                |               |               |               |               |               |          |                |               |               |              |               |               |          |            |            |                                 |                |                  |                             |
+#if defined(__gfx950__)
-       // Latency friendly
+#else        
+        // Latency friendly
        DeviceGemm_Xdl_CShuffle_Streamk_V3<  Row,     Row,     Row,     F16,   F16,  F16,   F32,     F16,      PassThrough, PassThrough, PassThrough,       GemmSpec,   128,    32,   16,    64,   8,   4,  16,   16,    1,    1,     S<8, 16, 1>,     S<1, 0, 2>,    S<1, 0, 2>,               2,              8,              8,          0,    S<16, 8, 1>,     S<0, 2, 1>,     S<0, 2, 1>,             1,              2,              4,          0,          1,           1,                   S<1, 16, 1, 8>,               2,  BlkGemmPipeSched, BlockGemmPipelineVersion::v1>,
        DeviceGemm_Xdl_CShuffle_Streamk_V3<  Row,     Row,     Row,     F16,   F16,  F16,   F32,     F16,      PassThrough, PassThrough, PassThrough,       GemmSpec,   128,    32,   16,    64,   4,   4,  16,   16,    1,    1,     S<16, 8, 1>,     S<1, 0, 2>,    S<1, 0, 2>,               2,              4,              4,          0,    S<16, 8, 1>,     S<0, 2, 1>,     S<0, 2, 1>,             1,              2,              4,          0,          1,           1,                   S<1, 16, 1, 8>,               2,  BlkGemmPipeSched, BlockGemmPipelineVersion::v1>,
        DeviceGemm_Xdl_CShuffle_Streamk_V3<  Row,     Row,     Row,     F16,   F16,  F16,   F32,     F16,      PassThrough, PassThrough, PassThrough,       GemmSpec,   128,    32,   16,    64,   2,   2,  16,   16,    1,    1,     S<32, 4, 1>,     S<1, 0, 2>,    S<1, 0, 2>,               2,              2,              2,          0,    S<32, 4, 1>,     S<0, 2, 1>,     S<0, 2, 1>,             1,              2,              2,          0,          1,           1,                   S<1, 16, 1, 8>,               2,  BlkGemmPipeSched, BlockGemmPipelineVersion::v1>,
@@ -95,6 +100,7 @@ using device_gemm_xdl_universal_streamk_f16_f16_f16_mk_kn_mn_mem_instances = std
        DeviceGemm_Xdl_CShuffle_Streamk_V3<  Row,     Row,     Row,     F16,   F16,  F16,   F32,     F16,      PassThrough, PassThrough, PassThrough,       GemmSpec,   256,    16,  256,    64,   8,   4,  16,   16,    1,    4,     S<8, 16, 1>,     S<1, 0, 2>,    S<1, 0, 2>,               2,              8,              8,          0,    S<8, 32, 1>,     S<0, 2, 1>,     S<0, 2, 1>,             1,              8,              4,          0,          1,           1,                   S<1, 16, 1, 16>,              4,  BlkGemmPipeSched, BlockGemmPipelineVersion::v2>,
        DeviceGemm_Xdl_CShuffle_Streamk_V3<  Row,     Row,     Row,     F16,   F16,  F16,   F32,     F16,      PassThrough, PassThrough, PassThrough,       GemmSpec,   256,    16,  256,    64,   4,   4,  16,   16,    1,    4,     S<16, 8, 1>,     S<1, 0, 2>,    S<1, 0, 2>,               2,              4,              4,          0,    S<8, 32, 1>,     S<0, 2, 1>,     S<0, 2, 1>,             1,              4,              4,          0,          1,           1,                   S<1, 16, 1, 16>,              4,  BlkGemmPipeSched, BlockGemmPipelineVersion::v2>,
        DeviceGemm_Xdl_CShuffle_Streamk_V3<  Row,     Row,     Row,     F16,   F16,  F16,   F32,     F16,      PassThrough, PassThrough, PassThrough,       GemmSpec,   256,    32,  256,    64,   8,   4,  32,   32,    1,    2,     S<8, 32, 1>,     S<1, 0, 2>,    S<1, 0, 2>,               2,              8,              8,          0,    S<8, 32, 1>,     S<0, 2, 1>,     S<0, 2, 1>,             1,              8,              4,          0,          1,           1,                   S<1, 16, 1, 16>,              8,  BlkGemmPipeSched, BlockGemmPipelineVersion::v2>
+#endif // defined(__gfx950__)
    // clang-format on
    >;
 } // namespace instance

--- a/library/src/tensor_operation_instance/gpu/gemm_universal_streamk/device_gemm_xdl_universal_streamk_f16_f16_f16/device_gemm_xdl_universal_streamk_f16_f16_f16_mk_nk_mn.hpp
+++ b/library/src/tensor_operation_instance/gpu/gemm_universal_streamk/device_gemm_xdl_universal_streamk_f16_f16_f16/device_gemm_xdl_universal_streamk_f16_f16_f16_mk_nk_mn.hpp
@@ -34,12 +34,13 @@ static constexpr auto Interwave = BlockGemmPipelineScheduler::Interwave;
 template <GemmSpecialization GemmSpec>
 using device_gemm_xdl_universal_streamk_f16_f16_f16_mk_nk_mn_comp_instances = std::tuple<
-    // clang-format off
+// clang-format off
        //#########################| ALayout| BLayout| CLayout|AData| BData| CData| AccData| Cshuffle|           A|           B|           C|          GEMM| Block|  MPer|  NPer|  KPer| AK1| BK1|MPer| NPer| MXdl| NXdl|  ABlockTransfer| ABlockTransfer| ABlockTransfer| ABlockTransfer| ABlockTransfer| ABlockTransfer| ABlockLds|  BBlockTransfer| BBlockTransfer| BBlockTransfer| BlockTransfer| BBlockTransfer| BBlockTransfer| BBlockLds|    CShuffle|    CShuffle|     CBlockTransferClusterLengths|  CBlockTransfer|                         Block-wiseGemm|               Block-wiseGemm|
        //#########################|        |        |        | Type|  Type|  Type|    Type|     Type| Elementwise| Elementwise| Elementwise|Specialization|  Size| Block| Block| Block|    |    | XDL|  XDL|  Per|  Per|   ThreadCluster|  ThreadCluster| SrcAccessOrder|   SrcVectorDim|      SrcScalar|      DstScalar| AddExtraM|   ThreadCluster|  ThreadCluster| SrcAccessOrder|  SrcVectorDim|      SrcScalar|      DstScalar| AddExtraN| MXdlPerWave| NXdlPerWave| _MBlock_MXdlPerWave_MWaveMPerXdl| ScalarPerVector|                               Pipeline|                     Pipeline|
        //#########################|        |        |        |     |      |      |        |         |   Operation|   Operation|   Operation|              |      |      |      |      |    |    |    |     | Wave| Wave| Lengths_K0_M_K1|   ArrangeOrder|               |               |      PerVector|   PerVector_K1|          | Lengths_K0_N_K1|   ArrangeOrder|               |              |      PerVector|   PerVector_K1|          |  PerShuffle|  PerShuffle| _NBlock_NXdlPerWave_NWaveNPerXdl|   _NWaveNPerXdl|                              Scheduler|                     Verision|
        //#########################|        |        |        |     |      |      |        |         |            |            |            |              |      |      |      |      |    |    |    |     |     |     |                |               |               |               |               |               |          |                |               |               |              |               |               |          |            |            |                                 |                |                                       |                             |
+#if defined(CK_USE_AMD_MFMA_GFX950)
+#else        
        // Compute friendly
        DeviceGemm_Xdl_CShuffle_Streamk_V3<  Row,     Col,     Row,     F16,   F16,  F16,   F32,     F16,      PassThrough, PassThrough, PassThrough,       GemmSpec,   256,   256,   256,    32,   8,   8,  32,   32,    4,    4,     S<4, 64, 1>,     S<1, 0, 2>,    S<1, 0, 2>,               2,              8,              8,          0,    S<4, 64, 1>,     S<1, 0, 2>,    S<1, 0, 2>,               2,              8,              8,          0,          1,           1,                   S<1, 32, 1, 8>,               8,  BlockGemmPipelineScheduler::Intrawave, BlockGemmPipelineVersion::v4>,
        DeviceGemm_Xdl_CShuffle_Streamk_V3<  Row,     Col,     Row,     F16,   F16,  F16,   F32,     F16,      PassThrough, PassThrough, PassThrough,       GemmSpec,   256,   256,   256,    32,   4,   4,  32,   32,    4,    4,     S<8, 32, 1>,     S<1, 0, 2>,    S<1, 0, 2>,               2,              4,              4,          0,    S<8, 32, 1>,     S<1, 0, 2>,    S<1, 0, 2>,               2,              4,              4,          0,          1,           1,                   S<1, 32, 1, 8>,               8,  BlockGemmPipelineScheduler::Intrawave, BlockGemmPipelineVersion::v4>,
@@ -64,18 +65,20 @@ using device_gemm_xdl_universal_streamk_f16_f16_f16_mk_nk_mn_comp_instances = st
        DeviceGemm_Xdl_CShuffle_Streamk_V3<  Row,     Col,     Row,     F16,   F16,  F16,   F32,     F16,      PassThrough, PassThrough, PassThrough,       GemmSpec,   256,   128,   256,    32,   8,   8,  32,   32,    2,    4,     S<4, 64, 1>,     S<1, 0, 2>,    S<1, 0, 2>,               2,              8,              8,          0,    S<4, 64, 1>,     S<1, 0, 2>,    S<1, 0, 2>,               2,              8,              8,          0,          1,           1,                   S<1, 32, 1, 8>,               8,  BlockGemmPipelineScheduler::Interwave, BlockGemmPipelineVersion::v1>,
        DeviceGemm_Xdl_CShuffle_Streamk_V3<  Row,     Col,     Row,     F16,   F16,  F16,   F32,     F16,      PassThrough, PassThrough, PassThrough,       GemmSpec,   256,   256,   128,    32,   8,   8,  32,   32,    4,    2,     S<4, 64, 1>,     S<1, 0, 2>,    S<1, 0, 2>,               2,              8,              8,          0,    S<4, 64, 1>,     S<1, 0, 2>,    S<1, 0, 2>,               2,              8,              8,          0,          1,           1,                   S<1, 32, 1, 8>,               8,  BlockGemmPipelineScheduler::Interwave, BlockGemmPipelineVersion::v1>,
        DeviceGemm_Xdl_CShuffle_Streamk_V3<  Row,     Col,     Row,     F16,   F16,  F16,   F32,     F16,      PassThrough, PassThrough, PassThrough,       GemmSpec,   256,   128,   128,    64,   8,   8,  32,   32,    2,    2,     S<8, 32, 1>,     S<1, 0, 2>,    S<1, 0, 2>,               2,              8,              8,          0,    S<8, 32, 1>,     S<1, 0, 2>,    S<1, 0, 2>,               2,              8,              8,          0,          1,           1,                   S<1, 32, 1, 8>,               8,  BlockGemmPipelineScheduler::Interwave, BlockGemmPipelineVersion::v1>
+#endif // defined(CK_USE_AMD_MFMA_GFX950)
    // clang-format on
    >;
 template <BlockGemmPipelineScheduler BlkGemmPipeSched, GemmSpecialization GemmSpec>
 using device_gemm_xdl_universal_streamk_f16_f16_f16_mk_nk_mn_mem_instances = std::tuple<
-    // clang-format off
+// clang-format off
        //#########################| ALayout| BLayout| CLayout|AData| BData| CData| AccData| Cshuffle|           A|           B|           C|          GEMM| Block|  MPer|  NPer|  KPer| AK1| BK1|MPer| NPer| MXdl| NXdl|  ABlockTransfer| ABlockTransfer| ABlockTransfer| ABlockTransfer| ABlockTransfer| ABlockTransfer| ABlockLds|  BBlockTransfer| BBlockTransfer| BBlockTransfer| BlockTransfer| BBlockTransfer| BBlockTransfer| BBlockLds|    CShuffle|    CShuffle|     CBlockTransferClusterLengths|  CBlockTransfer|    Block-wiseGemm|               Block-wiseGemm|
        //#########################|        |        |        | Type|  Type|  Type|    Type|     Type| Elementwise| Elementwise| Elementwise|Specialization|  Size| Block| Block| Block|    |    | XDL|  XDL|  Per|  Per|   ThreadCluster|  ThreadCluster| SrcAccessOrder|   SrcVectorDim|      SrcScalar|      DstScalar| AddExtraM|   ThreadCluster|  ThreadCluster| SrcAccessOrder|  SrcVectorDim|      SrcScalar|      DstScalar| AddExtraN| MXdlPerWave| NXdlPerWave| _MBlock_MXdlPerWave_MWaveMPerXdl| ScalarPerVector|          Pipeline|                     Pipeline|
        //#########################|        |        |        |     |      |      |        |         |   Operation|   Operation|   Operation|              |      |      |      |      |    |    |    |     | Wave| Wave| Lengths_K0_M_K1|   ArrangeOrder|               |               |      PerVector|   PerVector_K1|          | Lengths_K0_N_K1|   ArrangeOrder|               |              |      PerVector|   PerVector_K1|          |  PerShuffle|  PerShuffle| _NBlock_NXdlPerWave_NWaveNPerXdl|   _NWaveNPerXdl|         Scheduler|                     Verision|
        //#########################|        |        |        |     |      |      |        |         |            |            |            |              |      |      |      |      |    |    |    |     |     |     |                |               |               |               |               |               |          |                |               |               |              |               |               |          |            |            |                                 |                |                  |                             |
+#if defined(CK_USE_AMD_MFMA_GFX950)
-       // Latency friendly 
+#else        
+        // Latency friendly 
        DeviceGemm_Xdl_CShuffle_Streamk_V3<  Row,     Col,     Row,     F16,   F16,  F16,   F32,     F16,      PassThrough, PassThrough, PassThrough,       GemmSpec,   128,    32,   16,    64,   8,   8,  16,   16,    1,    1,     S<8, 16, 1>,     S<1, 0, 2>,    S<1, 0, 2>,               2,              8,              8,          0,    S<8, 16, 1>,     S<1, 0, 2>,    S<1, 0, 2>,               2,              8,              8,          0,          1,           1,                   S<1, 16, 1, 8>,               2,  BlkGemmPipeSched, BlockGemmPipelineVersion::v1>,
        DeviceGemm_Xdl_CShuffle_Streamk_V3<  Row,     Col,     Row,     F16,   F16,  F16,   F32,     F16,      PassThrough, PassThrough, PassThrough,       GemmSpec,   128,    32,   16,    64,   4,   4,  16,   16,    1,    1,     S<16, 8, 1>,     S<1, 0, 2>,    S<1, 0, 2>,               2,              4,              4,          0,    S<16, 8, 1>,     S<1, 0, 2>,    S<1, 0, 2>,               2,              4,              4,          0,          1,           1,                   S<1, 16, 1, 8>,               2,  BlkGemmPipeSched, BlockGemmPipelineVersion::v1>,
        DeviceGemm_Xdl_CShuffle_Streamk_V3<  Row,     Col,     Row,     F16,   F16,  F16,   F32,     F16,      PassThrough, PassThrough, PassThrough,       GemmSpec,   128,    32,   16,    64,   2,   2,  16,   16,    1,    1,     S<32, 4, 1>,     S<1, 0, 2>,    S<1, 0, 2>,               2,              2,              2,          0,    S<32, 4, 1>,     S<1, 0, 2>,    S<1, 0, 2>,               2,              2,              2,          0,          1,           1,                   S<1, 16, 1, 8>,               2,  BlkGemmPipeSched, BlockGemmPipelineVersion::v1>,
@@ -109,6 +112,7 @@ using device_gemm_xdl_universal_streamk_f16_f16_f16_mk_nk_mn_mem_instances = std
        DeviceGemm_Xdl_CShuffle_Streamk_V3<  Row,     Col,     Row,     F16,   F16,  F16,   F32,     F16,      PassThrough, PassThrough, PassThrough,       GemmSpec,   256,    32,  256,    64,   8,   8,  32,   32,    1,    2,     S<8, 32, 1>,     S<1, 0, 2>,    S<1, 0, 2>,               2,              8,              8,          0,    S<8, 32, 1>,     S<1, 0, 2>,    S<1, 0, 2>,               2,              8,              8,          0,          1,           1,                   S<1, 16, 1, 16>,              8,  BlkGemmPipeSched, BlockGemmPipelineVersion::v2>,
        DeviceGemm_Xdl_CShuffle_Streamk_V3<  Row,     Col,     Row,     F16,   F16,  F16,   F32,     F16,      PassThrough, PassThrough, PassThrough,       GemmSpec,   256,    32,  256,    64,   4,   4,  32,   32,    1,    2,     S<16,16, 1>,     S<1, 0, 2>,    S<1, 0, 2>,               2,              4,              4,          0,    S<16,16, 1>,     S<1, 0, 2>,    S<1, 0, 2>,               2,              4,              4,          0,          1,           1,                   S<1, 16, 1, 16>,              8,  BlkGemmPipeSched, BlockGemmPipelineVersion::v2>,
        DeviceGemm_Xdl_CShuffle_Streamk_V3<  Row,     Col,     Row,     F16,   F16,  F16,   F32,     F16,      PassThrough, PassThrough, PassThrough,       GemmSpec,   256,    32,  256,    64,   2,   2,  32,   32,    1,    2,     S<32, 8, 1>,     S<1, 0, 2>,    S<1, 0, 2>,               2,              2,              2,          0,    S<32, 8, 1>,     S<1, 0, 2>,    S<1, 0, 2>,               2,              2,              2,          0,          1,           1,                   S<1, 16, 1, 16>,              8,  BlkGemmPipeSched, BlockGemmPipelineVersion::v2>
+#endif // defined(CK_USE_AMD_MFMA_GFX950)
    // clang-format on
    >;
 } // namespace instance

--- a/library/src/tensor_operation_instance/gpu/grouped_gemm/device_grouped_gemm_xdl_splitk_f16_f8_f16_mk_kn_mn_irregular_instance.cpp
+++ b/library/src/tensor_operation_instance/gpu/grouped_gemm/device_grouped_gemm_xdl_splitk_f16_f8_f16_mk_kn_mn_irregular_instance.cpp
@@ -27,15 +27,18 @@ using S = ck::Sequence<Is...>;
 using Empty_Tuple = ck::Tuple<>;
-using PassThrough                    = ck::tensor_operation::element_wise::PassThrough;
+using PassThrough = ck::tensor_operation::element_wise::PassThrough;
+#if !defined(CK_USE_AMD_MFMA_GFX950)
 static constexpr auto GemmMNKPadding = ck::tensor_operation::device::GemmSpecialization::MNKPadding;
+#endif
 using device_grouped_gemm_xdl_splitk_f16_f8_f16_mk_kn_mn_irregular_tile_instances = std::tuple<
-    // clang-format off
+// clang-format off
        //################################|      A|      B|          Ds|      E| AData| BData| AccData| CShuffle|      DsData| EData|           A|           B|           C|           GEMM| NumGemmK| Block|  MPer|  NPer|  KPer| AK1| BK1| MPer| NPer| MXdl| NXdl|  ABlockTransfer| ABlockTransfer| ABlockTransfer| ABlockTransfer| ABlockTransfer| ABlockTransfer| ABlockLds|  BBlockTransfer| BBlockTransfer| BBlockTransfer| BlockTransfer| BBlockTransfer| BBlockTransfer| BBlockLds|    CShuffle|    CShuffle| CBlockTransferClusterLengths|  CBlockTransfer|
        //################################| Layout| Layout|      Layout| Layout|  Type|  Type|    Type| DataType|        Type|  Type| Elementwise| Elementwise| Elementwise| Spacialization| Prefetch|  Size| Block| Block| Block|    |    |  XDL|  XDL|  Per|  Per|   ThreadCluster|  ThreadCluster| SrcAccessOrder|   SrcVectorDim|      SrcScalar|      DstScalar| AddExtraM|   ThreadCluster|  ThreadCluster| SrcAccessOrder|  SrcVectorDim|      SrcScalar|      DstScalar| AddExtraN| MXdlPerWave| NXdlPerWave|         _MBlock_MWaveMPerXdl| ScalarPerVector|
        //################################|       |       |            |       |      |      |        |         |            |      |   Operation|   Operation|   Operation|               |    Stage|      |      |      |      |    |    |     |     | Wave| Wave| Lengths_K0_M_K1|   ArrangeOrder|               |               |      PerVector|   PerVector_K1|          | Lengths_K0_N_K1|   ArrangeOrder|               |              |      PerVector|   PerVector_K1|          |  PerShuffle|  PerShuffle|         _NBlock_NWaveNPerXdl|   _NWaveNPerXdl|
        //################################|       |       |            |       |      |      |        |         |            |      |            |            |            |               |         |      |      |      |      |    |    |     |     |     |     |                |               |               |               |               |               |          |                |               |               |              |               |               |          |            |            |                             |                |
+#if defined(CK_USE_AMD_MFMA_GFX950)
+#else
        DeviceGroupedGemmXdlSplitKCShuffle<    Row,    Row, Empty_Tuple,    Row,   F16,    F8,     F32,      F16, Empty_Tuple,   F16, PassThrough, PassThrough, PassThrough, GemmMNKPadding,        1,   256,   256,   128,    32,   8,   8,   32,   32,    4,    2,  S<1, 4, 64, 1>,  S<0, 2, 1, 3>,  S<0, 2, 1, 3>,              3,              8,              8,         1,  S<1, 4, 64, 1>,  S<0, 1, 3, 2>,  S<0, 1, 3, 2>,             2,              2,              8,         1,           1,           1,               S<1, 32, 1, 8>,              8, PipelineVersion::v1>,
        DeviceGroupedGemmXdlSplitKCShuffle<    Row,    Row, Empty_Tuple,    Row,   F16,    F8,     F32,      F16, Empty_Tuple,   F16, PassThrough, PassThrough, PassThrough, GemmMNKPadding,        1,   256,   128,   256,    32,   8,   8,   32,   32,    2,    4,  S<1, 4, 64, 1>,  S<0, 2, 1, 3>,  S<0, 2, 1, 3>,              3,              8,              8,         1,  S<1, 4, 64, 1>,  S<0, 1, 3, 2>,  S<0, 1, 3, 2>,             2,              4,              8,         1,           1,           1,               S<1, 32, 1, 8>,              8, PipelineVersion::v1>,
        DeviceGroupedGemmXdlSplitKCShuffle<    Row,    Row, Empty_Tuple,    Row,   F16,    F8,     F32,      F16, Empty_Tuple,   F16, PassThrough, PassThrough, PassThrough, GemmMNKPadding,        1,   256,   192,    64,    32,   8,   8,   32,   32,    3,    1,  S<1, 4, 64, 1>,  S<0, 2, 1, 3>,  S<0, 2, 1, 3>,              3,              8,              8,         1,  S<1, 4, 32, 1>,  S<0, 1, 3, 2>,  S<0, 1, 3, 2>,             2,              2,              8,         1,           1,           1,               S<1, 32, 1, 8>,              8, PipelineVersion::v1>,
@@ -98,6 +101,7 @@ using device_grouped_gemm_xdl_splitk_f16_f8_f16_mk_kn_mn_irregular_tile_instance
        DeviceGroupedGemmXdlSplitKCShuffle<    Row,    Row, Empty_Tuple,    Row,   F16,    F8,     F32,      F16, Empty_Tuple,   F16, PassThrough, PassThrough, PassThrough, GemmMNKPadding,        1,    64,    64,    64,    32,   8,   8,   32,   32,    2,    2,  S<1, 4, 16, 1>,  S<0, 2, 1, 3>,  S<0, 2, 1, 3>,              3,              8,              8,         1,  S<1, 4, 16, 1>,  S<0, 1, 3, 2>,  S<0, 1, 3, 2>,             2,              2,              8,         1,           1,           1,               S<1, 16, 1, 4>,              8, PipelineVersion::v2>,
        DeviceGroupedGemmXdlSplitKCShuffle<    Row,    Row, Empty_Tuple,    Row,   F16,    F8,     F32,      F16, Empty_Tuple,   F16, PassThrough, PassThrough, PassThrough, GemmMNKPadding,        1,    64,    64,    32,    32,   8,   8,   32,   32,    2,    1,  S<1, 4, 16, 1>,  S<0, 2, 1, 3>,  S<0, 2, 1, 3>,              3,              8,              8,         1,  S<1, 4, 16, 1>,  S<0, 1, 3, 2>,  S<0, 1, 3, 2>,             2,              2,              8,         1,           1,           1,               S<1, 16, 1, 4>,              8, PipelineVersion::v2>,
        DeviceGroupedGemmXdlSplitKCShuffle<    Row,    Row, Empty_Tuple,    Row,   F16,    F8,     F32,      F16, Empty_Tuple,   F16, PassThrough, PassThrough, PassThrough, GemmMNKPadding,        1,    64,    32,    64,    32,   8,   8,   32,   32,    1,    2,  S<1, 4, 16, 1>,  S<0, 2, 1, 3>,  S<0, 2, 1, 3>,              3,              8,              8,         1,  S<1, 4, 16, 1>,  S<0, 1, 3, 2>,  S<0, 1, 3, 2>,             2,              2,              8,         1,           1,           1,               S<1, 16, 1, 4>,              8, PipelineVersion::v2>
+#endif // defined(CK_USE_AMD_MFMA_GFX950)
    // clang-format on
    >;

--- a/library/src/tensor_operation_instance/gpu/grouped_gemm_tile_loop/device_grouped_gemm_xdl_tile_loop_multiply_bf16_i8_bf16_mk_kn_mn.hpp
+++ b/library/src/tensor_operation_instance/gpu/grouped_gemm_tile_loop/device_grouped_gemm_xdl_tile_loop_multiply_bf16_i8_bf16_mk_kn_mn.hpp
@@ -44,11 +44,14 @@ template <typename DsLayout,
          typename CDEElementwiseOp,
          GemmSpecialization GemmSpec = GemmMNKPadding>
 using device_grouped_gemm_xdl_tile_loop_bf16_i8_bf16_mk_kn_mn_comp_instances = std::tuple<
-    // clang-format off
+// clang-format off
        //###########################################|      A|      B|          Ds|      E| AData| BData| AccData| CShuffle|      DsData| EData|           A|           B|                C|           GEMM| NumGemmK| Block|  MPer|  NPer|  KPer| AK1| BK1| MPer| NPer| MXdl| NXdl|  ABlockTransfer| ABlockTransfer| ABlockTransfer| ABlockTransfer| ABlockTransfer| ABlockTransfer| ABlockLds|  BBlockTransfer| BBlockTransfer| BBlockTransfer| BlockTransfer| BBlockTransfer| BBlockTransfer| BBlockLds|    CShuffle|    CShuffle| CBlockTransferClusterLengths|  CBlockTransfer|
        //###########################################| Layout| Layout|      Layout| Layout|  Type|  Type|    Type| DataType|        Type|  Type| Elementwise| Elementwise|      Elementwise| Spacialization| Prefetch|  Size| Block| Block| Block|    |    |  XDL|  XDL|  Per|  Per|   ThreadCluster|  ThreadCluster| SrcAccessOrder|   SrcVectorDim|      SrcScalar|      DstScalar| AddExtraM|   ThreadCluster|  ThreadCluster| SrcAccessOrder|  SrcVectorDim|      SrcScalar|      DstScalar| AddExtraN| MXdlPerWave| NXdlPerWave|         _MBlock_MWaveMPerXdl| ScalarPerVector|
        //###########################################|       |       |            |       |      |      |        |         |            |      |   Operation|   Operation|        Operation|               |    Stage|      |      |      |      |    |    |     |     | Wave| Wave| Lengths_K0_M_K1|   ArrangeOrder|               |               |      PerVector|   PerVector_K1|          | Lengths_K0_N_K1|   ArrangeOrder|               |              |      PerVector|   PerVector_K1|          |  PerShuffle|  PerShuffle|         _NBlock_NWaveNPerXdl|   _NWaveNPerXdl|
        //###########################################|       |       |            |       |      |      |        |         |            |      |            |            |                 |               |         |      |      |      |      |    |    |     |     |     |     |                |               |               |               |               |               |          |                |               |               |              |               |               |          |            |            |                             |   S<C,D0...,D_N| 
+#if defined(CK_USE_AMD_MFMA_GFX950)
+        DeviceGroupedGemmMultipleDXdlCShuffleTileLoop<   Row,     Row,    DsLayout,    Row,  BF16,    I8,     F32,      F32,  DsDataType,  BF16, PassThrough, PassThrough, CDEElementwiseOp,       GemmSpec,        1,   256,   128,   128,    128,  16,   4,   32,   32,    2,    2,     S<8, 32, 1>,     S<1, 0, 2>,     S<1, 0, 2>,              2,              8,              8,         0,    S<16, 16, 1>,     S<0, 2, 1>,     S<0, 2, 1>,             1,              8,              4,         0,           1,           1,               S<1, 32, 1, 8>,        S<8,8,8>,  BlockGemmPipelineScheduler::Intrawave, BlockGemmPipelineVersion::v1>
+#else
        // DeviceGroupedGemmMultipleDXdlCShuffleTileLoop<   Row,     Row,    DsLayout,    Row,  BF16,    I8,     F32,      F32,  DsDataType,  BF16, PassThrough, PassThrough, CDEElementwiseOp,       GemmSpec,        1,   256,   256,   256,    32,   8,   4,   32,   32,    4,    4,     S<4, 64, 1>,     S<1, 0, 2>,     S<1, 0, 2>,              2,              8,              8,         0,     S<8, 32, 1>,     S<0, 2, 1>,     S<0, 2, 1>,             1,              8,              4,         0,           1,           1,               S<1, 32, 1, 8>,        S<8,8,1>,  BlockGemmPipelineScheduler::Intrawave, BlockGemmPipelineVersion::v4>,
        // DeviceGroupedGemmMultipleDXdlCShuffleTileLoop<   Row,     Row,    DsLayout,    Row,  BF16,    I8,     F32,      F32,  DsDataType,  BF16, PassThrough, PassThrough, CDEElementwiseOp,       GemmSpec,        1,   256,   128,   128,    64,   8,   4,   32,   32,    2,    2,     S<8, 32, 1>,     S<1, 0, 2>,     S<1, 0, 2>,              2,              8,              8,         0,    S<16, 16, 1>,     S<0, 2, 1>,     S<0, 2, 1>,             1,              8,              4,         0,           1,           1,               S<1, 32, 1, 8>,        S<8,8,1>,  BlockGemmPipelineScheduler::Intrawave, BlockGemmPipelineVersion::v4>,
        // DeviceGroupedGemmMultipleDXdlCShuffleTileLoop<   Row,     Row,    DsLayout,    Row,  BF16,    I8,     F32,      F32,  DsDataType,  BF16, PassThrough, PassThrough, CDEElementwiseOp,       GemmSpec,        1,   256,   256,   256,    32,   8,   4,   32,   32,    4,    4,     S<4, 64, 1>,     S<1, 0, 2>,     S<1, 0, 2>,              2,              8,              8,         0,     S<8, 32, 1>,     S<0, 2, 1>,     S<0, 2, 1>,             1,              8,              4,         0,           1,           1,               S<1, 32, 1, 8>,        S<8,8,1>,  BlockGemmPipelineScheduler::Intrawave, BlockGemmPipelineVersion::v5>,
@@ -57,6 +60,7 @@ using device_grouped_gemm_xdl_tile_loop_bf16_i8_bf16_mk_kn_mn_comp_instances = s
        // DeviceGroupedGemmMultipleDXdlCShuffleTileLoop<   Row,     Row,    DsLayout,    Row,  BF16,    I8,     F32,      F32,  DsDataType,  BF16, PassThrough, PassThrough, CDEElementwiseOp,       GemmSpec,        1,   256,   128,   128,    64,   8,   4,   32,   32,    2,    2,     S<8, 32, 1>,     S<1, 0, 2>,     S<1, 0, 2>,              2,              8,              8,         0,    S<16, 16, 1>,     S<0, 2, 1>,     S<0, 2, 1>,             1,              8,              4,         0,           1,           1,               S<1, 32, 1, 8>,        S<8,8,1>,  BlockGemmPipelineScheduler::Intrawave, BlockGemmPipelineVersion::v3>,
        // DeviceGroupedGemmMultipleDXdlCShuffleTileLoop<   Row,     Row,    DsLayout,    Row,  BF16,    I8,     F32,      F32,  DsDataType,  BF16, PassThrough, PassThrough, CDEElementwiseOp,       GemmSpec,        1,   256,   128,   256,    32,   8,   4,   32,   32,    2,    4,     S<4, 64, 1>,     S<1, 0, 2>,     S<1, 0, 2>,              2,              8,              8,         0,     S<8, 32, 1>,     S<0, 2, 1>,     S<0, 2, 1>,             1,              8,              4,         0,           1,           1,               S<1, 32, 1, 8>,        S<8,8,1>,  BlockGemmPipelineScheduler::Interwave, BlockGemmPipelineVersion::v1>,
        DeviceGroupedGemmMultipleDXdlCShuffleTileLoop<   Row,     Row,    DsLayout,    Row,  BF16,    I8,     F32,      F32,  DsDataType,  BF16, PassThrough, PassThrough, CDEElementwiseOp,       GemmSpec,        1,   256,   128,   128,    64,   8,   4,   32,   32,    2,    2,     S<8, 32, 1>,     S<1, 0, 2>,     S<1, 0, 2>,              2,              8,              8,         0,    S<16, 16, 1>,     S<0, 2, 1>,     S<0, 2, 1>,             1,              8,              4,         0,           1,           1,               S<1, 32, 1, 8>,        S<8,8,8>,  BlockGemmPipelineScheduler::Intrawave, BlockGemmPipelineVersion::v1>
+#endif // defined(CK_USE_AMD_MFMA_GFX950)
    // clang-format on
    >;
@@ -65,13 +69,14 @@ template <typename DsLayout,
          typename CDEElementwiseOp,
          GemmSpecialization GemmSpec                 = GemmMNKPadding,
          BlockGemmPipelineScheduler BlkGemmPipeSched = BlockGemmPipelineScheduler::Intrawave>
-using device_grouped_gemm_xdl_tile_loop_bf16_i8_bf16_mk_kn_mn_mem_instances =
+using device_grouped_gemm_xdl_tile_loop_bf16_i8_bf16_mk_kn_mn_mem_instances = std::tuple<
-    std::tuple<
+// clang-format off
-        // clang-format off
        //###########################################|      A|      B|          Ds|      E| AData| BData| AccData| CShuffle|      DsData| EData|           A|           B|                C|           GEMM| NumGemmK| Block|  MPer|  NPer|  KPer| AK1| BK1| MPer| NPer| MXdl| NXdl|  ABlockTransfer| ABlockTransfer| ABlockTransfer| ABlockTransfer| ABlockTransfer| ABlockTransfer| ABlockLds|  BBlockTransfer| BBlockTransfer| BBlockTransfer| BlockTransfer| BBlockTransfer| BBlockTransfer| BBlockLds|    CShuffle|    CShuffle| CBlockTransferClusterLengths|  CBlockTransfer|
        //###########################################| Layout| Layout|      Layout| Layout|  Type|  Type|    Type| DataType|        Type|  Type| Elementwise| Elementwise|      Elementwise| Spacialization| Prefetch|  Size| Block| Block| Block|    |    |  XDL|  XDL|  Per|  Per|   ThreadCluster|  ThreadCluster| SrcAccessOrder|   SrcVectorDim|      SrcScalar|      DstScalar| AddExtraM|   ThreadCluster|  ThreadCluster| SrcAccessOrder|  SrcVectorDim|      SrcScalar|      DstScalar| AddExtraN| MXdlPerWave| NXdlPerWave|         _MBlock_MWaveMPerXdl| ScalarPerVector|
        //###########################################|       |       |            |       |      |      |        |         |            |      |   Operation|   Operation|        Operation|               |    Stage|      |      |      |      |    |    |     |     | Wave| Wave| Lengths_K0_M_K1|   ArrangeOrder|               |               |      PerVector|   PerVector_K1|          | Lengths_K0_N_K1|   ArrangeOrder|               |              |      PerVector|   PerVector_K1|          |  PerShuffle|  PerShuffle|         _NBlock_NWaveNPerXdl|   _NWaveNPerXdl|
        //###########################################|       |       |            |       |      |      |        |         |            |      |            |            |                 |               |         |      |      |      |      |    |    |     |     |     |     |                |               |               |               |               |               |          |                |               |               |              |               |               |          |            |            |                             |   S<C,D0...,D_N| 
+#if defined(CK_USE_AMD_MFMA_GFX950)
+#else
        // Latency friendly
        // DeviceGroupedGemmMultipleDXdlCShuffleTileLoop<   Row,     Row,    DsLayout,    Row,  BF16,    I8,     F32,      F32,  DsDataType,  BF16, PassThrough, PassThrough, CDEElementwiseOp,       GemmSpec,        1,    64,    16,    16,   256,   8,   4,   16,   16,    1,    1,     S<32, 2, 1>,     S<1, 0, 2>,     S<1, 0, 2>,              2,              8,              8,         0,     S<64, 1, 1>,     S<0, 2, 1>,     S<0, 2, 1>,             1,             16,              4,          0,          1,           1,               S<1, 16, 1, 4>,        S<4,4,1>,  BlkGemmPipeSched, BlockGemmPipelineVersion::v1>,
        // DeviceGroupedGemmMultipleDXdlCShuffleTileLoop<   Row,     Row,    DsLayout,    Row,  BF16,    I8,     F32,      F32,  DsDataType,  BF16, PassThrough, PassThrough, CDEElementwiseOp,       GemmSpec,        1,   128,    16,    32,   256,   8,   4,   16,   16,    1,    1,     S<32, 4, 1>,     S<1, 0, 2>,     S<1, 0, 2>,              2,              8,              8,         0,     S<64, 2, 1>,     S<0, 2, 1>,     S<0, 2, 1>,             1,             16,              4,          0,          1,           1,               S<1, 16, 1, 8>,        S<4,4,1>,  BlkGemmPipeSched, BlockGemmPipelineVersion::v1>,
@@ -84,8 +89,9 @@ using device_grouped_gemm_xdl_tile_loop_bf16_i8_bf16_mk_kn_mn_mem_instances =
        // DeviceGroupedGemmMultipleDXdlCShuffleTileLoop<   Row,     Row,    DsLayout,    Row,  BF16,    I8,     F32,      F32,  DsDataType,  BF16, PassThrough, PassThrough, CDEElementwiseOp,       GemmSpec,        1,   128,    32,   128,    64,   8,   4,   32,   32,    1,    2,     S<8, 16, 1>,     S<1, 0, 2>,     S<1, 0, 2>,              2,              8,              8,         0,     S<16, 8, 1>,     S<0, 2, 1>,     S<0, 2, 1>,             1,             16,              4,          0,          1,           1,               S<1, 16, 1, 8>,        S<8,8,1>,  BlkGemmPipeSched, BlockGemmPipelineVersion::v2>,
        // DeviceGroupedGemmMultipleDXdlCShuffleTileLoop<   Row,     Row,    DsLayout,    Row,  BF16,    I8,     F32,      F32,  DsDataType,  BF16, PassThrough, PassThrough, CDEElementwiseOp,       GemmSpec,        1,   256,    16,   256,    64,   8,   4,   16,   16,    1,    4,     S<8, 16, 1>,     S<1, 0, 2>,     S<1, 0, 2>,              2,              8,              8,         0,    S<16, 16, 1>,     S<0, 2, 1>,     S<0, 2, 1>,             1,             16,              4,          0,          1,           1,              S<1, 16, 1, 16>,        S<4,4,1>,  BlkGemmPipeSched, BlockGemmPipelineVersion::v2>,
        // DeviceGroupedGemmMultipleDXdlCShuffleTileLoop<   Row,     Row,    DsLayout,    Row,  BF16,    I8,     F32,      F32,  DsDataType,  BF16, PassThrough, PassThrough, CDEElementwiseOp,       GemmSpec,        1,   256,    32,   256,    64,   8,   4,   32,   32,    1,    2,     S<8, 32, 1>,     S<1, 0, 2>,     S<1, 0, 2>,              2,              8,              8,         0,    S<16, 16, 1>,     S<0, 2, 1>,     S<0, 2, 1>,             1,             16,              4,          0,          1,           1,              S<1, 16, 1, 16>,        S<8,8,1>,  BlkGemmPipeSched, BlockGemmPipelineVersion::v2>
-        // clang-format on
+#endif // defined(CK_USE_AMD_MFMA_GFX950)
-        >;
+       // clang-format on
+    >;
 } // namespace instance
 } // namespace device

--- a/test/CMakeLists.txt
+++ b/test/CMakeLists.txt
@@ -100,7 +100,7 @@ function(add_test_executable TEST_NAME)
        if(ARGN MATCHES "_xdl")
             list(REMOVE_ITEM TEST_TARGETS gfx900 gfx906 gfx906:xnack- gfx1030 gfx1100 gfx1101 gfx1102 gfx1103 gfx1200 gfx1201 gfx10.3-generic gfx11-generic gfx12-generic)
        elseif(ARGN MATCHES "_wmma")
-             list(REMOVE_ITEM TEST_TARGETS gfx900 gfx906 gfx906:xnack- gfx908:xnack+ gfx908:xnack- gfx90a:xnack+ gfx90a:xnack- gfx908 gfx90a gfx940 gfx941 gfx942 gfx1030)
+             list(REMOVE_ITEM TEST_TARGETS gfx900 gfx906 gfx906:xnack- gfx908:xnack+ gfx908:xnack- gfx90a:xnack+ gfx90a:xnack- gfx908 gfx90a gfx940 gfx941 gfx942 gfx1030 gfx950)
        elseif(ARGN MATCHES "_smfmac")
             list(REMOVE_ITEM TEST_TARGETS gfx900 gfx906 gfx906:xnack- gfx1030 gfx1100 gfx1101 gfx1102 gfx1103 gfx908 gfx90a gfx1200 gfx1201 gfx10.3-generic gfx11-generic gfx12-generic)
        endif()
@@ -169,26 +169,38 @@ function(add_gtest_executable TEST_NAME)
            list(REMOVE_ITEM ARGN "${source}")
        endif()
    endforeach()
    foreach(source IN LISTS ARGN)
        if(NOT TEST_TARGETS MATCHES "gfx9" AND source MATCHES "xdl")
            message("removing xdl test ${source} ")
            list(REMOVE_ITEM ARGN "${source}")
        endif()
    endforeach()
+    foreach(source IN LISTS ARGN)
+    if(NOT TEST_TARGETS MATCHES "gfx95" AND source MATCHES "mx_")
+        message("removing microscaling test ${source} ")
+        list(REMOVE_ITEM ARGN "${source}")
+    endif()
+    endforeach()
    foreach(source IN LISTS ARGN)
 	if(NOT TEST_TARGETS MATCHES "gfx11" AND NOT TEST_TARGETS MATCHES "gfx12" AND source MATCHES "wmma")
            message("removing wmma test ${source} ")
            list(REMOVE_ITEM ARGN "${source}")
        endif()
    endforeach()
    #only continue if there are some source files left on the list
    if(ARGN)
        if(ARGN MATCHES "_xdl")
             list(REMOVE_ITEM TEST_TARGETS gfx900 gfx906 gfx906:xnack- gfx1030 gfx1100 gfx1101 gfx1102 gfx1103 gfx1200 gfx1201 gfx10.3-generic gfx11-generic gfx12-generic)
        elseif(ARGN MATCHES "_wmma")
-             list(REMOVE_ITEM TEST_TARGETS gfx900 gfx906 gfx906:xnack- gfx908:xnack+ gfx908:xnack- gfx90a:xnack+ gfx90a:xnack- gfx908 gfx90a gfx940 gfx941 gfx942 gfx1030)
+             list(REMOVE_ITEM TEST_TARGETS gfx900 gfx906 gfx906:xnack- gfx908:xnack+ gfx908:xnack- gfx90a:xnack+ gfx90a:xnack- gfx908 gfx90a gfx940 gfx941 gfx942 gfx1030 gfx950)
        elseif(ARGN MATCHES "_smfmac")
             list(REMOVE_ITEM TEST_TARGETS gfx900 gfx906 gfx906:xnack- gfx1030 gfx1100 gfx1101 gfx1102 gfx1103 gfx908 gfx90a gfx1200 gfx1201 gfx10.3-generic gfx11-generic gfx12-generic)
+        elseif(ARGN MATCHES "_mx") #only build mx example for gfx950
+             list(REMOVE_ITEM TEST_TARGETS gfx900 gfx906 gfx906:xnack- gfx908:xnack+ gfx908:xnack- gfx90a:xnack+ gfx90a:xnack- gfx908 gfx90a gfx940 gfx941 gfx942 gfx1030 gfx1100 gfx1101 gfx1102 gfx1103 gfx1200 gfx1201 gfx10.3-generic gfx11-generic gfx12-generic)
        endif()
        set_source_files_properties(${ARGN} PROPERTIES LANGUAGE HIP)
        add_executable(${TEST_NAME} ${ARGN})
@@ -258,8 +270,11 @@ add_subdirectory(wrapper)
 if(SUPPORTED_GPU_TARGETS MATCHES "gfx11")
    add_subdirectory(wmma_op)
 endif()
-if(SUPPORTED_GPU_TARGETS MATCHES "gfx942" AND CK_HIP_VERSION_MAJOR GREATER_EQUAL 6 AND CK_HIP_VERSION_MINOR GREATER_EQUAL 2) # smfmac needs ROCm6.2
+if(SUPPORTED_GPU_TARGETS MATCHES "gfx942" OR SUPPORTED_GPU_TARGETS MATCHES "gfx950") # smfmac needs ROCm6.2
    add_subdirectory(smfmac_op)
 endif()
+if(SUPPORTED_GPU_TARGETS MATCHES "gfx950") 
+    add_subdirectory(mx_mfma_op)
+endif()
 add_subdirectory(position_embedding)
 add_subdirectory(scatter_gather)
--- a/test/data_type/CMakeLists.txt
+++ b/test/data_type/CMakeLists.txt
@@ -9,11 +9,10 @@ if (USE_BITINT_EXTENSION_INT4)
  endif()
 endif()
 add_custom_target(test_fp8)
 if (CK_USE_OCP_FP8)
+  # add test for ocp data types
  add_gtest_executable(test_fp8_ocp test_fp8_ocp.cpp)
  if(result EQUAL 0)
    target_link_libraries(test_fp8_ocp PRIVATE utility)
@@ -43,6 +42,45 @@ if (CK_USE_FNUZ_FP8)
  add_dependencies(test_fp8 test_bf8_fnuz)
 endif()
+if(GPU_TARGETS MATCHES "gfx950")
+  add_custom_target(test_mx_data_types)
+  add_gtest_executable(test_fp4 test_fp4.cpp)
+  if(result EQUAL 0)
+    target_link_libraries(test_fp4 PRIVATE utility)
+  endif()
+  add_dependencies(test_mx_data_types test_fp4)
+  add_gtest_executable(test_fp6 test_fp6.cpp)
+  if(result EQUAL 0)
+    target_link_libraries(test_fp6 PRIVATE utility)
+  endif()
+  add_dependencies(test_mx_data_types test_fp6)
+  add_gtest_executable(test_bf6 test_bf6.cpp)
+  if(result EQUAL 0)
+    target_link_libraries(test_bf6 PRIVATE utility)
+  endif()
+  add_dependencies(test_mx_data_types test_bf6)
+  add_gtest_executable(test_mx_fp8 test_mx_fp8.cpp)
+  if(result EQUAL 0)
+    target_link_libraries(test_mx_fp8 PRIVATE utility)
+  endif()
+  add_dependencies(test_mx_data_types test_mx_fp8)
+  add_gtest_executable(test_mx_bf8 test_mx_bf8.cpp)
+  if(result EQUAL 0)
+    target_link_libraries(test_mx_bf8 PRIVATE utility)
+  endif()
+  add_dependencies(test_mx_data_types test_mx_bf8)
+  add_gtest_executable(test_e8m0 test_e8m0.cpp)
+  if(result EQUAL 0)
+    target_link_libraries(test_e8m0 PRIVATE utility)
+  endif()
+  add_dependencies(test_mx_data_types test_e8m0)
+endif()
 add_gtest_executable(test_custom_type test_custom_type.cpp)
 if(result EQUAL 0)
  target_link_libraries(test_custom_type PRIVATE utility)

--- a/test/data_type/test_bf6.cpp
+++ b/test/data_type/test_bf6.cpp
+// SPDX-License-Identifier: MIT
+// Copyright (c) 2025, Advanced Micro Devices, Inc. All rights reserved.
+#include "gtest/gtest.h"
+#include "ck/utility/data_type.hpp"
+#include "ck/utility/type_convert.hpp"
+#include "ck/utility/scaled_type_convert.hpp"
+using ck::bf6_convert_rne;
+using ck::bf6_convert_sr;
+using ck::bf6_t;
+using ck::bf6x16_pk_t;
+using ck::bf6x32_pk_t;
+using ck::e8m0_bexp_t;
+using ck::Number;
+using ck::scaled_type_convert;
+using ck::type_convert;
+using ck::vector_type;
+TEST(BF6, NumericLimits)
+{
+    EXPECT_EQ(ck::NumericLimits<bf6_t>::Min(), bf6_t(0b001000));
+    EXPECT_EQ(ck::NumericLimits<bf6_t>::Max(), bf6_t(0b011111));
+    EXPECT_EQ(ck::NumericLimits<bf6_t>::Lowest(), bf6_t(0b111111));
+    EXPECT_EQ(ck::NumericLimits<bf6_t>::MinSubnorm(), bf6_t(0b000001));
+    EXPECT_EQ(ck::NumericLimits<bf6_t>::MaxSubnorm(), bf6_t(0b000011));
+}
+TEST(BF6, ConvertFP32Nearest)
+{
+    // set maximum bf6 value
+    float max_bf6 = 28.0f;
+    // convert 0 float to bf6 and back, check if holds
+    ASSERT_NEAR(0.0f, type_convert<float>(bf6_convert_rne(0.0f)), 0.0f);
+    // convert max_bf6 to float and check if equal to max_bf6
+    ASSERT_NEAR(max_bf6, type_convert<float>(bf6_convert_rne(max_bf6)), 0.0f);
+    // convert maximal float to bf6 and back, check if clipped to max_bf6
+    ASSERT_NEAR(
+        max_bf6, type_convert<float>(bf6_convert_rne(std::numeric_limits<float>::max())), 0.0f);
+    // convert float Inf to bf6 and back, check if clipped to max_bf6
+    ASSERT_NEAR(max_bf6,
+                type_convert<float>(bf6_convert_rne(std::numeric_limits<float>::infinity())),
+                0.0f);
+    // convert float value less than bf6 subnorm to bf6 and back, check if equal to 0.0
+    float less_than_subnorm = 0.03125f;
+    ASSERT_NEAR(0.0f, type_convert<float>(bf6_convert_rne(less_than_subnorm)), 0.0f);
+    // convert float NaN to bf6 and back, check if clipped to max_bf6
+    ASSERT_NEAR(max_bf6,
+                type_convert<float>(bf6_convert_rne(std::numeric_limits<float>::quiet_NaN())),
+                0.0f);
+    // positive norm float value to bf6 and back, check if holds
+    float pos_float = 0.25f;
+    ASSERT_NEAR(pos_float, type_convert<float>(bf6_convert_rne(pos_float)), 0.0f);
+    // negative norm float value to bf6 and back, check if holds
+    float neg_float = -0.5f;
+    ASSERT_NEAR(neg_float, type_convert<float>(bf6_convert_rne(neg_float)), 0.0f);
+    // positive subnorm float value to bf6 and back, check if holds
+    pos_float = 0.1875f;
+    ASSERT_NEAR(pos_float, type_convert<float>(bf6_convert_rne(pos_float)), 0.0f);
+    // negative subnorm float value to bf6 and back, check if holds
+    neg_float = -0.0625f;
+    ASSERT_NEAR(neg_float, type_convert<float>(bf6_convert_rne(neg_float)), 0.0f);
+}
+TEST(BF6, ConvertFP32Stochastic)
+{
+    // fix the tolerance value
+    float abs_tol = 1e-6;
+    // set maximum bf6 value
+    float max_bf6 = 28.0f;
+    // convert 0 float to bf6 and back, check if holds
+    ASSERT_NEAR(0.0f, type_convert<float>(bf6_convert_sr(0.0f)), abs_tol);
+    // convert maximal bf6_t to float and check if equal to max_bf6
+    ASSERT_NEAR(max_bf6, type_convert<float>(bf6_convert_sr(max_bf6)), abs_tol);
+    // convert maximal float to bf6 and back, check if clipped to max_bf6
+    ASSERT_NEAR(
+        max_bf6, type_convert<float>(bf6_convert_sr(std::numeric_limits<float>::max())), abs_tol);
+    // convert float Inf to bf6 and back, check if clipped to max_bf6
+    ASSERT_NEAR(max_bf6,
+                type_convert<float>(bf6_convert_rne(std::numeric_limits<float>::infinity())),
+                0.0f);
+    // convert float NaN to bf6 and back, check if clipped to max_bf6
+    ASSERT_NEAR(max_bf6,
+                type_convert<float>(bf6_convert_rne(std::numeric_limits<float>::quiet_NaN())),
+                0.0f);
+    // positive norm float value to bf6 and back, check if holds
+    float pos_float = 0.25f;
+    ASSERT_NEAR(pos_float, type_convert<float>(bf6_convert_sr(pos_float)), abs_tol);
+    // negative norm float value to bf6 and back, check if holds
+    float neg_float = -0.5f;
+    ASSERT_NEAR(neg_float, type_convert<float>(bf6_convert_sr(neg_float)), abs_tol);
+    // positive subnorm float value to bf6 and back, check if holds
+    pos_float = 0.1875f;
+    ASSERT_NEAR(pos_float, type_convert<float>(bf6_convert_sr(pos_float)), abs_tol);
+    // negative subnorm float value to bf6 and back, check if holds
+    neg_float = -0.0625f;
+    ASSERT_NEAR(neg_float, type_convert<float>(bf6_convert_sr(neg_float)), abs_tol);
+}
+TEST(BF6, ScaledConvertFP32Nearest)
+{
+    // set maximum scale
+    float max_scale = type_convert<float>(ck::NumericLimits<e8m0_bexp_t>::Max()); // 0xFE -> float
+    // set minimum scale
+    float min_scale = type_convert<float>(ck::NumericLimits<e8m0_bexp_t>::Min()); // 0x00 -> float
+    // set arbitrary scale to 256.0
+    float test_scale = 256.0f; // 0b10000111
+    // convert 0 float to bf6 and back with maximal scale, check if holds
+    ASSERT_NEAR(
+        0.0f, scaled_type_convert<float>(e8m0_bexp_t(max_scale), bf6_convert_rne(0.0f)), 0.0f);
+    // convert 0 float to bf6 and back with minimal scale, check if holds
+    ASSERT_NEAR(
+        0.0f, scaled_type_convert<float>(e8m0_bexp_t(min_scale), bf6_convert_rne(0.0f)), 0.0f);
+    // positive norm float value to bf6 and back with various scales, check if holds
+    float pos_float = 0.25f;
+    ASSERT_NEAR(pos_float * test_scale,
+                scaled_type_convert<float>(e8m0_bexp_t(test_scale), bf6_convert_rne(pos_float)),
+                0.0f);
+    ASSERT_NEAR(pos_float * max_scale,
+                scaled_type_convert<float>(e8m0_bexp_t(max_scale), bf6_convert_rne(pos_float)),
+                0.0f);
+    ASSERT_NEAR(pos_float * min_scale,
+                scaled_type_convert<float>(e8m0_bexp_t(min_scale), bf6_convert_rne(pos_float)),
+                0.0f);
+    // negative norm float value to bf6 and back with various scales, check if holds
+    float neg_float = -0.5f;
+    ASSERT_NEAR(neg_float * test_scale,
+                scaled_type_convert<float>(e8m0_bexp_t(test_scale), bf6_convert_rne(neg_float)),
+                0.0f);
+    ASSERT_NEAR(neg_float * max_scale,
+                scaled_type_convert<float>(e8m0_bexp_t(max_scale), bf6_convert_rne(neg_float)),
+                0.0f);
+    ASSERT_NEAR(neg_float * min_scale,
+                scaled_type_convert<float>(e8m0_bexp_t(min_scale), bf6_convert_rne(neg_float)),
+                0.0f);
+    // positive subnorm float value to bf6 and back with various scales, check if holds
+    pos_float = 0.1875f;
+    ASSERT_NEAR(pos_float * test_scale,
+                scaled_type_convert<float>(e8m0_bexp_t(test_scale), bf6_convert_rne(pos_float)),
+                0.0f);
+    ASSERT_NEAR(pos_float * max_scale,
+                scaled_type_convert<float>(e8m0_bexp_t(max_scale), bf6_convert_rne(pos_float)),
+                0.0f);
+    ASSERT_NEAR(pos_float * min_scale,
+                scaled_type_convert<float>(e8m0_bexp_t(min_scale), bf6_convert_rne(pos_float)),
+                0.0f);
+    // negative subnorm float value to bf6 and back with various scales, check if holds
+    neg_float = -0.0625f;
+    ASSERT_NEAR(neg_float * test_scale,
+                scaled_type_convert<float>(e8m0_bexp_t(test_scale), bf6_convert_rne(neg_float)),
+                0.0f);
+    ASSERT_NEAR(neg_float * max_scale,
+                scaled_type_convert<float>(e8m0_bexp_t(max_scale), bf6_convert_rne(neg_float)),
+                0.0f);
+    ASSERT_NEAR(neg_float * min_scale,
+                scaled_type_convert<float>(e8m0_bexp_t(min_scale), bf6_convert_rne(neg_float)),
+                0.0f);
+}
+TEST(BF6, ScaledConvertFP32Stochastic)
+{
+    // fix the tolerance value
+    float abs_tol = 1e-6;
+    // set maximum scale
+    float max_scale = type_convert<float>(ck::NumericLimits<e8m0_bexp_t>::Max()); // 0xFE -> float
+    // set minimum scale
+    float min_scale = type_convert<float>(ck::NumericLimits<e8m0_bexp_t>::Min()); // 0x00 -> float
+    // set arbitrary scale to 256.0
+    float test_scale = 256.0f; // 0b10000111
+    // convert 0 float to bf6 and back with maximal scale, check if holds
+    ASSERT_NEAR(
+        0.0f, scaled_type_convert<float>(e8m0_bexp_t(max_scale), bf6_convert_sr(0.0f)), abs_tol);
+    // convert 0 float to bf6 and back with minimal scale, check if holds
+    ASSERT_NEAR(
+        0.0f, scaled_type_convert<float>(e8m0_bexp_t(min_scale), bf6_convert_sr(0.0f)), abs_tol);
+    // positive norm float value to bf6 and back with various scales, check if holds
+    float pos_float = 0.25f;
+    ASSERT_NEAR(pos_float * test_scale,
+                scaled_type_convert<float>(e8m0_bexp_t(test_scale), bf6_convert_sr(pos_float)),
+                abs_tol);
+    ASSERT_NEAR(pos_float * max_scale,
+                scaled_type_convert<float>(e8m0_bexp_t(max_scale), bf6_convert_sr(pos_float)),
+                abs_tol);
+    ASSERT_NEAR(pos_float * min_scale,
+                scaled_type_convert<float>(e8m0_bexp_t(min_scale), bf6_convert_sr(pos_float)),
+                abs_tol);
+    // negative norm float value to bf6 and back with various scales, check if holds
+    float neg_float = -0.5f;
+    ASSERT_NEAR(neg_float * test_scale,
+                scaled_type_convert<float>(e8m0_bexp_t(test_scale), bf6_convert_sr(neg_float)),
+                abs_tol);
+    ASSERT_NEAR(neg_float * max_scale,
+                scaled_type_convert<float>(e8m0_bexp_t(max_scale), bf6_convert_sr(neg_float)),
+                abs_tol);
+    ASSERT_NEAR(neg_float * min_scale,
+                scaled_type_convert<float>(e8m0_bexp_t(min_scale), bf6_convert_sr(neg_float)),
+                abs_tol);
+    // positive subnorm float value to bf6 and back with various scales, check if holds
+    pos_float = 0.1875f;
+    ASSERT_NEAR(pos_float * test_scale,
+                scaled_type_convert<float>(e8m0_bexp_t(test_scale), bf6_convert_sr(pos_float)),
+                abs_tol);
+    ASSERT_NEAR(pos_float * max_scale,
+                scaled_type_convert<float>(e8m0_bexp_t(max_scale), bf6_convert_sr(pos_float)),
+                abs_tol);
+    ASSERT_NEAR(pos_float * min_scale,
+                scaled_type_convert<float>(e8m0_bexp_t(min_scale), bf6_convert_sr(pos_float)),
+                abs_tol);
+    // negative subnorm float value to bf6 and back with various scales, check if holds
+    neg_float = -0.0625f;
+    ASSERT_NEAR(neg_float * test_scale,
+                scaled_type_convert<float>(e8m0_bexp_t(test_scale), bf6_convert_sr(neg_float)),
+                abs_tol);
+    ASSERT_NEAR(neg_float * max_scale,
+                scaled_type_convert<float>(e8m0_bexp_t(max_scale), bf6_convert_sr(neg_float)),
+                abs_tol);
+    ASSERT_NEAR(neg_float * min_scale,
+                scaled_type_convert<float>(e8m0_bexp_t(min_scale), bf6_convert_sr(neg_float)),
+                abs_tol);
+}
+TEST(BF6, TestSize)
+{
+    ASSERT_EQ(1, sizeof(bf6_t));
+    ASSERT_EQ(12, sizeof(bf6x16_pk_t));
+    ASSERT_EQ(24, sizeof(bf6x32_pk_t));
+    ASSERT_EQ(16, sizeof(vector_type<bf6x16_pk_t, 1>));
+    ASSERT_EQ(32, sizeof(vector_type<bf6x16_pk_t, 2>));
+    ASSERT_EQ(32, sizeof(vector_type<bf6x32_pk_t, 1>));
+}
+TEST(BF6, TestAlignment)
+{
+    ASSERT_EQ(1, alignof(bf6_t));
+    ASSERT_EQ(4, alignof(bf6x16_pk_t));
+    ASSERT_EQ(4, alignof(bf6x32_pk_t));
+    ASSERT_EQ(16, alignof(vector_type<bf6x16_pk_t, 1>));
+    ASSERT_EQ(32, alignof(vector_type<bf6x16_pk_t, 2>));
+    ASSERT_EQ(32, alignof(vector_type<bf6x32_pk_t, 1>));
+}
+// test vector of 1 bf6x16_pk_t, contains 16 bf6_t
+TEST(BF6, TestAsType16x1)
+{
+    // test size
+    const int vector_size = 1;
+    const int packed_size = 16;
+    typedef int8_t test_vec_t __attribute__((ext_vector_type(16)));
+    test_vec_t test_vec = {bf6_t(0b000000),
+                           bf6_t(0b100000),
+                           bf6_t(0b000001),
+                           bf6_t(0b100001),
+                           bf6_t(0b000010),
+                           bf6_t(0b100010),
+                           bf6_t(0b000011),
+                           bf6_t(0b100011),
+                           bf6_t(0b000100),
+                           bf6_t(0b100100),
+                           bf6_t(0b000101),
+                           bf6_t(0b100101),
+                           bf6_t(0b000110),
+                           bf6_t(0b100110),
+                           bf6_t(0b001011),
+                           bf6_t(0b101011)};
+    // reference vector
+    vector_type<bf6x16_pk_t, vector_size> right_vec;
+    // check default CTOR
+    ck::static_for<0, packed_size, 1>{}([&](auto i) {
+        ASSERT_EQ(
+            right_vec.template AsType<bf6x16_pk_t>()(Number<0>{}).template unpack<>(Number<i>{}),
+            0);
+    });
+    // assign test values to the vector
+    ck::static_for<0, vector_size, 1>{}([&](auto i) {
+        right_vec.template AsType<bf6x16_pk_t>()(Number<i>{}) = bf6x16_pk_t{}.pack(test_vec);
+    });
+    // copy the vector
+    vector_type<bf6x16_pk_t, vector_size> left_vec{right_vec};
+    // check if values were copied correctly
+    ck::static_for<0, packed_size, 1>{}([&](auto i) {
+        ASSERT_EQ(
+            left_vec.template AsType<bf6x16_pk_t>()(Number<0>{}).template unpack<>(Number<i>{}),
+            static_cast<bf6_t>(test_vec[static_cast<int>(i)]));
+    });
+}
+// test vector of 2 bf6x16_pk_t, contains 32 bf6_t
+TEST(BF6, TestAsType16x2)
+{
+    // test size
+    const int vector_size = 2;
+    const int packed_size = 16;
+    typedef int8_t test_vec_t __attribute__((ext_vector_type(16)));
+    test_vec_t test_vec[2];
+    test_vec[0] = {bf6_t(0b000000),
+                   bf6_t(0b100000),
+                   bf6_t(0b000001),
+                   bf6_t(0b100001),
+                   bf6_t(0b000010),
+                   bf6_t(0b100010),
+                   bf6_t(0b000011),
+                   bf6_t(0b100011),
+                   bf6_t(0b000100),
+                   bf6_t(0b100100),
+                   bf6_t(0b000101),
+                   bf6_t(0b100101),
+                   bf6_t(0b000110),
+                   bf6_t(0b100110),
+                   bf6_t(0b001011),
+                   bf6_t(0b101011)};
+    test_vec[1] = {bf6_t(0b010000),
+                   bf6_t(0b110000),
+                   bf6_t(0b010001),
+                   bf6_t(0b110001),
+                   bf6_t(0b010010),
+                   bf6_t(0b110010),
+                   bf6_t(0b010011),
+                   bf6_t(0b110011),
+                   bf6_t(0b010100),
+                   bf6_t(0b110100),
+                   bf6_t(0b010101),
+                   bf6_t(0b110101),
+                   bf6_t(0b010110),
+                   bf6_t(0b110110),
+                   bf6_t(0b011011),
+                   bf6_t(0b111011)};
+    // reference vector
+    vector_type<bf6x16_pk_t, vector_size> right_vec;
+    // check default CTOR
+    ck::static_for<0, vector_size, 1>{}([&](auto idx_vector) {
+        ck::static_for<0, packed_size, 1>{}([&](auto idx_element) {
+            ASSERT_EQ(right_vec.template AsType<bf6x16_pk_t>()(Number<idx_vector>{})
+                          .template unpack<>(Number<idx_element>{}),
+                      0);
+        });
+    });
+    // assign test values to the vector
+    ck::static_for<0, vector_size, 1>{}([&](auto i) {
+        right_vec.template AsType<bf6x16_pk_t>()(Number<i>{}) = bf6x16_pk_t{}.pack(test_vec[i]);
+    });
+    // copy the vector
+    vector_type<bf6x16_pk_t, vector_size> left_vec{right_vec};
+    // check if values were copied correctly
+    ck::static_for<0, vector_size, 1>{}([&](auto idx_vector) {
+        ck::static_for<0, packed_size, 1>{}([&](auto idx_element) {
+            ASSERT_EQ(left_vec.template AsType<bf6x16_pk_t>()(Number<idx_vector>{})
+                          .template unpack<>(Number<idx_element>{}),
+                      static_cast<bf6_t>(test_vec[idx_vector][static_cast<int>(idx_element)]));
+        });
+    });
+}
+// test vector of 1 bf6x32_pk_t, contains 32 bf6_t
+TEST(BF6, TestAsType32x1)
+{
+    // test size
+    const int vector_size = 1;
+    const int packed_size = 32;
+    typedef int8_t test_vec_t __attribute__((ext_vector_type(32)));
+    test_vec_t test_vec = {bf6_t(0b000000), bf6_t(0b100000), bf6_t(0b000001), bf6_t(0b100001),
+                           bf6_t(0b000010), bf6_t(0b100010), bf6_t(0b000011), bf6_t(0b100011),
+                           bf6_t(0b000100), bf6_t(0b100100), bf6_t(0b000101), bf6_t(0b100101),
+                           bf6_t(0b000110), bf6_t(0b100110), bf6_t(0b001011), bf6_t(0b101011),
+                           bf6_t(0b010000), bf6_t(0b110000), bf6_t(0b010001), bf6_t(0b110001),
+                           bf6_t(0b010010), bf6_t(0b110010), bf6_t(0b010011), bf6_t(0b110011),
+                           bf6_t(0b010100), bf6_t(0b110100), bf6_t(0b010101), bf6_t(0b110101),
+                           bf6_t(0b010110), bf6_t(0b110110), bf6_t(0b011011), bf6_t(0b111011)};
+    // reference vector
+    vector_type<bf6x32_pk_t, vector_size> right_vec;
+    // check default CTOR
+    ck::static_for<0, packed_size, 1>{}([&](auto i) {
+        ASSERT_EQ(
+            right_vec.template AsType<bf6x32_pk_t>()(Number<0>{}).template unpack<>(Number<i>{}),
+            0);
+    });
+    // assign test values to the vector
+    ck::static_for<0, vector_size, 1>{}([&](auto i) {
+        right_vec.template AsType<bf6x32_pk_t>()(Number<i>{}) = bf6x32_pk_t{}.pack(test_vec);
+    });
+    // copy the vector
+    vector_type<bf6x32_pk_t, vector_size> left_vec{right_vec};
+    // check if values were copied correctly
+    ck::static_for<0, packed_size, 1>{}([&](auto i) {
+        ASSERT_EQ(
+            left_vec.template AsType<bf6x32_pk_t>()(Number<0>{}).template unpack<>(Number<i>{}),
+            static_cast<bf6_t>(test_vec[static_cast<int>(i)]));
+    });
+}
--- a/test/data_type/test_e8m0.cpp
+++ b/test/data_type/test_e8m0.cpp
+#include <gtest/gtest.h>
+#include "ck/utility/e8m0.hpp"
+#include "ck/utility/data_type.hpp"
+#include "ck/utility/type_convert.hpp"
+using namespace ck;
+TEST(E8M0, DefaultConstructor)
+{
+    e8m0_bexp_t exp;
+    EXPECT_EQ(exp.data, 0);
+}
+TEST(E8M0, InitConstructor)
+{
+    e8m0_bexp_t exp(0x7F);
+    EXPECT_EQ(exp.data, 0x7F);
+}
+TEST(E8M0, FloatConstructor)
+{
+    e8m0_bexp_t exp(1.0f);
+    EXPECT_EQ(exp.data, 0x7F);
+}
+TEST(E8M0, FloatConstructorNaN)
+{
+    e8m0_bexp_t exp(std::numeric_limits<float>::quiet_NaN());
+    EXPECT_EQ(exp.data, 0xFF);
+}
+TEST(E8M0, FloatConstructorZero)
+{
+    e8m0_bexp_t exp(0.0f);
+    EXPECT_EQ(exp.data, 0);
+}
+TEST(E8M0, ConversionToFloat)
+{
+    e8m0_bexp_t exp(0x7F);
+    float value = type_convert<float>(exp);
+    EXPECT_EQ(value, 1.0f);
+}
+TEST(E8M0, ConversionToFloatNaN)
+{
+    e8m0_bexp_t exp(0xFF);
+    float value = type_convert<float>(exp);
+    EXPECT_TRUE(std::isnan(value));
+}
+TEST(E8M0, MinValue)
+{
+    e8m0_bexp_t exp(0);
+    EXPECT_TRUE(exp == ck::NumericLimits<e8m0_bexp_t>::Min());
+    float value = type_convert<float>(exp);
+    EXPECT_EQ(value, std::powf(2, -ck::NumericUtils<e8m0_bexp_t>::bias));
+}
+TEST(E8M0, MaxValue)
+{
+    e8m0_bexp_t exp(254);
+    EXPECT_TRUE(exp == ck::NumericLimits<e8m0_bexp_t>::Max());
+    float value = type_convert<float>(exp);
+    EXPECT_EQ(value,
+              std::powf(2,
+                        ck::NumericLimits<e8m0_bexp_t>::Max().data -
+                            ck::NumericUtils<e8m0_bexp_t>::bias));
+}
+TEST(E8M0, EqualityOperator)
+{
+    e8m0_bexp_t exp1(0x7F);
+    e8m0_bexp_t exp2(0x7F);
+    EXPECT_TRUE(exp1 == exp2);
+}
+TEST(E8M0, InequalityOperator)
+{
+    e8m0_bexp_t exp1(0x7F);
+    e8m0_bexp_t exp2(0x80);
+    EXPECT_FALSE(exp1 == exp2);
+}
+TEST(E8M0, EqualityOperatorNaN)
+{
+    e8m0_bexp_t exp1(0xFF);
+    e8m0_bexp_t exp2(0xFF);
+    EXPECT_FALSE(exp1 == exp2);
+}
+TEST(E8M0, GetExponentValue)
+{
+    e8m0_bexp_t exp(0x7F);
+    int value = ck::utils::get_exponent_value(exp);
+    EXPECT_EQ(value, 0x7F);
+}
--- a/test/data_type/test_fp4.cpp
+++ b/test/data_type/test_fp4.cpp
+// SPDX-License-Identifier: MIT
+// Copyright (c) 2018-2025, Advanced Micro Devices, Inc. All rights reserved.
+#include "gtest/gtest.h"
+#include "ck/utility/data_type.hpp"
+#include "ck/utility/type_convert.hpp"
+#include "ck/utility/scaled_type_convert.hpp"
+using ck::e8m0_bexp_t;
+using ck::f4_convert_rne;
+using ck::f4_convert_sr;
+using ck::f4_t;
+using ck::f4x2_pk_t;
+using ck::Number;
+using ck::scaled_type_convert;
+using ck::type_convert;
+using ck::vector_type;
+TEST(FP4, NumericLimits)
+{
+    EXPECT_EQ(ck::NumericLimits<f4_t>::Min(), f4_t{0x2});
+    EXPECT_EQ(ck::NumericLimits<f4_t>::Max(), f4_t{0x7});
+    EXPECT_EQ(ck::NumericLimits<f4_t>::Lowest(), f4_t{0xF});
+    EXPECT_EQ(ck::NumericLimits<f4_t>::MinSubnorm(), f4_t{0x1});
+    EXPECT_EQ(ck::NumericLimits<f4_t>::MaxSubnorm(), f4_t{0x1});
+}
+TEST(FP4, ConvertFP32Nearest)
+{
+    // fix the tolerance value
+    float abs_tol = 1e-6;
+    // set maximum fp4 value
+    float max_fp4 = 6.0f;
+    // convert 0 float to fp4 and back, check if holds
+    ASSERT_NEAR(0.0f, type_convert<float>(f4_convert_rne(0.0f)), abs_tol);
+    // convert maximal f4_t to float and check if equal to 6.0
+    ASSERT_NEAR(max_fp4, type_convert<float>(f4_convert_rne(max_fp4)), abs_tol);
+    // convert maximal float to fp4 and back, check if clipped to 6.0
+    ASSERT_NEAR(
+        max_fp4, type_convert<float>(f4_convert_rne(std::numeric_limits<float>::max())), abs_tol);
+    // positive norm float value to fp4 and back, check if holds
+    float pos_float = 1.0f;
+    ASSERT_NEAR(pos_float, type_convert<float>(f4_convert_rne(pos_float)), abs_tol);
+    // negative norm float value to fp4 and back, check if holds
+    float neg_float = -1.5f;
+    ASSERT_NEAR(neg_float, type_convert<float>(f4_convert_rne(neg_float)), abs_tol);
+    // positive subnorm float value to fp4 and back, check if holds
+    pos_float = 0.5f;
+    ASSERT_NEAR(pos_float, type_convert<float>(f4_convert_rne(pos_float)), abs_tol);
+    // negative subnorm float value to fp4 and back, check if holds
+    neg_float = -0.5f;
+    ASSERT_NEAR(neg_float, type_convert<float>(f4_convert_rne(neg_float)), abs_tol);
+}
+TEST(FP4, ConvertFP32Stochastic)
+{
+    // fix the tolerance value
+    float abs_tol = 1e-6;
+    // set maximum fp4 value
+    float max_fp4 = 6.0f;
+    // convert 0 float to fp4 and back, check if holds
+    ASSERT_NEAR(0.0f, type_convert<float>(f4_convert_sr(0.0f)), abs_tol);
+    // convert maximal f4_t to float and check if equal to 6.0
+    ASSERT_NEAR(max_fp4, type_convert<float>(f4_convert_sr(max_fp4)), abs_tol);
+    // convert maximal float to fp4 and back, check if clipped to 6.0
+    ASSERT_NEAR(
+        max_fp4, type_convert<float>(f4_convert_sr(std::numeric_limits<float>::max())), abs_tol);
+    // positive norm float value to fp4 and back, check if holds
+    float pos_float = 1.0f;
+    ASSERT_NEAR(pos_float, type_convert<float>(f4_convert_sr(pos_float)), abs_tol);
+    // negative norm float value to fp4 and back, check if holds
+    float neg_float = -1.5f;
+    ASSERT_NEAR(neg_float, type_convert<float>(f4_convert_sr(neg_float)), abs_tol);
+    // positive subnorm float value to fp4 and back, check if holds
+    pos_float = 0.5f;
+    ASSERT_NEAR(pos_float, type_convert<float>(f4_convert_sr(pos_float)), abs_tol);
+    // negative subnorm float value to fp4 and back, check if holds
+    neg_float = -0.5f;
+    ASSERT_NEAR(neg_float, type_convert<float>(f4_convert_sr(neg_float)), abs_tol);
+}
+TEST(FP4, ScaledConvertFP32Nearest)
+{
+    // fix the tolerance value
+    float abs_tol = 1e-6;
+    // set maximum scale
+    float max_scale = type_convert<float>(ck::NumericLimits<e8m0_bexp_t>::Max()); // 0xFE -> float
+    // set minimum scale
+    float min_scale = type_convert<float>(ck::NumericLimits<e8m0_bexp_t>::Min()); // 0x00 -> float
+    // set arbitrary scale to 256.0
+    float test_scale = 256.0f; // 0b10000111
+    // convert 0 float to fp4 and back with maximal scale, check if holds
+    ASSERT_NEAR(
+        0.0f, scaled_type_convert<float>(e8m0_bexp_t(max_scale), f4_convert_rne(0.0f)), abs_tol);
+    // convert 0 float to fp4 and back with minimal scale, check if holds
+    ASSERT_NEAR(
+        0.0f, scaled_type_convert<float>(e8m0_bexp_t(min_scale), f4_convert_rne(0.0f)), abs_tol);
+    // positive norm float value to fp4 and back with various scales, check if holds
+    float pos_float = 1.0f;
+    ASSERT_NEAR(pos_float * test_scale,
+                scaled_type_convert<float>(e8m0_bexp_t(test_scale), f4_convert_rne(pos_float)),
+                abs_tol);
+    ASSERT_NEAR(pos_float * max_scale,
+                scaled_type_convert<float>(e8m0_bexp_t(max_scale), f4_convert_rne(pos_float)),
+                abs_tol);
+    ASSERT_NEAR(pos_float * min_scale,
+                scaled_type_convert<float>(e8m0_bexp_t(min_scale), f4_convert_rne(pos_float)),
+                abs_tol);
+    // negative norm float value to fp4 and back with various scales, check if holds
+    float neg_float = -1.5f;
+    ASSERT_NEAR(neg_float * test_scale,
+                scaled_type_convert<float>(e8m0_bexp_t(test_scale), f4_convert_rne(neg_float)),
+                abs_tol);
+    ASSERT_NEAR(neg_float * max_scale,
+                scaled_type_convert<float>(e8m0_bexp_t(max_scale), f4_convert_rne(neg_float)),
+                abs_tol);
+    ASSERT_NEAR(neg_float * min_scale,
+                scaled_type_convert<float>(e8m0_bexp_t(min_scale), f4_convert_rne(neg_float)),
+                abs_tol);
+    // positive subnorm float value to fp4 and back with various scales, check if holds
+    pos_float = 0.5f;
+    ASSERT_NEAR(pos_float * test_scale,
+                scaled_type_convert<float>(e8m0_bexp_t(test_scale), f4_convert_rne(pos_float)),
+                abs_tol);
+    ASSERT_NEAR(pos_float * max_scale,
+                scaled_type_convert<float>(e8m0_bexp_t(max_scale), f4_convert_rne(pos_float)),
+                abs_tol);
+    ASSERT_NEAR(pos_float * min_scale,
+                scaled_type_convert<float>(e8m0_bexp_t(min_scale), f4_convert_rne(pos_float)),
+                abs_tol);
+    // negative subnorm float value to fp4 and back with various scales, check if holds
+    neg_float = -0.5f;
+    ASSERT_NEAR(neg_float * test_scale,
+                scaled_type_convert<float>(e8m0_bexp_t(test_scale), f4_convert_rne(neg_float)),
+                abs_tol);
+    ASSERT_NEAR(neg_float * max_scale,
+                scaled_type_convert<float>(e8m0_bexp_t(max_scale), f4_convert_rne(neg_float)),
+                abs_tol);
+    ASSERT_NEAR(neg_float * min_scale,
+                scaled_type_convert<float>(e8m0_bexp_t(min_scale), f4_convert_rne(neg_float)),
+                abs_tol);
+}
+TEST(FP4, ScaledConvertFP32Stochastic)
+{
+    // fix the tolerance value
+    float abs_tol = 1e-6;
+    // set maximum scale
+    float max_scale = type_convert<float>(ck::NumericLimits<e8m0_bexp_t>::Max()); // 0xFE -> float
+    // set minimum scale
+    float min_scale = type_convert<float>(ck::NumericLimits<e8m0_bexp_t>::Min()); // 0x00 -> float
+    // set arbitrary scale to 256.0
+    float test_scale = 256.0f; // 0b10000111
+    // convert 0 float to fp4 and back with maximal scale, check if holds
+    ASSERT_NEAR(
+        0.0f, scaled_type_convert<float>(e8m0_bexp_t(max_scale), f4_convert_sr(0.0f)), abs_tol);
+    // convert 0 float to fp4 and back with minimal scale, check if holds
+    ASSERT_NEAR(
+        0.0f, scaled_type_convert<float>(e8m0_bexp_t(min_scale), f4_convert_sr(0.0f)), abs_tol);
+    // positive norm float value to fp4 and back with various scales, check if holds
+    float pos_float = 1.0f;
+    ASSERT_NEAR(pos_float * test_scale,
+                scaled_type_convert<float>(e8m0_bexp_t(test_scale), f4_convert_sr(pos_float)),
+                abs_tol);
+    ASSERT_NEAR(pos_float * max_scale,
+                scaled_type_convert<float>(e8m0_bexp_t(max_scale), f4_convert_sr(pos_float)),
+                abs_tol);
+    ASSERT_NEAR(pos_float * min_scale,
+                scaled_type_convert<float>(e8m0_bexp_t(min_scale), f4_convert_sr(pos_float)),
+                abs_tol);
+    // negative norm float value to fp4 and back with various scales, check if holds
+    float neg_float = -1.5f;
+    ASSERT_NEAR(neg_float * test_scale,
+                scaled_type_convert<float>(e8m0_bexp_t(test_scale), f4_convert_sr(neg_float)),
+                abs_tol);
+    ASSERT_NEAR(neg_float * max_scale,
+                scaled_type_convert<float>(e8m0_bexp_t(max_scale), f4_convert_sr(neg_float)),
+                abs_tol);
+    ASSERT_NEAR(neg_float * min_scale,
+                scaled_type_convert<float>(e8m0_bexp_t(min_scale), f4_convert_sr(neg_float)),
+                abs_tol);
+    // positive subnorm float value to fp4 and back with various scales, check if holds
+    pos_float = 0.5f;
+    ASSERT_NEAR(pos_float * test_scale,
+                scaled_type_convert<float>(e8m0_bexp_t(test_scale), f4_convert_sr(pos_float)),
+                abs_tol);
+    ASSERT_NEAR(pos_float * max_scale,
+                scaled_type_convert<float>(e8m0_bexp_t(max_scale), f4_convert_sr(pos_float)),
+                abs_tol);
+    ASSERT_NEAR(pos_float * min_scale,
+                scaled_type_convert<float>(e8m0_bexp_t(min_scale), f4_convert_sr(pos_float)),
+                abs_tol);
+    // negative subnorm float value to fp4 and back with various scales, check if holds
+    neg_float = -0.5f;
+    ASSERT_NEAR(neg_float * test_scale,
+                scaled_type_convert<float>(e8m0_bexp_t(test_scale), f4_convert_sr(neg_float)),
+                abs_tol);
+    ASSERT_NEAR(neg_float * max_scale,
+                scaled_type_convert<float>(e8m0_bexp_t(max_scale), f4_convert_sr(neg_float)),
+                abs_tol);
+    ASSERT_NEAR(neg_float * min_scale,
+                scaled_type_convert<float>(e8m0_bexp_t(min_scale), f4_convert_sr(neg_float)),
+                abs_tol);
+}
+TEST(FP4, TestSize)
+{
+    ASSERT_EQ(1, sizeof(f4x2_pk_t));
+    ASSERT_EQ(1, sizeof(vector_type<f4x2_pk_t, 1>));
+    ASSERT_EQ(2, sizeof(vector_type<f4x2_pk_t, 2>));
+    ASSERT_EQ(4, sizeof(vector_type<f4x2_pk_t, 4>));
+    ASSERT_EQ(8, sizeof(vector_type<f4x2_pk_t, 8>));
+    ASSERT_EQ(16, sizeof(vector_type<f4x2_pk_t, 16>));
+    ASSERT_EQ(32, sizeof(vector_type<f4x2_pk_t, 32>));
+}
+TEST(FP4, TestAlignment)
+{
+    ASSERT_EQ(1, alignof(f4x2_pk_t));
+    ASSERT_EQ(1, alignof(vector_type<f4x2_pk_t, 1>));
+    ASSERT_EQ(2, alignof(vector_type<f4x2_pk_t, 2>));
+    ASSERT_EQ(4, alignof(vector_type<f4x2_pk_t, 4>));
+    ASSERT_EQ(8, alignof(vector_type<f4x2_pk_t, 8>));
+    ASSERT_EQ(16, alignof(vector_type<f4x2_pk_t, 16>));
+    ASSERT_EQ(32, alignof(vector_type<f4x2_pk_t, 32>));
+}
+// test vector of 1 f4x2_pk_t, contains 2 f4_t
+TEST(FP4, TestAsType1)
+{
+    // test size
+    const int size                        = 1;
+    std::vector<f4x2_pk_t::type> test_vec = {f4x2_pk_t::type{0b0010}, f4x2_pk_t::type{0b1001}};
+    // reference vector
+    vector_type<f4x2_pk_t, size> right_vec;
+    // check default CTOR
+    ck::static_for<0, size, 1>{}([&](auto i) {
+        ASSERT_EQ(
+            right_vec.template AsType<f4x2_pk_t>()(Number<i>{}).template unpack<>(Number<0>{}), 0);
+        ASSERT_EQ(
+            right_vec.template AsType<f4x2_pk_t>()(Number<i>{}).template unpack<>(Number<1>{}), 0);
+    });
+    // assign test values to the vector
+    ck::static_for<0, size, 1>{}([&](auto i) {
+        right_vec.template AsType<f4x2_pk_t>()(Number<i>{}) =
+            f4x2_pk_t{}.pack(test_vec.at(i), test_vec.at(i + 1));
+    });
+    // copy the vector
+    vector_type<f4x2_pk_t, size> left_vec{right_vec};
+    // check if values were copied correctly
+    ck::static_for<0, size, 1>{}([&](auto i) {
+        ASSERT_EQ(left_vec.template AsType<f4x2_pk_t>()(Number<i>{}).template unpack<>(Number<0>{}),
+                  test_vec.at(i));
+        ASSERT_EQ(left_vec.template AsType<f4x2_pk_t>()(Number<i>{}).template unpack<>(Number<1>{}),
+                  test_vec.at(i + 1));
+    });
+}
+// test vector of 2 f4x2_pk_t, contains 4 f4_t
+TEST(FP4, TestAsType2)
+{
+    // test size
+    const int size                        = 2;
+    std::vector<f4x2_pk_t::type> test_vec = {f4x2_pk_t::type{0b0010},
+                                             f4x2_pk_t::type{0b1001},
+                                             f4x2_pk_t::type{0b0001},
+                                             f4x2_pk_t::type{0b0111}};
+    // reference vector
+    vector_type<f4x2_pk_t, size> right_vec;
+    // check default CTOR
+    ck::static_for<0, size, 1>{}([&](auto i) {
+        ASSERT_EQ(
+            right_vec.template AsType<f4x2_pk_t>()(Number<i>{}).template unpack<>(Number<0>{}), 0);
+        ASSERT_EQ(
+            right_vec.template AsType<f4x2_pk_t>()(Number<i>{}).template unpack<>(Number<1>{}), 0);
+    });
+    // assign test values to the vector
+    ck::static_for<0, size, 1>{}([&](auto i) {
+        right_vec.template AsType<f4x2_pk_t>()(Number<i>{}) =
+            f4x2_pk_t{}.pack(test_vec.at(i), test_vec.at(i + 1));
+    });
+    // copy the vector
+    vector_type<f4x2_pk_t, size> left_vec{right_vec};
+    // check if values were copied correctly
+    ck::static_for<0, size, 1>{}([&](auto i) {
+        ASSERT_EQ(left_vec.template AsType<f4x2_pk_t>()(Number<i>{}).template unpack<>(Number<0>{}),
+                  test_vec.at(i));
+        ASSERT_EQ(left_vec.template AsType<f4x2_pk_t>()(Number<i>{}).template unpack<>(Number<1>{}),
+                  test_vec.at(i + 1));
+    });
+}
+// test vector of 4 f4x2_pk_t, contains 8 f4_t
+TEST(FP4, TestAsType4)
+{
+    // test size
+    const int size                        = 4;
+    std::vector<f4x2_pk_t::type> test_vec = {f4x2_pk_t::type{0b0010},
+                                             f4x2_pk_t::type{0b1001},
+                                             f4x2_pk_t::type{0b0001},
+                                             f4x2_pk_t::type{0b0111},
+                                             f4x2_pk_t::type{0b1010},
+                                             f4x2_pk_t::type{0b0001},
+                                             f4x2_pk_t::type{0b1001},
+                                             f4x2_pk_t::type{0b1111}};
+    // reference vector
+    vector_type<f4x2_pk_t, size> right_vec;
+    // check default CTOR
+    ck::static_for<0, size, 1>{}([&](auto i) {
+        ASSERT_EQ(
+            right_vec.template AsType<f4x2_pk_t>()(Number<i>{}).template unpack<>(Number<0>{}), 0);
+        ASSERT_EQ(
+            right_vec.template AsType<f4x2_pk_t>()(Number<i>{}).template unpack<>(Number<1>{}), 0);
+    });
+    // assign test values to the vector
+    ck::static_for<0, size, 1>{}([&](auto i) {
+        right_vec.template AsType<f4x2_pk_t>()(Number<i>{}) =
+            f4x2_pk_t{}.pack(test_vec.at(i), test_vec.at(i + 1));
+    });
+    // copy the vector
+    vector_type<f4x2_pk_t, size> left_vec{right_vec};
+    // check if values were copied correctly
+    ck::static_for<0, size, 1>{}([&](auto i) {
+        ASSERT_EQ(left_vec.template AsType<f4x2_pk_t>()(Number<i>{}).template unpack<>(Number<0>{}),
+                  test_vec.at(i));
+        ASSERT_EQ(left_vec.template AsType<f4x2_pk_t>()(Number<i>{}).template unpack<>(Number<1>{}),
+                  test_vec.at(i + 1));
+    });
+}
+// test vector of 8 f4x2_pk_t, contains 16 f4_t
+TEST(FP4, TestAsType8)
+{
+    // test size
+    const int size                        = 8;
+    std::vector<f4x2_pk_t::type> test_vec = {f4x2_pk_t::type{0b0010},
+                                             f4x2_pk_t::type{0b1001},
+                                             f4x2_pk_t::type{0b0001},
+                                             f4x2_pk_t::type{0b0111},
+                                             f4x2_pk_t::type{0b1010},
+                                             f4x2_pk_t::type{0b0001},
+                                             f4x2_pk_t::type{0b1001},
+                                             f4x2_pk_t::type{0b1111},
+                                             f4x2_pk_t::type{0b0001},
+                                             f4x2_pk_t::type{0b0111},
+                                             f4x2_pk_t::type{0b1010},
+                                             f4x2_pk_t::type{0b0001},
+                                             f4x2_pk_t::type{0b0010},
+                                             f4x2_pk_t::type{0b1001},
+                                             f4x2_pk_t::type{0b1001},
+                                             f4x2_pk_t::type{0b1111}};
+    // reference vector
+    vector_type<f4x2_pk_t, size> right_vec;
+    // check default CTOR
+    ck::static_for<0, size, 1>{}([&](auto i) {
+        ASSERT_EQ(
+            right_vec.template AsType<f4x2_pk_t>()(Number<i>{}).template unpack<>(Number<0>{}), 0);
+        ASSERT_EQ(
+            right_vec.template AsType<f4x2_pk_t>()(Number<i>{}).template unpack<>(Number<1>{}), 0);
+    });
+    // assign test values to the vector
+    ck::static_for<0, size, 1>{}([&](auto i) {
+        right_vec.template AsType<f4x2_pk_t>()(Number<i>{}) =
+            f4x2_pk_t{}.pack(test_vec.at(i), test_vec.at(i + 1));
+    });
+    // copy the vector
+    vector_type<f4x2_pk_t, size> left_vec{right_vec};
+    // check if values were copied correctly
+    ck::static_for<0, size, 1>{}([&](auto i) {
+        ASSERT_EQ(left_vec.template AsType<f4x2_pk_t>()(Number<i>{}).template unpack<>(Number<0>{}),
+                  test_vec.at(i));
+        ASSERT_EQ(left_vec.template AsType<f4x2_pk_t>()(Number<i>{}).template unpack<>(Number<1>{}),
+                  test_vec.at(i + 1));
+    });
+}
+// test vector of 16 f4x2_pk_t, contains 32 f4_t
+TEST(FP4, TestAsType16)
+{
+    // test size
+    const int size                        = 16;
+    std::vector<f4x2_pk_t::type> test_vec = {
+        f4x2_pk_t::type{0b0010}, f4x2_pk_t::type{0b1001}, f4x2_pk_t::type{0b0001},
+        f4x2_pk_t::type{0b0111}, f4x2_pk_t::type{0b1010}, f4x2_pk_t::type{0b0001},
+        f4x2_pk_t::type{0b1001}, f4x2_pk_t::type{0b1111}, f4x2_pk_t::type{0b0001},
+        f4x2_pk_t::type{0b0111}, f4x2_pk_t::type{0b1010}, f4x2_pk_t::type{0b0001},
+        f4x2_pk_t::type{0b0010}, f4x2_pk_t::type{0b1001}, f4x2_pk_t::type{0b1001},
+        f4x2_pk_t::type{0b1111}, f4x2_pk_t::type{0b0010}, f4x2_pk_t::type{0b1001},
+        f4x2_pk_t::type{0b0001}, f4x2_pk_t::type{0b0111}, f4x2_pk_t::type{0b1010},
+        f4x2_pk_t::type{0b0001}, f4x2_pk_t::type{0b1001}, f4x2_pk_t::type{0b1111},
+        f4x2_pk_t::type{0b0001}, f4x2_pk_t::type{0b0111}, f4x2_pk_t::type{0b1010},
+        f4x2_pk_t::type{0b0001}, f4x2_pk_t::type{0b0010}, f4x2_pk_t::type{0b1001},
+        f4x2_pk_t::type{0b1001}, f4x2_pk_t::type{0b1111}};
+    // reference vector
+    vector_type<f4x2_pk_t, size> right_vec;
+    // check default CTOR
+    ck::static_for<0, size, 1>{}([&](auto i) {
+        ASSERT_EQ(
+            right_vec.template AsType<f4x2_pk_t>()(Number<i>{}).template unpack<>(Number<0>{}), 0);
+        ASSERT_EQ(
+            right_vec.template AsType<f4x2_pk_t>()(Number<i>{}).template unpack<>(Number<1>{}), 0);
+    });
+    // assign test values to the vector
+    ck::static_for<0, size, 1>{}([&](auto i) {
+        right_vec.template AsType<f4x2_pk_t>()(Number<i>{}) =
+            f4x2_pk_t{}.pack(test_vec.at(i), test_vec.at(i + 1));
+    });
+    // copy the vector
+    vector_type<f4x2_pk_t, size> left_vec{right_vec};
+    // check if values were copied correctly
+    ck::static_for<0, size, 1>{}([&](auto i) {
+        ASSERT_EQ(left_vec.template AsType<f4x2_pk_t>()(Number<i>{}).template unpack<>(Number<0>{}),
+                  test_vec.at(i));
+        ASSERT_EQ(left_vec.template AsType<f4x2_pk_t>()(Number<i>{}).template unpack<>(Number<1>{}),
+                  test_vec.at(i + 1));
+    });
+}
+// test vector of 32 f4x2_pk_t, contains 64 f4_t
+TEST(FP4, TestAsType32)
+{
+    // test size
+    const int size                        = 32;
+    std::vector<f4x2_pk_t::type> test_vec = {
+        f4x2_pk_t::type{0b0010}, f4x2_pk_t::type{0b1001}, f4x2_pk_t::type{0b0001},
+        f4x2_pk_t::type{0b0111}, f4x2_pk_t::type{0b1010}, f4x2_pk_t::type{0b0001},
+        f4x2_pk_t::type{0b1001}, f4x2_pk_t::type{0b1111}, f4x2_pk_t::type{0b0001},
+        f4x2_pk_t::type{0b0111}, f4x2_pk_t::type{0b1010}, f4x2_pk_t::type{0b0001},
+        f4x2_pk_t::type{0b0010}, f4x2_pk_t::type{0b1001}, f4x2_pk_t::type{0b1001},
+        f4x2_pk_t::type{0b1111}, f4x2_pk_t::type{0b0010}, f4x2_pk_t::type{0b1001},
+        f4x2_pk_t::type{0b0001}, f4x2_pk_t::type{0b0111}, f4x2_pk_t::type{0b1010},
+        f4x2_pk_t::type{0b0001}, f4x2_pk_t::type{0b1001}, f4x2_pk_t::type{0b1111},
+        f4x2_pk_t::type{0b0001}, f4x2_pk_t::type{0b0111}, f4x2_pk_t::type{0b1010},
+        f4x2_pk_t::type{0b0001}, f4x2_pk_t::type{0b0010}, f4x2_pk_t::type{0b1001},
+        f4x2_pk_t::type{0b1001}, f4x2_pk_t::type{0b1111}, f4x2_pk_t::type{0b0010},
+        f4x2_pk_t::type{0b1001}, f4x2_pk_t::type{0b0001}, f4x2_pk_t::type{0b0111},
+        f4x2_pk_t::type{0b1010}, f4x2_pk_t::type{0b0001}, f4x2_pk_t::type{0b1001},
+        f4x2_pk_t::type{0b1111}, f4x2_pk_t::type{0b0001}, f4x2_pk_t::type{0b0111},
+        f4x2_pk_t::type{0b1010}, f4x2_pk_t::type{0b0001}, f4x2_pk_t::type{0b0010},
+        f4x2_pk_t::type{0b1001}, f4x2_pk_t::type{0b1001}, f4x2_pk_t::type{0b1111},
+        f4x2_pk_t::type{0b0010}, f4x2_pk_t::type{0b1001}, f4x2_pk_t::type{0b0001},
+        f4x2_pk_t::type{0b0111}, f4x2_pk_t::type{0b1010}, f4x2_pk_t::type{0b0001},
+        f4x2_pk_t::type{0b1001}, f4x2_pk_t::type{0b1111}, f4x2_pk_t::type{0b0001},
+        f4x2_pk_t::type{0b0111}, f4x2_pk_t::type{0b1010}, f4x2_pk_t::type{0b0001},
+        f4x2_pk_t::type{0b0010}, f4x2_pk_t::type{0b1001}, f4x2_pk_t::type{0b1001},
+        f4x2_pk_t::type{0b1111}};
+    // reference vector
+    vector_type<f4x2_pk_t, size> right_vec;
+    // check default CTOR
+    ck::static_for<0, size, 1>{}([&](auto i) {
+        ASSERT_EQ(
+            right_vec.template AsType<f4x2_pk_t>()(Number<i>{}).template unpack<>(Number<0>{}), 0);
+        ASSERT_EQ(
+            right_vec.template AsType<f4x2_pk_t>()(Number<i>{}).template unpack<>(Number<1>{}), 0);
+    });
+    // assign test values to the vector
+    ck::static_for<0, size, 1>{}([&](auto i) {
+        right_vec.template AsType<f4x2_pk_t>()(Number<i>{}) =
+            f4x2_pk_t{}.pack(test_vec.at(i), test_vec.at(i + 1));
+    });
+    // copy the vector
+    vector_type<f4x2_pk_t, size> left_vec{right_vec};
+    // check if values were copied correctly
+    ck::static_for<0, size, 1>{}([&](auto i) {
+        ASSERT_EQ(left_vec.template AsType<f4x2_pk_t>()(Number<i>{}).template unpack<>(Number<0>{}),
+                  test_vec.at(i));
+        ASSERT_EQ(left_vec.template AsType<f4x2_pk_t>()(Number<i>{}).template unpack<>(Number<1>{}),
+                  test_vec.at(i + 1));
+    });
+}
--- a/test/data_type/test_fp6.cpp
+++ b/test/data_type/test_fp6.cpp
+// SPDX-License-Identifier: MIT
+// Copyright (c) 2025, Advanced Micro Devices, Inc. All rights reserved.
+#include "gtest/gtest.h"
+#include "ck/utility/data_type.hpp"
+#include "ck/utility/type_convert.hpp"
+#include "ck/utility/scaled_type_convert.hpp"
+using ck::e8m0_bexp_t;
+using ck::f6_convert_rne;
+using ck::f6_convert_sr;
+using ck::f6_t;
+using ck::f6x16_pk_t;
+using ck::f6x32_pk_t;
+using ck::Number;
+using ck::scaled_type_convert;
+using ck::type_convert;
+using ck::vector_type;
+TEST(FP6, NumericLimits)
+{
+    EXPECT_EQ(ck::NumericLimits<f6_t>::Min(), f6_t(0b001000));
+    EXPECT_EQ(ck::NumericLimits<f6_t>::Max(), f6_t(0b011111));
+    EXPECT_EQ(ck::NumericLimits<f6_t>::Lowest(), f6_t(0b111111));
+    EXPECT_EQ(ck::NumericLimits<f6_t>::MinSubnorm(), f6_t(0b000001));
+    EXPECT_EQ(ck::NumericLimits<f6_t>::MaxSubnorm(), f6_t(0b000111));
+}
+TEST(FP6, ConvertFP32Nearest)
+{
+    // set maximum fp6 value
+    float max_fp6 = 7.5f;
+    // convert 0 float to fp6 and back, check if holds
+    ASSERT_NEAR(0.0f, type_convert<float>(f6_convert_rne(0.0f)), 0.0f);
+    // convert maximal f6_t to float and check if equal to max_fp6
+    ASSERT_NEAR(max_fp6, type_convert<float>(f6_convert_rne(max_fp6)), 0.0f);
+    // convert maximal float to fp6 and back, check if clipped to max_fp6
+    ASSERT_NEAR(
+        max_fp6, type_convert<float>(f6_convert_rne(std::numeric_limits<float>::max())), 0.0f);
+    // convert float Inf to fp6 and back, check if clipped to max_fp6
+    ASSERT_NEAR(
+        max_fp6, type_convert<float>(f6_convert_rne(std::numeric_limits<float>::infinity())), 0.0f);
+    // convert float value less than fp6 subnorm to fp6 and back, check if equal to 0.0
+    float less_than_subnorm = 0.0625f;
+    ASSERT_NEAR(0.0f, type_convert<float>(f6_convert_rne(less_than_subnorm)), 0.0f);
+    // convert float NaN to fp6 and back, check if clipped to max_fp6
+    ASSERT_NEAR(max_fp6,
+                type_convert<float>(f6_convert_rne(std::numeric_limits<float>::quiet_NaN())),
+                0.0f);
+    // positive norm float value to fp6 and back, check if holds
+    float pos_float = 1.0f;
+    ASSERT_NEAR(pos_float, type_convert<float>(f6_convert_rne(pos_float)), 0.0f);
+    // negative norm float value to fp6 and back, check if holds
+    float neg_float = -1.5f;
+    ASSERT_NEAR(neg_float, type_convert<float>(f6_convert_rne(neg_float)), 0.0f);
+    // positive subnorm float value to fp6 and back, check if holds
+    pos_float = 0.125f;
+    ASSERT_NEAR(pos_float, type_convert<float>(f6_convert_rne(pos_float)), 0.0f);
+    // negative subnorm float value to fp6 and back, check if holds
+    neg_float = -0.25f;
+    ASSERT_NEAR(neg_float, type_convert<float>(f6_convert_rne(neg_float)), 0.0f);
+}
+TEST(FP6, ConvertFP32Stochastic)
+{
+    // fix the tolerance value
+    float abs_tol = 1e-6;
+    // set maximum fp6 value
+    float max_fp6 = 7.5f;
+    // convert 0 float to fp6 and back, check if holds
+    ASSERT_NEAR(0.0f, type_convert<float>(f6_convert_sr(0.0f)), abs_tol);
+    // convert maximal f6_t to float and check if equal to max_fp6
+    ASSERT_NEAR(max_fp6, type_convert<float>(f6_convert_sr(max_fp6)), abs_tol);
+    // convert maximal float to fp6 and back, check if clipped to max_fp6
+    ASSERT_NEAR(
+        max_fp6, type_convert<float>(f6_convert_sr(std::numeric_limits<float>::max())), abs_tol);
+    // convert float Inf to fp6 and back, check if clipped to max_fp6
+    ASSERT_NEAR(max_fp6,
+                type_convert<float>(f6_convert_sr(std::numeric_limits<float>::infinity())),
+                abs_tol);
+    // convert float NaN to fp6 and back, check if clipped to max_fp6
+    ASSERT_NEAR(max_fp6,
+                type_convert<float>(f6_convert_sr(std::numeric_limits<float>::quiet_NaN())),
+                abs_tol);
+    // positive norm float value to fp6 and back, check if holds
+    float pos_float = 1.0f;
+    ASSERT_NEAR(pos_float, type_convert<float>(f6_convert_sr(pos_float)), abs_tol);
+    // negative norm float value to fp6 and back, check if holds
+    float neg_float = -1.5f;
+    ASSERT_NEAR(neg_float, type_convert<float>(f6_convert_sr(neg_float)), abs_tol);
+    // positive subnorm float value to fp6 and back, check if holds
+    pos_float = 0.125f;
+    ASSERT_NEAR(pos_float, type_convert<float>(f6_convert_sr(pos_float)), abs_tol);
+    // negative subnorm float value to fp6 and back, check if holds
+    neg_float = -0.25f;
+    ASSERT_NEAR(neg_float, type_convert<float>(f6_convert_sr(neg_float)), abs_tol);
+}
+TEST(FP6, ScaledConvertFP32Nearest)
+{
+    // set maximum scale
+    float max_scale = type_convert<float>(ck::NumericLimits<e8m0_bexp_t>::Max()); // 0xFE -> float
+    // set minimum scale
+    float min_scale = type_convert<float>(ck::NumericLimits<e8m0_bexp_t>::Min()); // 0x00 -> float
+    // set arbitrary scale to 256.0
+    float test_scale = 256.0f; // 0b10000111
+    // convert 0 float to fp6 and back with maximal scale, check if holds
+    ASSERT_NEAR(
+        0.0f, scaled_type_convert<float>(e8m0_bexp_t(max_scale), f6_convert_rne(0.0f)), 0.0f);
+    // convert 0 float to fp6 and back with minimal scale, check if holds
+    ASSERT_NEAR(
+        0.0f, scaled_type_convert<float>(e8m0_bexp_t(min_scale), f6_convert_rne(0.0f)), 0.0f);
+    // positive norm float value to fp6 and back with various scales, check if holds
+    float pos_float = 1.0f;
+    ASSERT_NEAR(pos_float * test_scale,
+                scaled_type_convert<float>(e8m0_bexp_t(test_scale), f6_convert_rne(pos_float)),
+                0.0f);
+    ASSERT_NEAR(pos_float * max_scale,
+                scaled_type_convert<float>(e8m0_bexp_t(max_scale), f6_convert_rne(pos_float)),
+                0.0f);
+    ASSERT_NEAR(pos_float * min_scale,
+                scaled_type_convert<float>(e8m0_bexp_t(min_scale), f6_convert_rne(pos_float)),
+                0.0f);
+    // negative norm float value to fp6 and back with various scales, check if holds
+    float neg_float = -1.5f;
+    ASSERT_NEAR(neg_float * test_scale,
+                scaled_type_convert<float>(e8m0_bexp_t(test_scale), f6_convert_rne(neg_float)),
+                0.0f);
+    ASSERT_NEAR(neg_float * max_scale,
+                scaled_type_convert<float>(e8m0_bexp_t(max_scale), f6_convert_rne(neg_float)),
+                0.0f);
+    ASSERT_NEAR(neg_float * min_scale,
+                scaled_type_convert<float>(e8m0_bexp_t(min_scale), f6_convert_rne(neg_float)),
+                0.0f);
+    // positive subnorm float value to fp6 and back with various scales, check if holds
+    pos_float = 0.125f;
+    ASSERT_NEAR(pos_float * test_scale,
+                scaled_type_convert<float>(e8m0_bexp_t(test_scale), f6_convert_rne(pos_float)),
+                0.0f);
+    ASSERT_NEAR(pos_float * max_scale,
+                scaled_type_convert<float>(e8m0_bexp_t(max_scale), f6_convert_rne(pos_float)),
+                0.0f);
+    ASSERT_NEAR(pos_float * min_scale,
+                scaled_type_convert<float>(e8m0_bexp_t(min_scale), f6_convert_rne(pos_float)),
+                0.0f);
+    // negative subnorm float value to fp6 and back with various scales, check if holds
+    neg_float = -0.25f;
+    ASSERT_NEAR(neg_float * test_scale,
+                scaled_type_convert<float>(e8m0_bexp_t(test_scale), f6_convert_rne(neg_float)),
+                0.0f);
+    ASSERT_NEAR(neg_float * max_scale,
+                scaled_type_convert<float>(e8m0_bexp_t(max_scale), f6_convert_rne(neg_float)),
+                0.0f);
+    ASSERT_NEAR(neg_float * min_scale,
+                scaled_type_convert<float>(e8m0_bexp_t(min_scale), f6_convert_rne(neg_float)),
+                0.0f);
+}
+TEST(FP6, ScaledConvertFP32Stochastic)
+{
+    // fix the tolerance value
+    float abs_tol = 1e-6;
+    // set maximum scale
+    float max_scale = type_convert<float>(ck::NumericLimits<e8m0_bexp_t>::Max()); // 0xFE -> float
+    // set minimum scale
+    float min_scale = type_convert<float>(ck::NumericLimits<e8m0_bexp_t>::Min()); // 0x00 -> float
+    // set arbitrary scale to 256.0
+    float test_scale = 256.0f; // 0b10000111
+    // convert 0 float to fp6 and back with maximal scale, check if holds
+    ASSERT_NEAR(
+        0.0f, scaled_type_convert<float>(e8m0_bexp_t(max_scale), f6_convert_sr(0.0f)), abs_tol);
+    // convert 0 float to fp6 and back with minimal scale, check if holds
+    ASSERT_NEAR(
+        0.0f, scaled_type_convert<float>(e8m0_bexp_t(min_scale), f6_convert_sr(0.0f)), abs_tol);
+    // positive norm float value to fp6 and back with various scales, check if holds
+    float pos_float = 1.0f;
+    ASSERT_NEAR(pos_float * test_scale,
+                scaled_type_convert<float>(e8m0_bexp_t(test_scale), f6_convert_sr(pos_float)),
+                abs_tol);
+    ASSERT_NEAR(pos_float * max_scale,
+                scaled_type_convert<float>(e8m0_bexp_t(max_scale), f6_convert_sr(pos_float)),
+                abs_tol);
+    ASSERT_NEAR(pos_float * min_scale,
+                scaled_type_convert<float>(e8m0_bexp_t(min_scale), f6_convert_sr(pos_float)),
+                abs_tol);
+    // negative norm float value to fp6 and back with various scales, check if holds
+    float neg_float = -1.5f;
+    ASSERT_NEAR(neg_float * test_scale,
+                scaled_type_convert<float>(e8m0_bexp_t(test_scale), f6_convert_sr(neg_float)),
+                abs_tol);
+    ASSERT_NEAR(neg_float * max_scale,
+                scaled_type_convert<float>(e8m0_bexp_t(max_scale), f6_convert_sr(neg_float)),
+                abs_tol);
+    ASSERT_NEAR(neg_float * min_scale,
+                scaled_type_convert<float>(e8m0_bexp_t(min_scale), f6_convert_sr(neg_float)),
+                abs_tol);
+    // positive subnorm float value to fp6 and back with various scales, check if holds
+    pos_float = 0.125f;
+    ASSERT_NEAR(pos_float * test_scale,
+                scaled_type_convert<float>(e8m0_bexp_t(test_scale), f6_convert_sr(pos_float)),
+                abs_tol);
+    ASSERT_NEAR(pos_float * max_scale,
+                scaled_type_convert<float>(e8m0_bexp_t(max_scale), f6_convert_sr(pos_float)),
+                abs_tol);
+    ASSERT_NEAR(pos_float * min_scale,
+                scaled_type_convert<float>(e8m0_bexp_t(min_scale), f6_convert_sr(pos_float)),
+                abs_tol);
+    // negative subnorm float value to fp6 and back with various scales, check if holds
+    neg_float = -0.25f;
+    ASSERT_NEAR(neg_float * test_scale,
+                scaled_type_convert<float>(e8m0_bexp_t(test_scale), f6_convert_sr(neg_float)),
+                abs_tol);
+    ASSERT_NEAR(neg_float * max_scale,
+                scaled_type_convert<float>(e8m0_bexp_t(max_scale), f6_convert_sr(neg_float)),
+                abs_tol);
+    ASSERT_NEAR(neg_float * min_scale,
+                scaled_type_convert<float>(e8m0_bexp_t(min_scale), f6_convert_sr(neg_float)),
+                abs_tol);
+}
+TEST(FP6, TestSize)
+{
+    ASSERT_EQ(1, sizeof(f6_t));
+    ASSERT_EQ(12, sizeof(f6x16_pk_t));
+    ASSERT_EQ(24, sizeof(f6x32_pk_t));
+    ASSERT_EQ(16, sizeof(vector_type<f6x16_pk_t, 1>));
+    ASSERT_EQ(32, sizeof(vector_type<f6x16_pk_t, 2>));
+    ASSERT_EQ(32, sizeof(vector_type<f6x32_pk_t, 1>));
+}
+TEST(FP6, TestAlignment)
+{
+    ASSERT_EQ(1, alignof(f6_t));
+    ASSERT_EQ(4, alignof(f6x16_pk_t));
+    ASSERT_EQ(4, alignof(f6x32_pk_t));
+    ASSERT_EQ(16, alignof(vector_type<f6x16_pk_t, 1>));
+    ASSERT_EQ(32, alignof(vector_type<f6x16_pk_t, 2>));
+    ASSERT_EQ(32, alignof(vector_type<f6x32_pk_t, 1>));
+}
+// test vector of 1 f6x16_pk_t, contains 16 f6_t
+TEST(FP6, TestAsType16x1)
+{
+    // test size
+    const int vector_size = 1;
+    const int packed_size = 16;
+    typedef int8_t test_vec_t __attribute__((ext_vector_type(16)));
+    test_vec_t test_vec = {f6_t(0b000000),
+                           f6_t(0b100000),
+                           f6_t(0b000001),
+                           f6_t(0b100001),
+                           f6_t(0b000010),
+                           f6_t(0b100010),
+                           f6_t(0b000011),
+                           f6_t(0b100011),
+                           f6_t(0b000100),
+                           f6_t(0b100100),
+                           f6_t(0b000101),
+                           f6_t(0b100101),
+                           f6_t(0b000110),
+                           f6_t(0b100110),
+                           f6_t(0b001011),
+                           f6_t(0b101011)};
+    // reference vector
+    vector_type<f6x16_pk_t, vector_size> right_vec;
+    // check default CTOR
+    ck::static_for<0, packed_size, 1>{}([&](auto i) {
+        ASSERT_EQ(
+            right_vec.template AsType<f6x16_pk_t>()(Number<0>{}).template unpack<>(Number<i>{}), 0);
+    });
+    // assign test values to the vector
+    ck::static_for<0, vector_size, 1>{}([&](auto i) {
+        right_vec.template AsType<f6x16_pk_t>()(Number<i>{}) = f6x16_pk_t{}.pack(test_vec);
+    });
+    // copy the vector
+    vector_type<f6x16_pk_t, vector_size> left_vec{right_vec};
+    // check if values were copied correctly
+    ck::static_for<0, packed_size, 1>{}([&](auto i) {
+        ASSERT_EQ(
+            left_vec.template AsType<f6x16_pk_t>()(Number<0>{}).template unpack<>(Number<i>{}),
+            static_cast<f6_t>(test_vec[static_cast<int>(i)]));
+    });
+}
+// test vector of 2 f6x16_pk_t, contains 32 f6_t
+TEST(FP6, TestAsType16x2)
+{
+    // test size
+    const int vector_size = 2;
+    const int packed_size = 16;
+    typedef int8_t test_vec_t __attribute__((ext_vector_type(16)));
+    test_vec_t test_vec[2];
+    test_vec[0] = {f6_t(0b000000),
+                   f6_t(0b100000),
+                   f6_t(0b000001),
+                   f6_t(0b100001),
+                   f6_t(0b000010),
+                   f6_t(0b100010),
+                   f6_t(0b000011),
+                   f6_t(0b100011),
+                   f6_t(0b000100),
+                   f6_t(0b100100),
+                   f6_t(0b000101),
+                   f6_t(0b100101),
+                   f6_t(0b000110),
+                   f6_t(0b100110),
+                   f6_t(0b001011),
+                   f6_t(0b101011)};
+    test_vec[1] = {f6_t(0b010000),
+                   f6_t(0b110000),
+                   f6_t(0b010001),
+                   f6_t(0b110001),
+                   f6_t(0b010010),
+                   f6_t(0b110010),
+                   f6_t(0b010011),
+                   f6_t(0b110011),
+                   f6_t(0b010100),
+                   f6_t(0b110100),
+                   f6_t(0b010101),
+                   f6_t(0b110101),
+                   f6_t(0b010110),
+                   f6_t(0b110110),
+                   f6_t(0b011011),
+                   f6_t(0b111011)};
+    // reference vector
+    vector_type<f6x16_pk_t, vector_size> right_vec;
+    // check default CTOR
+    ck::static_for<0, vector_size, 1>{}([&](auto idx_vector) {
+        ck::static_for<0, packed_size, 1>{}([&](auto idx_element) {
+            ASSERT_EQ(right_vec.template AsType<f6x16_pk_t>()(Number<idx_vector>{})
+                          .template unpack<>(Number<idx_element>{}),
+                      0);
+        });
+    });
+    // assign test values to the vector
+    ck::static_for<0, vector_size, 1>{}([&](auto i) {
+        right_vec.template AsType<f6x16_pk_t>()(Number<i>{}) = f6x16_pk_t{}.pack(test_vec[i]);
+    });
+    // copy the vector
+    vector_type<f6x16_pk_t, vector_size> left_vec{right_vec};
+    // check if values were copied correctly
+    ck::static_for<0, vector_size, 1>{}([&](auto idx_vector) {
+        ck::static_for<0, packed_size, 1>{}([&](auto idx_element) {
+            ASSERT_EQ(left_vec.template AsType<f6x16_pk_t>()(Number<idx_vector>{})
+                          .template unpack<>(Number<idx_element>{}),
+                      static_cast<f6_t>(test_vec[idx_vector][static_cast<int>(idx_element)]));
+        });
+    });
+}
+// test vector of 1 f6x32_pk_t, contains 32 f6_t
+TEST(FP6, TestAsType32x1)
+{
+    // test size
+    const int vector_size = 1;
+    const int packed_size = 32;
+    typedef int8_t test_vec_t __attribute__((ext_vector_type(32)));
+    test_vec_t test_vec = {f6_t(0b000000), f6_t(0b100000), f6_t(0b000001), f6_t(0b100001),
+                           f6_t(0b000010), f6_t(0b100010), f6_t(0b000011), f6_t(0b100011),
+                           f6_t(0b000100), f6_t(0b100100), f6_t(0b000101), f6_t(0b100101),
+                           f6_t(0b000110), f6_t(0b100110), f6_t(0b001011), f6_t(0b101011),
+                           f6_t(0b010000), f6_t(0b110000), f6_t(0b010001), f6_t(0b110001),
+                           f6_t(0b010010), f6_t(0b110010), f6_t(0b010011), f6_t(0b110011),
+                           f6_t(0b010100), f6_t(0b110100), f6_t(0b010101), f6_t(0b110101),
+                           f6_t(0b010110), f6_t(0b110110), f6_t(0b011011), f6_t(0b111011)};
+    // reference vector
+    vector_type<f6x32_pk_t, vector_size> right_vec;
+    // check default CTOR
+    ck::static_for<0, packed_size, 1>{}([&](auto i) {
+        ASSERT_EQ(
+            right_vec.template AsType<f6x32_pk_t>()(Number<0>{}).template unpack<>(Number<i>{}), 0);
+    });
+    // assign test values to the vector
+    ck::static_for<0, vector_size, 1>{}([&](auto i) {
+        right_vec.template AsType<f6x32_pk_t>()(Number<i>{}) = f6x32_pk_t{}.pack(test_vec);
+    });
+    // copy the vector
+    vector_type<f6x32_pk_t, vector_size> left_vec{right_vec};
+    // check if values were copied correctly
+    ck::static_for<0, packed_size, 1>{}([&](auto i) {
+        ASSERT_EQ(
+            left_vec.template AsType<f6x32_pk_t>()(Number<0>{}).template unpack<>(Number<i>{}),
+            static_cast<f6_t>(test_vec[static_cast<int>(i)]));
+    });
+}
--- a/test/data_type/test_fp8_ocp.cpp
+++ b/test/data_type/test_fp8_ocp.cpp
@@ -60,8 +60,8 @@ TEST(FP8OCP, ConvertFP32Nearest)
    float neg_float = -0.015625f; //-2^-6
    ASSERT_NEAR(neg_float, type_convert<float>(f8_convert_rne<f8_ocp_t>(neg_float)), 0.0f);
-    // positive subnorm float value to fp8 and back, check if holds
+    // positive subnorm fp8 value to fp8 and back, check if holds
-    pos_float = 0.00390625f;
+    pos_float = 0.00390625f; // 2^-8
    ASSERT_NEAR(pos_float, type_convert<float>(f8_convert_rne<f8_ocp_t>(pos_float)), abs_tol);
    // min subnorm fp8 value to fp8 and back, check if holds

--- a/test/data_type/test_mx_bf8.cpp
+++ b/test/data_type/test_mx_bf8.cpp
+// SPDX-License-Identifier: MIT
+// Copyright (c) 2024-2025, Advanced Micro Devices, Inc. All rights reserved.
+#include "gtest/gtest.h"
+#include "ck/library/utility/device_memory.hpp"
+#include "ck/utility/scaled_type_convert.hpp"
+using ck::bf8_ocp_t;
+using ck::bf8x16_ocp_t;
+using ck::bf8x2_ocp_t;
+using ck::bf8x32_ocp_t;
+using ck::e8m0_bexp_t;
+using ck::float16_t;
+using ck::float2_t;
+using ck::float32_t;
+using ck::mxf8_convert_rne;
+using ck::mxf8_convert_sr;
+using ck::scaled_type_convert;
+using ck::type_convert;
+constexpr uint64_t test_size = 256 * 256 + 2 + 4 + 6;
+/**
+ * @brief  Tests conversion of BF8 values to float using E8M0 exponent scaling.
+ *
+ * This function performs a series of conversions from BF8 values to float values using
+ * E8M0 exponent scaling. It handles all possible combinations of E8M0 and BF8 values,
+ * as well as specific vector and rounding conversions.
+ *
+ * @param N The maximum number of conversions to perform.
+ * @param p_test Pointer to the output array where the converted float values will be stored.
+ * @param p_completed Pointer to a variable that tracks the number of completed conversions.
+ *
+ * @note If either p_test or p_completed is nullptr, the function will return immediately.
+ * @note The function will stop converting if the number of conversions reaches N.
+ * @note First 256*256 conversions are for all possible combinations of E8M0 and BF8 values that are
+ * stored in memory sequentially with BF8 values varying faster.
+ *
+ * The function performs the following conversions:
+ * - All possible combinations of E8M0 and BF8 values. [256x256]
+ * - Vector conversions bf8x2 -> f32x2. [2]
+ * - Vector conversions  f32x2 -> bf8x2 rne. [2]
+ * - Vector conversions  f32x2 -> bf8x2 sr. [2]
+ * - Round to nearest even conversions for specific float values. [6]
+ *
+ * The results are stored in the p_test array, and the number of completed conversions
+ * is updated in the p_completed variable.
+ */
+__host__ __device__ void
+test_mx_bf8_scaled_convert(uint64_t N, float* p_test, uint64_t* p_completed)
+{
+    if(p_completed == nullptr)
+    {
+        return;
+    }
+    uint64_t& i = *p_completed;
+    i           = 0;
+    if(p_test == nullptr)
+    {
+        return;
+    }
+    // All possible combinations of E8M0 and BF8
+    for(ck::index_t exp_id = 0; exp_id < 256; exp_id++)
+    {
+        for(ck::index_t bf8_id = 0; bf8_id < 256; bf8_id++)
+        {
+            uint8_t bf8_uid = static_cast<uint8_t>(bf8_id);
+            auto v          = scaled_type_convert<float>(e8m0_bexp_t(exp_id), bf8_ocp_t{bf8_uid});
+            p_test[i]       = v;
+            i++;
+            if(i >= N)
+            {
+                return;
+            }
+        }
+    }
+    /// Test vector conversions
+    // bf8x2 -> f32x2
+    bf8x2_ocp_t bf8x2{bf8x2_ocp_t::data_v{0b10000100, 0b00000001}}; //-2^-14, 2^-16
+    auto scale = e8m0_bexp_t(8.0f);
+    float2_t f32x2 = scaled_type_convert<float2_t>(scale, bf8x2);
+    p_test[i++]    = f32x2[0];
+    if(i >= N)
+    {
+        return;
+    }
+    p_test[i++] = f32x2[1];
+    if(i >= N)
+    {
+        return;
+    }
+    // f32x2 -> bf8x2
+    f32x2       = {-8.0f, 4.0f};
+    auto scale2 = e8m0_bexp_t(2.0f);
+    bf8x2 = mxf8_convert_rne<bf8x2_ocp_t>(f32x2, type_convert<float>(scale2)); // expect {-4, 2}
+    p_test[i++] = type_convert<float>(bf8x2.AsType<bf8_ocp_t>()(ck::Number<0>{})); //-4f
+    if(i >= N)
+    {
+        return;
+    }
+    p_test[i++] = type_convert<float>(bf8x2.AsType<bf8_ocp_t>()(ck::Number<1>{})); // 2f
+    if(i >= N)
+    {
+        return;
+    }
+    auto scale4 = e8m0_bexp_t(4.0f);
+    bf8x2 = mxf8_convert_sr<bf8x2_ocp_t>(f32x2, type_convert<float>(scale4)); // expect {-2, 1}
+    p_test[i++] = type_convert<float>(bf8x2.AsType<bf8_ocp_t>()(ck::Number<0>{})); //-2f
+    if(i >= N)
+    {
+        return;
+    }
+    p_test[i++] = type_convert<float>(bf8x2.AsType<bf8_ocp_t>()(ck::Number<1>{})); // 1f
+    if(i >= N)
+    {
+        return;
+    }
+    /// Test round to nearest even
+    p_test[i++] = type_convert<float>(mxf8_convert_rne<bf8_ocp_t>(1024.0f, 4.0f)); // 1024/4
+    if(i >= N)
+    {
+        return;
+    }
+    p_test[i++] = type_convert<float>(
+        mxf8_convert_rne<bf8_ocp_t>(std::numeric_limits<float>::quiet_NaN(), 4.0f)); // => NaN
+    if(i >= N)
+    {
+        return;
+    }
+    p_test[i++] = type_convert<float>(mxf8_convert_rne<bf8_ocp_t>(
+        std::numeric_limits<float>::infinity(), 2.0f)); // => BF8 Inf on device
+    if(i >= N)
+    {
+        return;
+    }
+    // 31000/0.5 > 57344 => BF8 Inf on device
+    p_test[i++] = type_convert<float>(mxf8_convert_rne<bf8_ocp_t>(31000.0f, 0.5f));
+    if(i >= N)
+    {
+        return;
+    }
+    // -31000/0.5 < -57344  => -BF8 Inf on device
+    p_test[i++] = type_convert<float>(mxf8_convert_rne<bf8_ocp_t>(-31000.0f, 0.5f));
+    if(i >= N)
+    {
+        return;
+    }
+    p_test[i++] = type_convert<float>(
+        mxf8_convert_rne<bf8_ocp_t>(powf(2.0f, 16.0f), 4.0f)); // 2^16/4 = 65536/4
+    if(i >= N)
+    {
+        return;
+    }
+}
+TEST(MXBF8, HostScaledConvert)
+{
+    std::vector<float> out(test_size, -1.0f);
+    uint64_t completed = 0;
+    test_mx_bf8_scaled_convert(test_size, out.data(), &completed);
+    // V = X * P; X - E8M0 scale, P - BF8
+    // If X = NaN, then V = NaN regardless of P
+    uint8_t e8m0_nan_id = ck::NumericLimits<e8m0_bexp_t>::QuietNaN().data;
+    for(ck::index_t bf8_id = 0; bf8_id < 256; bf8_id++)
+    {
+        auto idx = e8m0_nan_id * 256 + bf8_id;
+        ASSERT_TRUE(std::isnan(out[idx]));
+    }
+    // If P in {Inf, NaN}, then V = P
+    std::set<uint8_t> bf8_spec_ids;
+    bf8_spec_ids.insert(0b11111111); // -NaN
+    bf8_spec_ids.insert(0b01111111); // +NaN
+    bf8_spec_ids.insert(0b11111101); // -NaN
+    bf8_spec_ids.insert(0b01111101); // +NaN
+    bf8_spec_ids.insert(0b11111110); // -NaN
+    bf8_spec_ids.insert(0b01111110); // +NaN
+    bf8_spec_ids.insert(0b11111100); // -inf
+    bf8_spec_ids.insert(0b01111100); // +inf
+    for(ck::index_t exp_id = 0; exp_id < 256; exp_id++)
+    {
+        if(exp_id == e8m0_nan_id)
+            continue;
+        for(auto bf8_spec_id : bf8_spec_ids)
+        {
+            auto idx = exp_id * 256 + bf8_spec_id;
+            if(std::isnan(type_convert<float>(bf8_ocp_t{bf8_spec_id})))
+            {
+                ASSERT_TRUE(std::isnan(out[idx]))
+                    << "exp_id: " << exp_id << " bf8_id: " << bf8_spec_id << std::endl
+                    << type_convert<float>(e8m0_bexp_t(exp_id)) << " * "
+                    << type_convert<float>(bf8_ocp_t{bf8_spec_id}) << " != " << out[idx];
+            }
+            else
+            {
+                ASSERT_EQ(out[idx], type_convert<float>(bf8_ocp_t{bf8_spec_id}))
+                    << "exp_id: " << exp_id << " bf8_id: " << bf8_spec_id << std::endl
+                    << type_convert<float>(e8m0_bexp_t(exp_id)) << " * "
+                    << type_convert<float>(bf8_ocp_t{bf8_spec_id}) << " != " << out[idx];
+            }
+        }
+    }
+    // V = X * P; X, P - finite
+    for(ck::index_t exp_id = 0; exp_id < 256; exp_id++)
+    {
+        if(exp_id == e8m0_nan_id)
+            continue;
+        for(ck::index_t bf8_id = 0; bf8_id < 256; bf8_id++)
+        {
+            if(bf8_spec_ids.find(bf8_id) != bf8_spec_ids.end())
+                continue;
+            uint8_t bf8_uid = static_cast<uint8_t>(bf8_id);
+            auto idx        = exp_id * 256 + bf8_uid;
+            ASSERT_FLOAT_EQ(out[idx],
+                            type_convert<float>(e8m0_bexp_t(exp_id)) *
+                                type_convert<float>(bf8_ocp_t{bf8_uid}))
+                << "exp_id: " << exp_id << " bf8_id: " << bf8_uid << std::endl
+                << type_convert<float>(e8m0_bexp_t(exp_id)) << " * "
+                << type_convert<float>(bf8_ocp_t{bf8_uid});
+        }
+    }
+    /// Test vector conversions
+    auto i = 256 * 256;
+    // bf8x2 -> f32x2
+    EXPECT_EQ(out[i++], -powf(2.0f, -11.0f));
+    EXPECT_EQ(out[i++], powf(2.0f, -13.0f));
+    // f32x2 -> bf8x2
+    // RNE
+    EXPECT_EQ(out[i++], -4.0f);
+    EXPECT_EQ(out[i++], 2.0f);
+    // SR
+    EXPECT_EQ(out[i++], -2.0f);
+    EXPECT_EQ(out[i++], 1.0f);
+    /// Test round to nearest even
+    EXPECT_EQ(out[i++], 1024.0f / 4.0f) << "out[i-1]: " << out[i - 1];
+    EXPECT_TRUE(std::isnan(out[i++])) << "out[i-1]: " << out[i - 1];
+    EXPECT_EQ(out[i++], type_convert<float>(ck::NumericLimits<bf8_ocp_t>::Max()))
+        << "out[i-1]: " << out[i - 1];
+    EXPECT_EQ(out[i++], type_convert<float>(ck::NumericLimits<bf8_ocp_t>::Max()))
+        << "out[i-1]: " << out[i - 1];
+    EXPECT_EQ(out[i++], type_convert<float>(ck::NumericLimits<bf8_ocp_t>::Lowest()))
+        << "out[i-1]: " << out[i - 1];
+    EXPECT_EQ(out[i++], powf(2.0f, 14.0f)) << "out[i-1]: " << out[i - 1];
+    EXPECT_EQ(test_size, completed);
+    EXPECT_EQ(test_size, i);
+}
+__global__ void test_mx_bf8_device_scaled_convert(uint64_t N, float* p_test, uint64_t* p_completed)
+{
+    test_mx_bf8_scaled_convert(N, p_test, p_completed);
+}
+TEST(MXBF8, DeviceScaledConvert)
+{
+    std::vector<float> out(test_size, -1.0f);
+    DeviceMem device_out(test_size * sizeof(float));
+    DeviceMem device_completed(sizeof(uint64_t));
+    device_out.SetValue(-21.0f);
+    device_completed.SetValue(-21.0f);
+    test_mx_bf8_device_scaled_convert<<<1, 1>>>(
+        test_size,
+        static_cast<float*>(device_out.GetDeviceBuffer()),
+        static_cast<uint64_t*>(device_completed.GetDeviceBuffer()));
+    uint64_t completed = 0;
+    device_completed.FromDevice(&completed);
+    device_out.FromDevice(out.data());
+    // V = X * P; X - E8M0 scale, P - BF8
+    // If X = NaN, then V = NaN regardless of P
+    uint8_t e8m0_nan_id = ck::NumericLimits<e8m0_bexp_t>::QuietNaN().data;
+    for(ck::index_t bf8_id = 0; bf8_id < 256; bf8_id++)
+    {
+        auto idx = e8m0_nan_id * 256 + bf8_id;
+        ASSERT_TRUE(std::isnan(out[idx])) << "idx: " << idx << " out[idx]: " << out[idx];
+    }
+    // If P in {Inf, NaN}, then V = P
+    std::set<uint8_t> bf8_spec_ids;
+    bf8_spec_ids.insert(0b11111111); //-NaN
+    bf8_spec_ids.insert(0b01111111); // +NaN
+    bf8_spec_ids.insert(0b11111101); //-NaN
+    bf8_spec_ids.insert(0b01111101); // +NaN
+    bf8_spec_ids.insert(0b11111110); //-NaN
+    bf8_spec_ids.insert(0b01111110); // +NaN
+    bf8_spec_ids.insert(0b11111100); //-inf
+    bf8_spec_ids.insert(0b01111100); // +inf
+    for(ck::index_t exp_id = 0; exp_id < 256; exp_id++)
+    {
+        if(exp_id == e8m0_nan_id)
+            continue;
+        for(auto bf8_spec_id : bf8_spec_ids)
+        {
+            auto idx = exp_id * 256 + bf8_spec_id;
+            if(std::isnan(type_convert<float>(bf8_ocp_t{bf8_spec_id})))
+            {
+                ASSERT_TRUE(std::isnan(out[idx]))
+                    << "exp_id: " << exp_id << " bf8_id: " << bf8_spec_id << std::endl
+                    << type_convert<float>(e8m0_bexp_t(exp_id)) << " * "
+                    << type_convert<float>(bf8_ocp_t{bf8_spec_id}) << " != " << out[idx];
+            }
+            else
+            {
+                ASSERT_EQ(out[idx], type_convert<float>(bf8_ocp_t{bf8_spec_id}))
+                    << "exp_id: " << exp_id << " bf8_id: " << bf8_spec_id << std::endl
+                    << type_convert<float>(e8m0_bexp_t(exp_id)) << " * "
+                    << type_convert<float>(bf8_ocp_t{bf8_spec_id}) << " != " << out[idx];
+            }
+        }
+    }
+    for(ck::index_t exp_id = 0; exp_id < 256; exp_id++)
+    {
+        if(exp_id == e8m0_nan_id)
+            continue;
+        for(ck::index_t bf8_id = 0; bf8_id < 256; bf8_id++)
+        {
+            if(bf8_spec_ids.find(bf8_id) != bf8_spec_ids.end())
+                continue;
+            uint8_t bf8_uid = static_cast<uint8_t>(bf8_id);
+            auto idx        = exp_id * 256 + bf8_uid;
+            ASSERT_FLOAT_EQ(out[idx],
+                            type_convert<float>(e8m0_bexp_t(exp_id)) *
+                                type_convert<float>(bf8_ocp_t{bf8_uid}))
+                << "exp_id: " << exp_id << " bf8_id: " << bf8_uid << std::endl
+                << type_convert<float>(e8m0_bexp_t(exp_id)) << " * "
+                << type_convert<float>(bf8_ocp_t{bf8_uid});
+        }
+    }
+    /// Test vector conversions
+    auto i = 256 * 256;
+    // bf8x2 -> f32x2
+    EXPECT_EQ(out[i++], -powf(2.0f, -11.0f));
+    EXPECT_EQ(out[i++], powf(2.0f, -13.0f));
+    // f32x2 -> bf8x2
+    // RNE
+    EXPECT_EQ(out[i++], -4.0f);
+    EXPECT_EQ(out[i++], 2.0f);
+    // SR
+    EXPECT_EQ(out[i++], -2.0f);
+    EXPECT_EQ(out[i++], 1.0f);
+    /// Test round to nearest even
+    EXPECT_EQ(out[i++], 1024.0f / 4.0f) << "out[i-1]: " << out[i - 1];
+    EXPECT_TRUE(std::isnan(out[i++])) << "out[i-1]: " << out[i - 1];
+#if 1
+    EXPECT_TRUE(std::isinf(out[i++])) << "out[i-1]: " << out[i - 1];
+    EXPECT_TRUE(std::isinf(out[i++])) << "out[i-1]: " << out[i - 1];
+    EXPECT_TRUE(std::isinf(out[i++])) << "out[i-1]: " << out[i - 1];
+#else
+    // NOTE: Host and Device have different behavior.
+    // Device returns Infs, while Host returns Max (saturation to finite value).
+    EXPECT_EQ(out[i++], type_convert<float>(ck::NumericLimits<bf8_ocp_t>::Max()))
+        << "out[i-1]: " << out[i - 1];
+    EXPECT_EQ(out[i++], type_convert<float>(ck::NumericLimits<bf8_ocp_t>::Max()))
+        << "out[i-1]: " << out[i - 1];
+    EXPECT_EQ(out[i++], type_convert<float>(ck::NumericLimits<bf8_ocp_t>::Lowest()))
+        << "out[i-1]: " << out[i - 1];
+#endif
+    EXPECT_EQ(out[i++], powf(2.0f, 14.0f)) << "out[i-1]: " << out[i - 1];
+    EXPECT_EQ(test_size, completed);
+    EXPECT_EQ(test_size, i);
+}
+__host__ __device__ float vec16_generator(ck::index_t i) { return powf(-1.0f, i) * powf(2.0f, i); }
+__global__ void test_mx_bf8x16_device_scaled_convert(float* p_test, uint64_t* p_completed)
+{
+    constexpr int N = 16;
+    if(p_completed == nullptr)
+    {
+        return;
+    }
+    uint64_t& i = *p_completed;
+    i           = 0;
+    if(p_test == nullptr)
+    {
+        return;
+    }
+    auto scale2 = e8m0_bexp_t(2.0f);
+    bf8x16_ocp_t bf8x16{};
+    float16_t float16{};
+    ck::static_for<0, N, 1>{}(
+        [&](auto ii) { float16[static_cast<int>(ii)] = vec16_generator(ii); });
+    bf8x16 = scaled_type_convert<bf8x16_ocp_t>(scale2, float16);
+    ck::static_for<0, N, 1>{}([&](auto ii) {
+        p_test[i++] = type_convert<float>(bf8x16.AsType<bf8_ocp_t>()(ck::Number<ii>{}));
+    });
+}
+TEST(MXBF8, DeviceF32x16ToBF8x16ScaledConvert)
+{
+    constexpr int N = 16;
+    std::vector<float> out(N, -1.0f);
+    DeviceMem device_out(N * sizeof(float));
+    DeviceMem device_completed(sizeof(uint64_t));
+    device_out.SetValue(-21.0f);
+    device_completed.SetValue(-21.0f);
+    test_mx_bf8x16_device_scaled_convert<<<1, 1>>>(
+        static_cast<float*>(device_out.GetDeviceBuffer()),
+        static_cast<uint64_t*>(device_completed.GetDeviceBuffer()));
+    uint64_t completed = 0;
+    device_completed.FromDevice(&completed);
+    device_out.FromDevice(out.data());
+    auto i = 0;
+    ck::static_for<0, N, 1>{}([&](auto ii) {
+        EXPECT_EQ(out[i++], vec16_generator(ii) / 2.0f) << "ii: " << ii << std::endl;
+    });
+    EXPECT_EQ(N, completed);
+    EXPECT_EQ(N, i);
+}
+__host__ __device__ float vec32_generator(ck::index_t i)
+{
+    if(i < 16)
+    {
+        return vec16_generator(i % 16);
+    }
+    else
+    {
+        return 1.5f * vec16_generator(i % 16);
+    }
+}
+__global__ void test_mx_bf8x32_device_scaled_convert(float* p_test, uint64_t* p_completed)
+{
+    constexpr int N = 32;
+    if(p_completed == nullptr)
+    {
+        return;
+    }
+    uint64_t& i = *p_completed;
+    i           = 0;
+    if(p_test == nullptr)
+    {
+        return;
+    }
+    auto scale2 = e8m0_bexp_t(2.0f);
+    bf8x32_ocp_t bf8x32{};
+    float32_t float32{};
+    ck::static_for<0, N, 1>{}(
+        [&](auto ii) { float32[static_cast<int>(ii)] = vec32_generator(ii); });
+    bf8x32 = mxf8_convert_rne<bf8x32_ocp_t>(float32, type_convert<float>(scale2));
+    ck::static_for<0, N, 1>{}([&](auto ii) {
+        p_test[i++] = type_convert<float>(bf8x32.AsType<bf8_ocp_t>()(ck::Number<ii>{}));
+    });
+}
+TEST(MXBF8, DeviceF32x32ToBF8x32ScaledConvert)
+{
+    constexpr int N = 32;
+    std::vector<float> out(N, -1.0f);
+    DeviceMem device_out(N * sizeof(float));
+    DeviceMem device_completed(sizeof(uint64_t));
+    device_out.SetValue(-21.0f);
+    device_completed.SetValue(-21.0f);
+    test_mx_bf8x32_device_scaled_convert<<<1, 1>>>(
+        static_cast<float*>(device_out.GetDeviceBuffer()),
+        static_cast<uint64_t*>(device_completed.GetDeviceBuffer()));
+    uint64_t completed = 0;
+    device_completed.FromDevice(&completed);
+    device_out.FromDevice(out.data());
+    auto i = 0;
+    ck::static_for<0, N, 1>{}([&](auto ii) {
+        EXPECT_EQ(out[i++], vec32_generator(ii) / 2.0f) << "ii: " << ii << std::endl;
+    });
+    EXPECT_EQ(N, completed);
+    EXPECT_EQ(N, i);
+}
+__global__ void test_mx_bf8x32_device_scaled_convert_sr(float* p_test, uint64_t* p_completed)
+{
+    constexpr int N = 32;
+    if(p_completed == nullptr)
+    {
+        return;
+    }
+    uint64_t& i = *p_completed;
+    i           = 0;
+    if(p_test == nullptr)
+    {
+        return;
+    }
+    auto scale2 = e8m0_bexp_t(8.0f);
+    bf8x32_ocp_t bf8x32{};
+    float32_t float32{};
+    ck::static_for<0, N, 1>{}(
+        [&](auto ii) { float32[static_cast<int>(ii)] = vec32_generator(ii); });
+    bf8x32 = mxf8_convert_sr<bf8x32_ocp_t>(float32, type_convert<float>(scale2));
+    ck::static_for<0, N, 1>{}([&](auto ii) {
+        p_test[i++] = type_convert<float>(bf8x32.AsType<bf8_ocp_t>()(ck::Number<ii>{}));
+    });
+}
+TEST(MXBF8, DeviceF32x32ToBF8x32ScaledConvertSR)
+{
+    constexpr int N = 32;
+    std::vector<float> out(N, -1.0f);
+    DeviceMem device_out(N * sizeof(float));
+    DeviceMem device_completed(sizeof(uint64_t));
+    device_out.SetValue(-21.0f);
+    device_completed.SetValue(-21.0f);
+    test_mx_bf8x32_device_scaled_convert_sr<<<1, 1>>>(
+        static_cast<float*>(device_out.GetDeviceBuffer()),
+        static_cast<uint64_t*>(device_completed.GetDeviceBuffer()));
+    uint64_t completed = 0;
+    device_completed.FromDevice(&completed);
+    device_out.FromDevice(out.data());
+    auto i = 0;
+    ck::static_for<0, N, 1>{}([&](auto ii) {
+        EXPECT_EQ(out[i++], vec32_generator(ii) / 8.0f) << "ii: " << ii << std::endl;
+    });
+    EXPECT_EQ(N, completed);
+    EXPECT_EQ(N, i);
+}
+__global__ void test_mx_f32x32_device_scaled_convert(float* p_test, uint64_t* p_completed)
+{
+    constexpr int N = 32;
+    if(p_completed == nullptr)
+    {
+        return;
+    }
+    uint64_t& i = *p_completed;
+    i           = 0;
+    if(p_test == nullptr)
+    {
+        return;
+    }
+    auto scale2 = e8m0_bexp_t(4.0f);
+    bf8x32_ocp_t bf8x32{};
+    float32_t float32{};
+    ck::static_for<0, N, 1>{}([&](auto ii) {
+        bf8x32.AsType<bf8_ocp_t>()(ii) = type_convert<bf8_ocp_t>(vec32_generator(ii) / 16.0f);
+    });
+    float32 = scaled_type_convert<float32_t>(scale2, bf8x32);
+    ck::static_for<0, N, 1>{}([&](auto ii) { p_test[i++] = float32[static_cast<int>(ii)]; });
+}
+TEST(MXBF8, DeviceBF8x32ToF32x32ScaledConvert)
+{
+    constexpr int N = 32;
+    std::vector<float> out(N, -1.0f);
+    DeviceMem device_out(N * sizeof(float));
+    DeviceMem device_completed(sizeof(uint64_t));
+    device_out.SetValue(-21.0f);
+    device_completed.SetValue(-21.0f);
+    test_mx_f32x32_device_scaled_convert<<<1, 1>>>(
+        static_cast<float*>(device_out.GetDeviceBuffer()),
+        static_cast<uint64_t*>(device_completed.GetDeviceBuffer()));
+    uint64_t completed = 0;
+    device_completed.FromDevice(&completed);
+    device_out.FromDevice(out.data());
+    auto i = 0;
+    ck::static_for<0, N, 1>{}([&](auto ii) {
+        EXPECT_EQ(out[i++], vec32_generator(ii) / 4.0f) << "ii: " << ii << std::endl;
+    });
+    EXPECT_EQ(N, completed);
+    EXPECT_EQ(N, i);
+}
--- a/test/data_type/test_mx_fp8.cpp
+++ b/test/data_type/test_mx_fp8.cpp
+// SPDX-License-Identifier: MIT
+// Copyright (c) 2024-2025, Advanced Micro Devices, Inc. All rights reserved.
+#include "gtest/gtest.h"
+#include "ck/library/utility/device_memory.hpp"
+#include "ck/utility/scaled_type_convert.hpp"
+using ck::e8m0_bexp_t;
+using ck::f8_ocp_t;
+using ck::f8x16_ocp_t;
+using ck::f8x2_ocp_t;
+using ck::f8x32_ocp_t;
+using ck::float16_t;
+using ck::float2_t;
+using ck::float32_t;
+using ck::mxf8_convert_rne;
+using ck::mxf8_convert_sr;
+using ck::scaled_type_convert;
+using ck::type_convert;
+using ck::fp8_impl::fp8x2_storage_t;
+constexpr uint64_t test_size = 256 * 256 + 2 + 4 + 6;
+/**
+ * @brief Tests conversion of FP8 values to float using E8M0 exponent scaling.
+ *
+ * This function performs a series of conversions from FP8 values to float values using
+ * E8M0 exponent scaling. It handles all possible combinations of E8M0 and FP8 values,
+ * as well as specific vector and rounding conversions.
+ *
+ * @param N The maximum number of conversions to perform.
+ * @param p_test Pointer to the output array where the converted float values will be stored.
+ * @param p_completed Pointer to a variable that tracks the number of completed conversions.
+ *
+ * @note If either p_test or p_completed is nullptr, the function will return immediately.
+ * @note The function will stop converting if the number of conversions reaches N.
+ * @note First 256*256 conversions are for all possible combinations of E8M0 and FP8 values that are
+ * stored in memory sequentially with FP8 values varying faster.
+ *
+ * The function performs the following conversions:
+ * - All possible combinations of E8M0 and FP8 values. [256x256]
+ * - Vector conversions f8x2 -> f32x2. [2]
+ * - Vector conversions  f32x2 -> f8x2 rne. [2]
+ * - Vector conversions  f32x2 -> f8x2 sr. [2]
+ * - Round to nearest even conversions for specific float values. [6]
+ *
+ * The results are stored in the p_test array, and the number of completed conversions
+ * is updated in the p_completed variable.
+ */
+__host__ __device__ void
+test_mx_fp8_scaled_convert(uint64_t N, float* p_test, uint64_t* p_completed)
+{
+    if(p_completed == nullptr)
+    {
+        return;
+    }
+    uint64_t& i = *p_completed;
+    i           = 0;
+    if(p_test == nullptr)
+    {
+        return;
+    }
+    // All possible combinations of E8M0 and FP8
+    for(ck::index_t exp_id = 0; exp_id < 256; exp_id++)
+    {
+        for(ck::index_t fp8_id = 0; fp8_id < 256; fp8_id++)
+        {
+            uint8_t fp8_uid = static_cast<uint8_t>(fp8_id);
+            auto v          = scaled_type_convert<float>(e8m0_bexp_t(exp_id), f8_ocp_t{fp8_uid});
+            p_test[i]       = v;
+            i++;
+            if(i >= N)
+            {
+                return;
+            }
+        }
+    }
+    /// Test vector conversions
+    // f8x2 -> f32x2
+    f8x2_ocp_t fp8x2{f8x2_ocp_t::data_v{0b10001000, 0b00000001}}; //-2^-6, 2^-9
+    auto scale2 = e8m0_bexp_t(2.0f);
+    float2_t f32x2 = scaled_type_convert<float2_t>(scale2, fp8x2);
+    p_test[i++]    = f32x2[0];
+    if(i >= N)
+    {
+        return;
+    }
+    p_test[i++] = f32x2[1];
+    if(i >= N)
+    {
+        return;
+    }
+    // f32x2 -> f8x2
+    f32x2 = {-8.0f, 4.0f};
+    fp8x2 = mxf8_convert_rne<f8x2_ocp_t>(f32x2, type_convert<float>(scale2)); // expect {-4, 2}
+    p_test[i++] = type_convert<float>(fp8x2.AsType<f8_ocp_t>()(ck::Number<0>{})); //-4f
+    if(i >= N)
+    {
+        return;
+    }
+    p_test[i++] = type_convert<float>(fp8x2.AsType<f8_ocp_t>()(ck::Number<1>{})); // 2f
+    if(i >= N)
+    {
+        return;
+    }
+    auto scale4 = e8m0_bexp_t(4.0f);
+    fp8x2 = mxf8_convert_sr<f8x2_ocp_t>(f32x2, type_convert<float>(scale4)); // expect {-2, 1}
+    p_test[i++] = type_convert<float>(fp8x2.AsType<f8_ocp_t>()(ck::Number<0>{})); //-2f
+    if(i >= N)
+    {
+        return;
+    }
+    p_test[i++] = type_convert<float>(fp8x2.AsType<f8_ocp_t>()(ck::Number<1>{})); // 1f
+    if(i >= N)
+    {
+        return;
+    }
+    /// Test round to nearest even
+    p_test[i++] = type_convert<float>(mxf8_convert_rne<f8_ocp_t>(1024.0f, 4.0f)); // 1024/4
+    if(i >= N)
+    {
+        return;
+    }
+    p_test[i++] = type_convert<float>(
+        mxf8_convert_rne<f8_ocp_t>(std::numeric_limits<float>::quiet_NaN(), 4.0f)); // => NaN
+    if(i >= N)
+    {
+        return;
+    }
+    // Inf/2 > 448 => NaN on device
+    p_test[i++] = type_convert<float>(
+        mxf8_convert_rne<f8_ocp_t>(std::numeric_limits<float>::infinity(), 2.0f));
+    if(i >= N)
+    {
+        return;
+    }
+    // 256/0.5  > 448 => NaN on device
+    p_test[i++] = type_convert<float>(mxf8_convert_rne<f8_ocp_t>(256.0f, 0.5f));
+    if(i >= N)
+    {
+        return;
+    }
+    // -256/0.5  < -448 => NaN on device
+    p_test[i++] = type_convert<float>(mxf8_convert_rne<f8_ocp_t>(-256.0f, 0.5f));
+    if(i >= N)
+    {
+        return;
+    }
+    // proper scale selection 2^13 < 10000; 2^8 < 448 => scale = 2^(13-8) = 2^5
+    p_test[i++] =
+        type_convert<float>(mxf8_convert_rne<f8_ocp_t>(10000.0f, 32.0f)); // 10000/32 = 312.5
+    if(i >= N)
+    {
+        return;
+    }
+}
+TEST(MXFP8, HostScaledConvert)
+{
+    std::vector<float> out(test_size, -1.0f);
+    uint64_t completed = 0;
+    test_mx_fp8_scaled_convert(test_size, out.data(), &completed);
+    // V = X * P; X - E8M0 scale, P - FP8
+    // If X = NaN, then V = NaN regardless of P
+    uint8_t e8m0_nan_id = ck::NumericLimits<e8m0_bexp_t>::QuietNaN().data;
+    for(ck::index_t fp8_id = 0; fp8_id < 256; fp8_id++)
+    {
+        auto idx = e8m0_nan_id * 256 + fp8_id;
+        ASSERT_TRUE(std::isnan(out[idx]));
+    }
+    // If P in {Inf, NaN}, then V = P
+    std::set<uint8_t> fp8_nan_ids;
+    fp8_nan_ids.insert(0b11111111); //-NaN
+    fp8_nan_ids.insert(0b01111111); // +NaN
+    for(ck::index_t exp_id = 0; exp_id < 256; exp_id++)
+    {
+        if(exp_id == e8m0_nan_id)
+            continue;
+        for(auto fp8_nan_id : fp8_nan_ids)
+        {
+            auto idx = exp_id * 256 + fp8_nan_id;
+            ASSERT_TRUE(std::isnan(out[idx]));
+        }
+    }
+    for(ck::index_t exp_id = 0; exp_id < 256; exp_id++)
+    {
+        if(exp_id == e8m0_nan_id)
+            continue;
+        for(ck::index_t fp8_id = 0; fp8_id < 256; fp8_id++)
+        {
+            if(fp8_nan_ids.find(fp8_id) != fp8_nan_ids.end())
+                continue;
+            uint8_t fp8_uid = static_cast<uint8_t>(fp8_id);
+            auto idx        = exp_id * 256 + fp8_uid;
+            ASSERT_FLOAT_EQ(out[idx],
+                            type_convert<float>(e8m0_bexp_t(exp_id)) *
+                                type_convert<float>(f8_ocp_t{fp8_uid}))
+                << "exp_id: " << exp_id << " fp8_id: " << fp8_id << std::endl
+                << type_convert<float>(e8m0_bexp_t(exp_id)) << " * "
+                << type_convert<float>(f8_ocp_t{fp8_uid});
+        }
+    }
+    /// Test vector conversions
+    auto i = 256 * 256;
+    // f8x2 -> f32x2
+    EXPECT_EQ(out[i++], -powf(2.0f, -5.0f));
+    EXPECT_EQ(out[i++], powf(2.0f, -8.0f));
+    // f32x2 -> fp8x2
+    // RNE
+    EXPECT_EQ(out[i++], -4.0f);
+    EXPECT_EQ(out[i++], 2.0f);
+    // SR
+    EXPECT_EQ(out[i++], -2.0f);
+    EXPECT_EQ(out[i++], 1.0f);
+    /// Test round to nearest even
+    EXPECT_EQ(out[i++], 1024.0f / 4.0f) << "out[i-1]: " << out[i - 1];
+    EXPECT_TRUE(std::isnan(out[i++])) << "out[i-1]: " << out[i - 1];
+    EXPECT_EQ(out[i++], type_convert<float>(ck::NumericLimits<f8_ocp_t>::Max()))
+        << "out[i-1]: " << out[i - 1];
+    EXPECT_EQ(out[i++], type_convert<float>(ck::NumericLimits<f8_ocp_t>::Max()))
+        << "out[i-1]: " << out[i - 1];
+    EXPECT_EQ(out[i++], type_convert<float>(ck::NumericLimits<f8_ocp_t>::Lowest()))
+        << "out[i-1]: " << out[i - 1];
+    EXPECT_EQ(out[i++], type_convert<float>(type_convert<f8_ocp_t>(312.5f)))
+        << "out[i-1]: " << out[i - 1];
+    EXPECT_EQ(test_size, completed);
+    EXPECT_EQ(test_size, i);
+}
+__global__ void test_mx_fp8_device_scaled_convert(uint64_t N, float* p_test, uint64_t* p_completed)
+{
+    test_mx_fp8_scaled_convert(N, p_test, p_completed);
+}
+TEST(MXFP8, DeviceScaledConvert)
+{
+    std::vector<float> out(test_size, -1.0f);
+    DeviceMem device_out(test_size * sizeof(float));
+    DeviceMem device_completed(sizeof(uint64_t));
+    device_out.SetValue(-21.0f);
+    device_completed.SetValue(-21.0f);
+    test_mx_fp8_device_scaled_convert<<<1, 1>>>(
+        test_size,
+        static_cast<float*>(device_out.GetDeviceBuffer()),
+        static_cast<uint64_t*>(device_completed.GetDeviceBuffer()));
+    uint64_t completed = 0;
+    device_completed.FromDevice(&completed);
+    device_out.FromDevice(out.data());
+    // V = X * P; X - E8M0 scale, P - FP8
+    // If X = NaN, then V = NaN regardless of P
+    uint8_t e8m0_nan_id = ck::NumericLimits<e8m0_bexp_t>::QuietNaN().data;
+    for(ck::index_t fp8_id = 0; fp8_id < 256; fp8_id++)
+    {
+        auto idx = e8m0_nan_id * 256 + fp8_id;
+        ASSERT_TRUE(std::isnan(out[idx])) << "idx: " << idx << " out[idx]: " << out[idx];
+    }
+    // If P in {Inf, NaN}, then V = P
+    std::set<uint8_t> fp8_nan_ids;
+    fp8_nan_ids.insert(0b11111111); //-NaN
+    fp8_nan_ids.insert(0b01111111); // +NaN
+    for(ck::index_t exp_id = 0; exp_id < 256; exp_id++)
+    {
+        if(exp_id == e8m0_nan_id)
+            continue;
+        for(auto fp8_nan_id : fp8_nan_ids)
+        {
+            auto idx = exp_id * 256 + fp8_nan_id;
+            ASSERT_TRUE(std::isnan(out[idx])) << "idx: " << idx << " out[idx]: " << out[idx];
+        }
+    }
+    for(ck::index_t exp_id = 0; exp_id < 256; exp_id++)
+    {
+        if(exp_id == e8m0_nan_id)
+            continue;
+        for(ck::index_t fp8_id = 0; fp8_id < 256; fp8_id++)
+        {
+            if(fp8_nan_ids.find(fp8_id) != fp8_nan_ids.end())
+                continue;
+            uint8_t fp8_uid = static_cast<uint8_t>(fp8_id);
+            auto idx        = exp_id * 256 + fp8_uid;
+            ASSERT_FLOAT_EQ(out[idx],
+                            type_convert<float>(e8m0_bexp_t(exp_id)) *
+                                type_convert<float>(f8_ocp_t{fp8_uid}))
+                << "exp_id: " << exp_id << " fp8_id: " << fp8_id << std::endl
+                << type_convert<float>(e8m0_bexp_t(exp_id)) << " * "
+                << type_convert<float>(f8_ocp_t{fp8_uid});
+        }
+    }
+    /// Test vector conversions
+    auto i = 256 * 256;
+    // f8x2 -> f32x2
+    EXPECT_EQ(out[i++], -powf(2.0f, -5.0f));
+    EXPECT_EQ(out[i++], powf(2.0f, -8.0f));
+    // f32x2 -> fp8x2
+    // RNE
+    EXPECT_EQ(out[i++], -4.0f);
+    EXPECT_EQ(out[i++], 2.0f);
+    // SR
+    EXPECT_EQ(out[i++], -2.0f);
+    EXPECT_EQ(out[i++], 1.0f);
+    /// Test round to nearest even
+    EXPECT_EQ(out[i++], 1024.0f / 4.0f) << "out[i-1]: " << out[i - 1];
+    EXPECT_TRUE(std::isnan(out[i++])) << "out[i-1]: " << out[i - 1];
+#if 1
+    EXPECT_TRUE(std::isnan(out[i++])) << "out[i-1]: " << out[i - 1];
+    EXPECT_TRUE(std::isnan(out[i++])) << "out[i-1]: " << out[i - 1];
+    EXPECT_TRUE(std::isnan(out[i++])) << "out[i-1]: " << out[i - 1];
+#else
+    // NOTE: Host and Device have different behavior.
+    // Device returns NaN, while Host returns Max (saturation to finite value).
+    EXPECT_EQ(out[i++], type_convert<float>(ck::NumericLimits<f8_ocp_t>::Max()))
+        << "out[i-1]: " << out[i - 1];
+    EXPECT_EQ(out[i++], type_convert<float>(ck::NumericLimits<f8_ocp_t>::Max()))
+        << "out[i-1]: " << out[i - 1];
+    EXPECT_EQ(out[i++], type_convert<float>(ck::NumericLimits<f8_ocp_t>::Lowest()))
+        << "out[i-1]: " << out[i - 1];
+#endif
+    EXPECT_EQ(out[i++], type_convert<float>(type_convert<f8_ocp_t>(312.5f)))
+        << "out[i-1]: " << out[i - 1];
+    EXPECT_EQ(test_size, completed);
+    EXPECT_EQ(test_size, i);
+}
+__host__ __device__ float vec16_generator(ck::index_t i)
+{
+    return (i < 8 ? -1.0 : 1.0) * powf(2.0f, i % 8);
+}
+__global__ void test_mx_fp8x16_device_scaled_convert(float* p_test, uint64_t* p_completed)
+{
+    constexpr int N = 16;
+    if(p_completed == nullptr)
+    {
+        return;
+    }
+    uint64_t& i = *p_completed;
+    i           = 0;
+    if(p_test == nullptr)
+    {
+        return;
+    }
+    auto scale2 = e8m0_bexp_t(2.0f);
+    f8x16_ocp_t fp8x16{};
+    float16_t float16{};
+    ck::static_for<0, N, 1>{}(
+        [&](auto ii) { float16[static_cast<int>(ii)] = vec16_generator(ii); });
+    fp8x16 = scaled_type_convert<ck::f8x16_ocp_t>(scale2, float16);
+    ck::static_for<0, N, 1>{}([&](auto ii) {
+        p_test[i++] = type_convert<float>(fp8x16.AsType<f8_ocp_t>()(ck::Number<ii>{}));
+    });
+}
+TEST(MXFP8, DeviceF32x16ToF8x16ScaledConvert)
+{
+    constexpr int N = 16;
+    std::vector<float> out(N, -1.0f);
+    DeviceMem device_out(N * sizeof(float));
+    DeviceMem device_completed(sizeof(uint64_t));
+    device_out.SetValue(-21.0f);
+    device_completed.SetValue(-21.0f);
+    test_mx_fp8x16_device_scaled_convert<<<1, 1>>>(
+        static_cast<float*>(device_out.GetDeviceBuffer()),
+        static_cast<uint64_t*>(device_completed.GetDeviceBuffer()));
+    uint64_t completed = 0;
+    device_completed.FromDevice(&completed);
+    device_out.FromDevice(out.data());
+    auto i = 0;
+    ck::static_for<0, N, 1>{}([&](auto ii) {
+        EXPECT_EQ(out[i++], vec16_generator(ii) / 2.0f) << "ii: " << ii << std::endl;
+    });
+    EXPECT_EQ(N, completed);
+    EXPECT_EQ(N, i);
+}
+__host__ __device__ float vec32_generator(ck::index_t i)
+{
+    if(i < 16)
+    {
+        return vec16_generator(i % 16);
+    }
+    else
+    {
+        return 1.5f * vec16_generator(i % 16);
+    }
+}
+__global__ void test_mx_fp8x32_device_scaled_convert(float* p_test, uint64_t* p_completed)
+{
+    constexpr int N = 32;
+    if(p_completed == nullptr)
+    {
+        return;
+    }
+    uint64_t& i = *p_completed;
+    i           = 0;
+    if(p_test == nullptr)
+    {
+        return;
+    }
+    auto scale2 = e8m0_bexp_t(2.0f);
+    f8x32_ocp_t fp8x32{};
+    float32_t float32{};
+    ck::static_for<0, N, 1>{}(
+        [&](auto ii) { float32[static_cast<int>(ii)] = vec32_generator(ii); });
+    fp8x32 = mxf8_convert_rne<f8x32_ocp_t>(float32, type_convert<float>(scale2));
+    ck::static_for<0, N, 1>{}(
+        [&](auto ii) { p_test[i++] = type_convert<float>(fp8x32.AsType<f8_ocp_t>()(ii)); });
+}
+TEST(MXFP8, DeviceF32x32ToF8x32ScaledConvert)
+{
+    constexpr int N = 32;
+    std::vector<float> out(N, -1.0f);
+    DeviceMem device_out(N * sizeof(float));
+    DeviceMem device_completed(sizeof(uint64_t));
+    device_out.SetValue(-21.0f);
+    device_completed.SetValue(-21.0f);
+    test_mx_fp8x32_device_scaled_convert<<<1, 1>>>(
+        static_cast<float*>(device_out.GetDeviceBuffer()),
+        static_cast<uint64_t*>(device_completed.GetDeviceBuffer()));
+    uint64_t completed = 0;
+    device_completed.FromDevice(&completed);
+    device_out.FromDevice(out.data());
+    auto i = 0;
+    ck::static_for<0, N, 1>{}([&](auto ii) {
+        EXPECT_EQ(out[i++], vec32_generator(ii) / 2.0f) << "ii: " << ii << std::endl;
+    });
+    EXPECT_EQ(N, completed);
+    EXPECT_EQ(N, i);
+}
+__global__ void test_mx_fp8x32_device_scaled_convert_sr(float* p_test, uint64_t* p_completed)
+{
+    constexpr int N = 32;
+    if(p_completed == nullptr)
+    {
+        return;
+    }
+    uint64_t& i = *p_completed;
+    i           = 0;
+    if(p_test == nullptr)
+    {
+        return;
+    }
+    auto scale2 = e8m0_bexp_t(8.0f);
+    f8x32_ocp_t fp8x32{};
+    float32_t float32{};
+    ck::static_for<0, N, 1>{}(
+        [&](auto ii) { float32[static_cast<int>(ii)] = vec32_generator(ii); });
+    fp8x32 = mxf8_convert_sr<f8x32_ocp_t>(float32, type_convert<float>(scale2));
+    ck::static_for<0, N, 1>{}(
+        [&](auto ii) { p_test[i++] = type_convert<float>(fp8x32.AsType<f8_ocp_t>()(ii)); });
+}
+TEST(MXFP8, DeviceF32x32ToF8x32ScaledConvertSR)
+{
+    constexpr int N = 32;
+    std::vector<float> out(N, -1.0f);
+    DeviceMem device_out(N * sizeof(float));
+    DeviceMem device_completed(sizeof(uint64_t));
+    device_out.SetValue(-21.0f);
+    device_completed.SetValue(-21.0f);
+    test_mx_fp8x32_device_scaled_convert_sr<<<1, 1>>>(
+        static_cast<float*>(device_out.GetDeviceBuffer()),
+        static_cast<uint64_t*>(device_completed.GetDeviceBuffer()));
+    uint64_t completed = 0;
+    device_completed.FromDevice(&completed);
+    device_out.FromDevice(out.data());
+    auto i = 0;
+    ck::static_for<0, N, 1>{}([&](auto ii) {
+        EXPECT_EQ(out[i++], vec32_generator(ii) / 8.0f) << "ii: " << ii << std::endl;
+    });
+    EXPECT_EQ(N, completed);
+    EXPECT_EQ(N, i);
+}
+__global__ void test_mx_f32x32_device_scaled_convert(float* p_test, uint64_t* p_completed)
+{
+    constexpr int N = 32;
+    if(p_completed == nullptr)
+    {
+        return;
+    }
+    uint64_t& i = *p_completed;
+    i           = 0;
+    if(p_test == nullptr)
+    {
+        return;
+    }
+    auto scale2 = e8m0_bexp_t(4.0f);
+    f8x32_ocp_t fp8x32{};
+    float32_t float32{};
+    ck::static_for<0, N, 1>{}([&](auto ii) {
+        fp8x32.AsType<f8_ocp_t>()(ii) = type_convert<f8_ocp_t>(vec32_generator(ii) / 16.0f);
+    });
+    float32 = scaled_type_convert<float32_t>(scale2, fp8x32);
+    ck::static_for<0, N, 1>{}([&](auto ii) { p_test[i++] = float32[static_cast<int>(ii)]; });
+}
+TEST(MXFP8, DeviceF8x32ToF32x32ScaledConvert)
+{
+    constexpr int N = 32;
+    std::vector<float> out(N, -1.0f);
+    DeviceMem device_out(N * sizeof(float));
+    DeviceMem device_completed(sizeof(uint64_t));
+    device_out.SetValue(-21.0f);
+    device_completed.SetValue(-21.0f);
+    test_mx_f32x32_device_scaled_convert<<<1, 1>>>(
+        static_cast<float*>(device_out.GetDeviceBuffer()),
+        static_cast<uint64_t*>(device_completed.GetDeviceBuffer()));
+    uint64_t completed = 0;
+    device_completed.FromDevice(&completed);
+    device_out.FromDevice(out.data());
+    auto i = 0;
+    ck::static_for<0, N, 1>{}([&](auto ii) {
+        EXPECT_EQ(out[i++], vec32_generator(ii) / 4.0f) << "ii: " << ii << std::endl;
+    });
+    EXPECT_EQ(N, completed);
+    EXPECT_EQ(N, i);
+}
--- a/test/mx_mfma_op/CMakeLists.txt
+++ b/test/mx_mfma_op/CMakeLists.txt
+add_custom_target(test_mx_mfma)
+add_gtest_executable(test_mx_mfma_op mx_mfma_op.cpp)
+if(result EQUAL 0)
+    target_link_libraries(test_mx_mfma_op PRIVATE utility)
+endif()
+add_dependencies(test_mx_mfma test_mx_mfma_op)
--- a/test/mx_mfma_op/mx_mfma_op.cpp
+++ b/test/mx_mfma_op/mx_mfma_op.cpp
+// SPDX-License-Identifier: MIT
+// Copyright (c) 2025, Advanced Micro Devices, Inc. All rights reserved.
+#include "gtest/gtest.h"
+#include "mx_mfma_op.hpp"
+using ck::e8m0_bexp_t;
+using ck::f8_t;
+using ck::half_t;
+using ck::type_convert;
+/**
+ * @brief Run the test for the given MFMA instruction
+ *
+ * @param init - selects initialization algorithm for A and B tensors
+ */
+template <typename AType, typename BType, typename CType, ck::MFMA_F8F6F4 mfma>
+bool run_mfma_test(ck::index_t init)
+{
+    using ALayout = ck::tensor_layout::gemm::ColumnMajor;
+    using BLayout = ck::tensor_layout::gemm::ColumnMajor;
+    using CLayout = ck::tensor_layout::gemm::ColumnMajor;
+    using AccType    = float; // only MFMA_F32 instructions supported
+    using CPUAccType = AccType;
+    ck::mfma_type<static_cast<ck::MfmaInstr>(mfma)> mfma_instr;
+    constexpr auto BLOCK_M = mfma_instr.m_per_blk;
+    constexpr auto BLOCK_N = mfma_instr.n_per_blk;
+    constexpr auto BLOCK_K = mfma_instr.num_input_blks * mfma_instr.k_per_blk;
+    const auto mx_mfma_kernel = ck::matmul<AType, BType, CType, AccType, BLOCK_M, BLOCK_N, BLOCK_K>;
+    bool pass = true;
+    pass = ck::mfma_test::TestMFMA<decltype(mx_mfma_kernel),
+                                   AType,
+                                   BType,
+                                   CType,
+                                   AccType,
+                                   CPUAccType,
+                                   ALayout,
+                                   BLayout,
+                                   CLayout,
+                                   BLOCK_M,
+                                   BLOCK_N,
+                                   BLOCK_K>{}(mx_mfma_kernel, init);
+    return pass;
+}
+TEST(MFMA, FP8MFMA16x16x128)
+{
+    auto AB_init = 0;
+    auto pass    = run_mfma_test<f8_t, f8_t, half_t, ck::MFMA_F8F6F4::F32_16x16x128>(AB_init);
+    EXPECT_TRUE(pass);
+}
+TEST(MFMA, FP8MFMA32x32x64)
+{
+    auto AB_init = 0;
+    auto pass    = run_mfma_test<f8_t, f8_t, float, ck::MFMA_F8F6F4::F32_32x32x64>(AB_init);
+    EXPECT_TRUE(pass);
+}
--- a/test/mx_mfma_op/mx_mfma_op.hpp
+++ b/test/mx_mfma_op/mx_mfma_op.hpp
+#pragma once
+#include "ck/ck.hpp"
+#include "ck/library/utility/host_tensor.hpp"
+#include "ck/library/utility/device_memory.hpp"
+#include "ck/tensor_operation/gpu/device/tensor_layout.hpp"
+#include "ck/tensor_operation/gpu/warp/xdlops_gemm.hpp"
+#include "ck/library/utility/host_tensor_generator.hpp"
+#include "ck/library/reference_tensor_operation/cpu/reference_gemm.hpp"
+#include "ck/library/utility/check_err.hpp"
+namespace ck {
+// MFMA instructions supported in this test
+enum class MFMA_F8F6F4
+{
+    F32_16x16x128 =
+        static_cast<int>(MfmaInstr::mfma_f32_16x16x128f8f6f4), // V_MFMA_F32_16X16X128_F8F6F4
+    F32_32x32x64 =
+        static_cast<int>(MfmaInstr::mfma_f32_32x32x64f8f6f4) // V_MFMA_F32_32X32X64_F8F6F4
+};
+template <typename AFragT, typename BFragT, typename AccumFragT, int32_t BLOCK_M, int32_t BLOCK_N>
+struct mfma_type_selector;
+template <typename AFragT, typename BFragT, typename AccumFragT>
+struct mfma_type_selector<AFragT, BFragT, AccumFragT, 16, 16>
+{
+    __device__ void operator()(AFragT const& fragA, BFragT const& fragB, AccumFragT& fragAcc)
+    {
+        auto op = mfma_type<MfmaInstr::mfma_f32_16x16x128f8f6f4>{};
+        op.template run<16, 16, AFragT, BFragT, AccumFragT>(fragA, fragB, fragAcc);
+    }
+};
+template <typename AFragT, typename BFragT, typename AccumFragT>
+struct mfma_type_selector<AFragT, BFragT, AccumFragT, 32, 32>
+{
+    __device__ void operator()(AFragT const& fragA, BFragT const& fragB, AccumFragT& fragAcc)
+    {
+        auto op = mfma_type<MfmaInstr::mfma_f32_32x32x64f8f6f4>{};
+        op.template run<32, 32, AFragT, BFragT, AccumFragT>(fragA, fragB, fragAcc);
+    }
+};
+template <typename VecT>
+static constexpr int32_t vectorSize(const VecT&)
+{
+    return scalar_type<VecT>::vector_size;
+}
+// Define a load function for input A blocks:
+// Size: (BLOCK_M x BLOCK_K)
+// ASSUMPTION:
+// - We want contiguous BLOCK_M sized column neighbors in register.
+// - Data is in col_major format
+// This means:
+// - From A we will load K columns of size BLOCK_M to satisfy our input data
+template <typename AType, typename AFragT, int32_t BLOCK_M, int32_t BLOCK_K>
+__device__ AFragT load_A_col_major(AType const* input_ptr)
+{
+    // clang-format off
+    // Register Mapping for 16x128:                                                        ||    Register Mapping for 32x64:
+    // Size              |   BLOCK_M  |   BLOCK_M   |   BLOCK_M  |   BLOCK_M   |           ||    Size              |   BLOCK_M  |   BLOCK_M   |
+    // M                 | 0  ...  15 |  0  ...  15 | 0  ...  15 |  0  ...  15 |           ||    M                 | 0  ...  31 |  0  ...  31 |
+    // Thread Id         | 0  ...  15 | 16  ...  31 | 32  ... 47 | 48  ...  63 | Vector    ||    Thread Id         | 0  ...  31 | 32  ...  63 | Vector
+    // Register Element   ------------ ------------- ------------ -------------  Element   ||    Register Element   ------------ -------------  Element
+    // Reg 0 [0:7]       |     K0     |     K32     |     K64    |     K96     |  v[0]     ||    Reg 0 [0:7]       |     K0     |     K32     |  v[0]
+    // Reg 0 [8:15]      |     K1     |     K33     |     K65    |     K97     |  v[1]     ||    Reg 0 [8:15]      |     K1     |     K33     |  v[1]
+    // Reg 0 [16:23]     |     K2     |     K34     |     K66    |     K98     |  v[2]     ||    Reg 0 [16:23]     |     K2     |     K34     |  v[2]
+    // Reg 0 [24:31]     |     K3     |     K35     |     K67    |     K99     |  v[3]     ||    Reg 0 [24:31]     |     K3     |     K35     |  v[3]
+    // Reg 1 [0:7]       |     K4     |     K36     |     K68    |     K100    |  v[4]     ||    Reg 1 [0:7]       |     K4     |     K36     |  v[4]
+    // Reg 1 [8:15]      |     K5     |     K37     |     K69    |     K101    |  v[5]     ||    Reg 1 [8:15]      |     K5     |     K37     |  v[5]
+    // Reg 1 [16:23]     |     K6     |     K38     |     K70    |     K102    |  v[6]     ||    Reg 1 [16:23]     |     K6     |     K38     |  v[6]
+    // Reg 1 [24:31]     |     K7     |     K39     |     K71    |     K103    |  v[7]     ||    Reg 1 [24:31]     |     K7     |     K39     |  v[7]
+    // Reg 2 [0:7]       |     K8     |     K40     |     K72    |     K104    |  v[8]     ||    Reg 2 [0:7]       |     K8     |     K40     |  v[8]
+    // Reg 2 [8:15]      |     K9     |     K41     |     K73    |     K105    |  v[9]     ||    Reg 2 [8:15]      |     K9     |     K41     |  v[9]
+    // Reg 2 [16:23]     |     K10    |     K42     |     K74    |     K106    |  v[10]    ||    Reg 2 [16:23]     |     K10    |     K42     |  v[10]
+    // Reg 2 [24:31]     |     K11    |     K43     |     K75    |     K107    |  v[11]    ||    Reg 2 [24:31]     |     K11    |     K43     |  v[11]
+    // Reg 3 [0:7]       |     K12    |     K44     |     K76    |     K108    |  v[12]    ||    Reg 3 [0:7]       |     K12    |     K44     |  v[12]
+    // Reg 3 [8:15]      |     K13    |     K45     |     K77    |     K109    |  v[13]    ||    Reg 3 [8:15]      |     K13    |     K45     |  v[13]
+    // Reg 3 [16:23]     |     K14    |     K46     |     K78    |     K110    |  v[14]    ||    Reg 3 [16:23]     |     K14    |     K46     |  v[14]
+    // Reg 3 [24:31]     |     K15    |     K47     |     K79    |     K111    |  v[15]    ||    Reg 3 [24:31]     |     K15    |     K47     |  v[15]
+    // Reg 4 [0:7]       |     K16    |     K48     |     K80    |     K112    |  v[16]    ||    Reg 4 [0:7]       |     K16    |     K48     |  v[16]
+    // Reg 4 [8:15]      |     K17    |     K49     |     K81    |     K113    |  v[17]    ||    Reg 4 [8:15]      |     K17    |     K49     |  v[17]
+    // Reg 4 [16:23]     |     K18    |     K50     |     K82    |     K114    |  v[18]    ||    Reg 4 [16:23]     |     K18    |     K50     |  v[18]
+    // Reg 4 [24:31]     |     K19    |     K51     |     K83    |     K115    |  v[19]    ||    Reg 4 [24:31]     |     K19    |     K51     |  v[19]
+    // Reg 5 [0:7]       |     K20    |     K52     |     K84    |     K116    |  v[20]    ||    Reg 5 [0:7]       |     K20    |     K52     |  v[20]
+    // Reg 5 [8:15]      |     K21    |     K53     |     K85    |     K117    |  v[21]    ||    Reg 5 [8:15]      |     K21    |     K53     |  v[21]
+    // Reg 5 [16:23]     |     K22    |     K54     |     K86    |     K118    |  v[22]    ||    Reg 5 [16:23]     |     K22    |     K54     |  v[22]
+    // Reg 5 [24:31]     |     K23    |     K55     |     K87    |     K119    |  v[23]    ||    Reg 5 [24:31]     |     K23    |     K55     |  v[23]
+    // Reg 6 [0:7]       |     K24    |     K56     |     K88    |     K120    |  v[24]    ||    Reg 6 [0:7]       |     K24    |     K56     |  v[24]
+    // Reg 6 [8:15]      |     K25    |     K57     |     K89    |     K121    |  v[25]    ||    Reg 6 [8:15]      |     K25    |     K57     |  v[25]
+    // Reg 6 [16:23]     |     K26    |     K58     |     K90    |     K122    |  v[26]    ||    Reg 6 [16:23]     |     K26    |     K58     |  v[26]
+    // Reg 6 [24:31]     |     K27    |     K59     |     K91    |     K123    |  v[27]    ||    Reg 6 [24:31]     |     K27    |     K59     |  v[27]
+    // Reg 7 [0:7]       |     K28    |     K60     |     K92    |     K124    |  v[28]    ||    Reg 7 [0:7]       |     K28    |     K60     |  v[28]
+    // Reg 7 [8:15]      |     K29    |     K61     |     K93    |     K125    |  v[29]    ||    Reg 7 [8:15]      |     K29    |     K61     |  v[29]
+    // Reg 7 [16:23]     |     K30    |     K62     |     K94    |     K126    |  v[30]    ||    Reg 7 [16:23]     |     K30    |     K62     |  v[30]
+    // Reg 7 [24:31]     |     K31    |     K63     |     K95    |     K127    |  v[31]    ||    Reg 7 [24:31]     |     K31    |     K63     |  v[31]
+    // clang-format on
+    // Here we want to load a BLOCK_M x BLOCK_K block of data.
+    static constexpr uint32_t VW = vectorSize(AFragT{});
+    using ARawT                  = typename scalar_type<AFragT>::type;
+    using AScalarFragT           = vector_type<ARawT, VW>::type;
+    // To start the loading process, let's visualize in 2D coords.
+    // Each thread will load 32 elements.
+    // We need to know where they start, and where the next elements are.
+    auto startCoord2D = std::make_pair(threadIdx.x % BLOCK_M,         // Row
+                                       (threadIdx.x / BLOCK_M) * VW); // Col
+    auto stepCoord2D  = std::make_pair(0u, 1u);
+    // Flatten to 1D col_major offsets.
+    auto col_major = [](auto const& coord, auto ld) { return coord.first + coord.second * ld; };
+    // BLOCK_M is a stride in A matrix
+    auto startOffset = col_major(startCoord2D, BLOCK_M);
+    auto kOffset     = col_major(stepCoord2D, BLOCK_M);
+    // kOffset == BLOCK_M
+    // This means every BLOCK_M element is loaded into output vector
+    auto fragA = AScalarFragT{};
+#pragma unroll VW
+    for(uint32_t i = 0; i < VW; i++)
+    {
+        fragA[i] = bit_cast<ARawT>(input_ptr[startOffset + i * kOffset]);
+    }
+    return fragA;
+}
+// Define a load function for input B blocks:
+// Size: (BLOCK_K x BLOCK_N)
+// ASSUMPTION:
+// - We want contiguous BLOCK_N sized row neighbors in register.
+// - Data is in row_major format
+// This means:
+// - From B we will load K rows of size BLOCK_N to satisfy our input data
+template <typename BType, typename BFragT, int32_t BLOCK_K, int32_t BLOCK_N>
+__device__ BFragT load_B_col_major(BType const* input_ptr)
+{
+    // clang-format off
+    // Register Mapping for 128x16:                                                        ||    Register Mapping for 64x32:
+    // Size              |   BLOCK_N  |   BLOCK_N   |   BLOCK_N  |   BLOCK_N   |           ||    Size              |   BLOCK_N  |   BLOCK_N   |
+    // N                 | 0  ...  15 |  0  ...  15 | 0  ...  15 |  0  ...  15 |           ||    N                 | 0  ...  31 |  0  ...  31 |
+    // Thread Id         | 0  ...  15 | 16  ...  31 | 32  ... 47 | 48  ...  63 | Vector    ||    Thread Id         | 0  ...  31 | 32  ...  63 | Vector
+    // Register Element   ------------ ------------- ------------ -------------  Element   ||    Register Element   ------------ -------------  Element
+    // Reg 0 [0:7]       |     K0     |     K32     |     K64    |     K96     |  v[0]     ||    Reg 0 [0:7]       |     K0     |     K32     |  v[0]
+    // Reg 0 [8:15]      |     K1     |     K33     |     K65    |     K97     |  v[1]     ||    Reg 0 [8:15]      |     K1     |     K33     |  v[1]
+    // Reg 0 [16:23]     |     K2     |     K34     |     K66    |     K98     |  v[2]     ||    Reg 0 [16:23]     |     K2     |     K34     |  v[2]
+    // Reg 0 [24:31]     |     K3     |     K35     |     K67    |     K99     |  v[3]     ||    Reg 0 [24:31]     |     K3     |     K35     |  v[3]
+    // Reg 1 [0:7]       |     K4     |     K36     |     K68    |     K100    |  v[4]     ||    Reg 1 [0:7]       |     K4     |     K36     |  v[4]
+    // Reg 1 [8:15]      |     K5     |     K37     |     K69    |     K101    |  v[5]     ||    Reg 1 [8:15]      |     K5     |     K37     |  v[5]
+    // Reg 1 [16:23]     |     K6     |     K38     |     K70    |     K102    |  v[6]     ||    Reg 1 [16:23]     |     K6     |     K38     |  v[6]
+    // Reg 1 [24:31]     |     K7     |     K39     |     K71    |     K103    |  v[7]     ||    Reg 1 [24:31]     |     K7     |     K39     |  v[7]
+    // Reg 2 [0:7]       |     K8     |     K40     |     K72    |     K104    |  v[8]     ||    Reg 2 [0:7]       |     K8     |     K40     |  v[8]
+    // Reg 2 [8:15]      |     K9     |     K41     |     K73    |     K105    |  v[9]     ||    Reg 2 [8:15]      |     K9     |     K41     |  v[9]
+    // Reg 2 [16:23]     |     K10    |     K42     |     K74    |     K106    |  v[10]    ||    Reg 2 [16:23]     |     K10    |     K42     |  v[10]
+    // Reg 2 [24:31]     |     K11    |     K43     |     K75    |     K107    |  v[11]    ||    Reg 2 [24:31]     |     K11    |     K43     |  v[11]
+    // Reg 3 [0:7]       |     K12    |     K44     |     K76    |     K108    |  v[12]    ||    Reg 3 [0:7]       |     K12    |     K44     |  v[12]
+    // Reg 3 [8:15]      |     K13    |     K45     |     K77    |     K109    |  v[13]    ||    Reg 3 [8:15]      |     K13    |     K45     |  v[13]
+    // Reg 3 [16:23]     |     K14    |     K46     |     K78    |     K110    |  v[14]    ||    Reg 3 [16:23]     |     K14    |     K46     |  v[14]
+    // Reg 3 [24:31]     |     K15    |     K47     |     K79    |     K111    |  v[15]    ||    Reg 3 [24:31]     |     K15    |     K47     |  v[15]
+    // Reg 4 [0:7]       |     K16    |     K48     |     K80    |     K112    |  v[16]    ||    Reg 4 [0:7]       |     K16    |     K48     |  v[16]
+    // Reg 4 [8:15]      |     K17    |     K49     |     K81    |     K113    |  v[17]    ||    Reg 4 [8:15]      |     K17    |     K49     |  v[17]
+    // Reg 4 [16:23]     |     K18    |     K50     |     K82    |     K114    |  v[18]    ||    Reg 4 [16:23]     |     K18    |     K50     |  v[18]
+    // Reg 4 [24:31]     |     K19    |     K51     |     K83    |     K115    |  v[19]    ||    Reg 4 [24:31]     |     K19    |     K51     |  v[19]
+    // Reg 5 [0:7]       |     K20    |     K52     |     K84    |     K116    |  v[20]    ||    Reg 5 [0:7]       |     K20    |     K52     |  v[20]
+    // Reg 5 [8:15]      |     K21    |     K53     |     K85    |     K117    |  v[21]    ||    Reg 5 [8:15]      |     K21    |     K53     |  v[21]
+    // Reg 5 [16:23]     |     K22    |     K54     |     K86    |     K118    |  v[22]    ||    Reg 5 [16:23]     |     K22    |     K54     |  v[22]
+    // Reg 5 [24:31]     |     K23    |     K55     |     K87    |     K119    |  v[23]    ||    Reg 5 [24:31]     |     K23    |     K55     |  v[23]
+    // Reg 6 [0:7]       |     K24    |     K56     |     K88    |     K120    |  v[24]    ||    Reg 6 [0:7]       |     K24    |     K56     |  v[24]
+    // Reg 6 [8:15]      |     K25    |     K57     |     K89    |     K121    |  v[25]    ||    Reg 6 [8:15]      |     K25    |     K57     |  v[25]
+    // Reg 6 [16:23]     |     K26    |     K58     |     K90    |     K122    |  v[26]    ||    Reg 6 [16:23]     |     K26    |     K58     |  v[26]
+    // Reg 6 [24:31]     |     K27    |     K59     |     K91    |     K123    |  v[27]    ||    Reg 6 [24:31]     |     K27    |     K59     |  v[27]
+    // Reg 7 [0:7]       |     K28    |     K60     |     K92    |     K124    |  v[28]    ||    Reg 7 [0:7]       |     K28    |     K60     |  v[28]
+    // Reg 7 [8:15]      |     K29    |     K61     |     K93    |     K125    |  v[29]    ||    Reg 7 [8:15]      |     K29    |     K61     |  v[29]
+    // Reg 7 [16:23]     |     K30    |     K62     |     K94    |     K126    |  v[30]    ||    Reg 7 [16:23]     |     K30    |     K62     |  v[30]
+    // Reg 7 [24:31]     |     K31    |     K63     |     K95    |     K127    |  v[31]    ||    Reg 7 [24:31]     |     K31    |     K63     |  v[31]
+    // clang-format on
+    // Here we want to load a BLOCK_K x BLOCK_N block of data.
+    static constexpr uint32_t VW = vectorSize(BFragT{});
+    // To start the loading process, let's visualize in 2D coords.
+    // Each thread will load 32 elements.
+    // We need to know where they start, and where the next elements are.
+    auto startCoord2D = std::make_pair((threadIdx.x / BLOCK_N) * VW, // Row
+                                       threadIdx.x % BLOCK_N);       // Col
+    // Flatten to 1D col_major offsets.
+    auto col_major = [](auto const& coord, auto ld) { return coord.first + coord.second * ld; };
+    auto startOffset = col_major(startCoord2D, BLOCK_K);
+    auto const* fragPtr = reinterpret_cast<BFragT const*>(input_ptr + startOffset);
+    return *fragPtr;
+}
+// Define a store function for C
+// Size: (BLOCK_M x BLOCK_N)
+// ASSUMPTION:
+// - We want contiguous BLOCK_N sized row neighbors in register.
+// - Data is in col_major format
+// This means:
+// - From C we will load BLOCK_M rows of size BLOCK_N to satisfy our input data
+template <typename CType, typename CFragT, int32_t BLOCK_M, int32_t BLOCK_N>
+struct store_C_col_major;
+// Here we want to store a 16x16 block of data.
+//
+// Size              |   BLOCK_N  |   BLOCK_N   |   BLOCK_N   |   BLOCK_N   |
+// N                 | 0  ...  15 |  0  ...  15 | 0  ...  15  |  0  ...  15 |
+// Thread Id         | 0  ...  15 | 16  ...  31 | 32  ... 47  | 48  ...  63 | Vector
+// Register Element   ------------ ------------- ------------ -------------- Element
+// Reg0              |     M0     |     M4      |     M8      |     M12     | v[0]
+// Reg1              |     M1     |     M5      |     M9      |     M13     | v[1]
+// Reg2              |     M2     |     M6      |     M10     |     M14     | v[2]
+// Reg3              |     M3     |     M7      |     M11     |     M15     | v[3]
+template <typename CType, typename CFragT>
+struct store_C_col_major<CType, CFragT, 16, 16>
+{
+    __device__ void operator()(CType* output, CFragT cFrag)
+    {
+        static constexpr uint32_t VW  = vectorSize(cFrag); // 4
+        static constexpr uint32_t Dim = 16;
+        // Each thread will load 4 elements.
+        // We need to know where they start, and where the next elements are.
+        auto startCoord2D = std::make_pair((threadIdx.x / Dim) * VW, // Row
+                                           threadIdx.x % Dim);       // Col
+        // Flatten to 1D col_major offsets.
+        auto col_major = [](auto const& coord, auto ld) { return coord.first + coord.second * ld; };
+        auto startOffset = col_major(startCoord2D, 16);
+        auto* fragPtr = reinterpret_cast<CFragT*>(output + startOffset);
+        *fragPtr      = cFrag;
+    }
+};
+// Here we want to store a 32x32 block of data.
+// Register Mapping:
+// Size              |   BLOCK_N  |   BLOCK_N   |
+// N                 | 0  ...  31 |  0  ...  31 |
+// Thread Id         | 0  ...  31 | 32  ...  63 | Vector
+// Register Element   ------------ -------------  Element
+// Reg0              |     M0     |     M4      | v[0]
+// Reg1              |     M1     |     M5      | v[1]
+// Reg2              |     M2     |     M6      | v[2]
+// Reg3              |     M3     |     M7      | v[3]
+//                    ____________ _____________
+// Reg4              |     M8     |     M12     | v[4]
+// Reg5              |     M9     |     M13     | v[5]
+// Reg6              |     M10    |     M14     | v[6]
+// Reg7              |     M11    |     M15     | v[7]
+//                    ____________ _____________
+// Reg8              |     M16    |     M20     | v[8]
+// Reg9              |     M17    |     M21     | v[9]
+// Reg10             |     M18    |     M22     | v[10]
+// Reg11             |     M19    |     M23     | v[11]
+//                    ____________ _____________
+// Reg12             |     M24    |     M28     | v[12]
+// Reg13             |     M25    |     M29     | v[13]
+// Reg14             |     M26    |     M30     | v[14]
+// Reg15             |     M27    |     M31     | v[15]
+template <typename CType, typename CFragT>
+struct store_C_col_major<CType, CFragT, 32, 32>
+{
+    __device__ void operator()(CType* output, CFragT cFrag)
+    {
+        static constexpr uint32_t WAVE_SIZE      = 64;
+        static constexpr uint32_t VW             = 4;
+        static constexpr uint32_t Dim            = 32;
+        static constexpr uint32_t M_PER_VW_CHUNK = VW * WAVE_SIZE / 32; // 8
+        auto startCoord2D = std::make_pair((threadIdx.x / Dim) * VW, // Row
+                                           threadIdx.x % Dim);       // Col
+        // Major step between 'chunks'
+        auto majorStepCoord2D = std::make_pair(M_PER_VW_CHUNK, 0);
+        // Flatten to 1D col_major offsets.
+        auto col_major = [](auto const& coord, auto ld) { return coord.first + coord.second * ld; };
+        auto startOffset  = col_major(startCoord2D, 32);
+        auto kMajorOffset = col_major(majorStepCoord2D, 32); // 8
+        // we can vector store 4 contiguous elements at a time.
+        using CRawT        = typename scalar_type<CFragT>::type;
+        using CScalarFragT = vector_type<CRawT, VW>::type;
+        union
+        {
+            CFragT frag;
+            CScalarFragT chunks[vectorSize(CFragT{}) / VW];
+        } fragC{cFrag}; // Initialize with input fragment
+        *(reinterpret_cast<CScalarFragT*>(output + startOffset))                = fragC.chunks[0];
+        *(reinterpret_cast<CScalarFragT*>(output + startOffset + kMajorOffset)) = fragC.chunks[1];
+        *(reinterpret_cast<CScalarFragT*>(output + startOffset + 2 * kMajorOffset)) =
+            fragC.chunks[2];
+        *(reinterpret_cast<CScalarFragT*>(output + startOffset + 3 * kMajorOffset)) =
+            fragC.chunks[3];
+    }
+};
+template <typename AType,
+          typename BType,
+          typename CType,
+          typename AccType,
+          int32_t BLOCK_M,
+          int32_t BLOCK_N,
+          int32_t BLOCK_K>
+__global__ void matmul(const AType* a, const BType* b, CType* c)
+{
+    constexpr int WAVE_SIZE = 64;
+    assert(threadIdx.x < WAVE_SIZE);
+    assert(blockDim.x == 1 && blockDim.y == 1 && blockDim.z == 1);
+    using AFragT        = vector_type<AType, BLOCK_M * BLOCK_K / WAVE_SIZE>::type;
+    using BFragT        = vector_type<BType, BLOCK_K * BLOCK_N / WAVE_SIZE>::type;
+    using CFragT        = vector_type<CType, BLOCK_M * BLOCK_N / WAVE_SIZE>::type;
+    using AccumFragT    = vector_type<AccType, BLOCK_M * BLOCK_N / WAVE_SIZE>;
+    using RawAccumFragT = vector_type<AccType, BLOCK_M * BLOCK_N / WAVE_SIZE>::type;
+    // Create frags
+    auto fragA   = AFragT{};
+    auto fragB   = BFragT{};
+    auto fragC   = CFragT{};
+    auto fragAcc = AccumFragT{0};
+    // Load the inputs.
+    // A = col major, BLOCK_M x BLOCK_K
+    fragA = load_A_col_major<AType, AFragT, BLOCK_M, BLOCK_K>(a);
+    // B = col major, BLOCK_K x BLOCK_N
+    fragB = load_B_col_major<BType, BFragT, BLOCK_K, BLOCK_N>(b);
+    // Matrix multiply-accumulate using MFMA units
+    // Accumulation intermediate = BLOCK_M x BLOCK_N
+    mfma_type_selector<AFragT, BFragT, AccumFragT, BLOCK_M, BLOCK_N>{}(fragA, fragB, fragAcc);
+    for(int i = 0; i < vectorSize(fragC); ++i)
+    {
+        fragC[i] = type_convert<CType>(fragAcc.template AsType<RawAccumFragT>()[Number<0>{}][i]);
+    }
+    auto storeC = store_C_col_major<CType, CFragT, BLOCK_M, BLOCK_N>{};
+    storeC(c, fragC);
+}
+/**
+ * @brief Structure to hold dimension parameters for GEMM tensors.
+ *
+ * M Number of rows in matrix A and matrix C.
+ * N Number of columns in matrix B and matrix C.
+ * K Number of columns in matrix A and number of rows in matrix B.
+ * StrideA Stride (leading dimension) of matrix A.
+ * StrideB Stride (leading dimension) of matrix B.
+ * StrideC Stride (leading dimension) of matrix C.
+ */
+struct GemmParams
+{
+    ck::index_t M = 16;
+    ck::index_t N = 16;
+    ck::index_t K = 128;
+    ck::index_t StrideA = -1;
+    ck::index_t StrideB = -1;
+    ck::index_t StrideC = -1;
+};
+namespace mfma_test {
+template <typename GemmInstance,
+          typename ADataType,
+          typename BDataType,
+          typename CDataType,
+          typename AElementwiseOperation,
+          typename BElementwiseOperation,
+          typename CElementwiseOperation>
+void RunHostGEMM(const Tensor<ADataType>& A,
+                 const Tensor<BDataType>& B,
+                 Tensor<CDataType>& C,
+                 AElementwiseOperation a_element_op,
+                 BElementwiseOperation b_element_op,
+                 CElementwiseOperation c_element_op)
+{
+    auto ref_gemm     = GemmInstance{};
+    auto ref_invoker  = ref_gemm.MakeInvoker();
+    auto ref_argument = ref_gemm.MakeArgument(A, B, C, a_element_op, b_element_op, c_element_op);
+    ref_invoker.Run(ref_argument);
+}
+template <typename KernelType, typename ADataType, typename BDataType, typename CDataType>
+bool RunDeviceGEMM(KernelType kernel,
+                   const Tensor<ADataType>& A,
+                   const Tensor<BDataType>& B,
+                   Tensor<CDataType>& C)
+{
+    DeviceMem a_m_k_device_buf(sizeof(ADataType) * A.mDesc.GetElementSpaceSize());
+    DeviceMem b_n_k_device_buf(sizeof(BDataType) * B.mDesc.GetElementSpaceSize());
+    DeviceMem c_m_n_device_buf(sizeof(CDataType) * C.mDesc.GetElementSpaceSize());
+    a_m_k_device_buf.ToDevice(A.mData.data());
+    b_n_k_device_buf.ToDevice(B.mData.data());
+    kernel<<<1, 64>>>(static_cast<const ADataType*>(a_m_k_device_buf.GetDeviceBuffer()),
+                      static_cast<const BDataType*>(b_n_k_device_buf.GetDeviceBuffer()),
+                      static_cast<CDataType*>(c_m_n_device_buf.GetDeviceBuffer()));
+    c_m_n_device_buf.FromDevice(C.mData.data());
+    return true;
+}
+template <typename DeviceMFMA,
+          typename ADataType,
+          typename BDataType,
+          typename CDataType,
+          typename GPUAccDataType,
+          typename CPUAccDataType,
+          typename ALayout,
+          typename BLayout,
+          typename CLayout,
+          index_t BLOCK_M,
+          index_t BLOCK_N,
+          index_t BLOCK_K>
+struct TestMFMA
+{
+    auto PrepareGemmTensors(const GemmParams& params, index_t init)
+    {
+        auto f_host_tensor_descriptor =
+            [](std::size_t row, std::size_t col, std::size_t stride, auto layout) {
+                if(std::is_same<decltype(layout), ck::tensor_layout::gemm::RowMajor>::value)
+                {
+                    return HostTensorDescriptor(std::vector<std::size_t>({row, col}),
+                                                std::vector<std::size_t>({stride, 1}));
+                }
+                else
+                {
+                    return HostTensorDescriptor(std::vector<std::size_t>({row, col}),
+                                                std::vector<std::size_t>({1, stride}));
+                }
+            };
+        Tensor<ADataType> a_m_k(
+            f_host_tensor_descriptor(params.M, params.K, params.StrideA, ALayout{}));
+        Tensor<BDataType> b_n_k(
+            f_host_tensor_descriptor(params.K, params.N, params.StrideB, BLayout{}));
+        Tensor<CDataType> c_m_n_host_result(
+            f_host_tensor_descriptor(params.M, params.N, params.StrideC, CLayout{}));
+        Tensor<CDataType> c_m_n_device_result(
+            f_host_tensor_descriptor(params.M, params.N, params.StrideC, CLayout{}));
+        switch(init)
+        {
+        case 0:
+            a_m_k.GenerateTensorValue(GeneratorTensor_1<ADataType>{0.015625f});
+            // NOTE: not all numbers are representable in FP8, BF8, etc.
+            b_n_k.GenerateTensorValue(GeneratorTensor_Sequential<BDataType, 1>{});
+            break;
+        case 1:
+            // results in C = {K}
+            a_m_k.GenerateTensorValue(GeneratorTensor_1<ADataType>{1.0f});
+            b_n_k.GenerateTensorValue(GeneratorTensor_1<BDataType>{1.0f});
+            break;
+        case 2:
+            // expect small round off errors
+            a_m_k.GenerateTensorValue(GeneratorTensor_3<ADataType>{-5, 5});
+            b_n_k.GenerateTensorValue(GeneratorTensor_3<BDataType>{-5, 5});
+            break;
+        case 3:
+            // expect small round off errors
+            a_m_k.GenerateTensorValue(GeneratorTensor_4<ADataType>(-1, 3));
+            b_n_k.GenerateTensorValue(GeneratorTensor_4<BDataType>(1, 3));
+            break;
+        default:
+            // all initial values are representable in FP8, BF8
+            a_m_k.GenerateTensorValue(GeneratorTensor_2<ADataType>{-5, 6});
+            b_n_k.GenerateTensorValue(GeneratorTensor_2<BDataType>{-5, 6});
+            break;
+        }
+        return std::make_tuple(a_m_k, b_n_k, c_m_n_host_result, c_m_n_device_result);
+    }
+    auto operator()(const DeviceMFMA& mfma_kernel, index_t init)
+    {
+        std::cout << "ALayout = " << ALayout{}.name << ", BLayout = " << BLayout{}.name
+                  << ", CLayout = " << CLayout{}.name << std::endl;
+        // Arrange
+        GemmParams params;
+        params.M = BLOCK_M;
+        params.N = BLOCK_N;
+        params.K = BLOCK_K;
+        auto f_get_default_stride = [](std::size_t row,
+                                       std::size_t col,
+                                       ck::index_t stride,
+                                       auto layout) {
+            if(stride == -1)
+            {
+                // give a chance if stride is -1, return a default packed stride
+                if constexpr(std::is_same_v<decltype(layout), ck::tensor_layout::gemm::RowMajor>)
+                {
+                    return static_cast<std::size_t>(col);
+                }
+                else
+                {
+                    return static_cast<std::size_t>(row);
+                }
+            }
+            else
+                return static_cast<std::size_t>(stride);
+        };
+        params.StrideA = f_get_default_stride(BLOCK_M, BLOCK_K, params.StrideA, ALayout{});
+        params.StrideB = f_get_default_stride(BLOCK_K, BLOCK_N, params.StrideB, BLayout{});
+        params.StrideC = f_get_default_stride(BLOCK_M, BLOCK_N, params.StrideC, CLayout{});
+        auto host_tensors = PrepareGemmTensors(params, init);
+        const Tensor<ADataType>& a  = std::get<0>(host_tensors);
+        const Tensor<BDataType>& b  = std::get<1>(host_tensors);
+        Tensor<CDataType>& c_host   = std::get<2>(host_tensors);
+        Tensor<CDataType>& c_device = std::get<3>(host_tensors);
+        using PassThrough = ck::tensor_operation::element_wise::PassThrough;
+        auto a_element_op = PassThrough{};
+        auto b_element_op = PassThrough{};
+        auto c_element_op = PassThrough{};
+        using ReferenceGemmInstance = ck::tensor_operation::host::ReferenceGemm<ADataType,
+                                                                                BDataType,
+                                                                                CDataType,
+                                                                                CPUAccDataType,
+                                                                                PassThrough,
+                                                                                PassThrough,
+                                                                                PassThrough>;
+        RunHostGEMM<ReferenceGemmInstance>(a, b, c_host, a_element_op, b_element_op, c_element_op);
+        RunDeviceGEMM(mfma_kernel, a, b, c_device);
+        bool res = false;
+        if constexpr(std::is_same<CDataType, float>::value ||
+                     std::is_same<CDataType, half_t>::value)
+        {
+            res = ck::utils::check_err(c_device.mData, c_host.mData);
+            std::cout << (res ? "SUCCESS" : "FAILURE") << std::endl;
+        }
+        else
+        {
+            std::cout << "UNSUPPORTED CDataType" << std::endl;
+        }
+        return res;
+    }
+};
+} // namespace mfma_test
+} // namespace ck
--- a/test/smfmac_op/smfmac_op_xdl.cpp
+++ b/test/smfmac_op/smfmac_op_xdl.cpp
@@ -40,7 +40,7 @@ class TestSmfmac : public ::testing::Test
    void Run()
    {
        bool pass = true;
-        if(ck::get_device_name() == "gfx942")
+        if(ck::get_device_name() == "gfx942" || ck::get_device_name() == "gfx950")
        {
            constexpr auto matmul_default = ck::smfmac_op_util::matmul<Src1Type,
                                                                       Src1VecSize,