"git@developer.sourcefind.cn:dadigang/Ventoy.git" did not exist on "9c3e1a688066371c895d1e15e5d654a5f2a4d7af"
  • Chao Liu's avatar
    Added bwd data v3r1 v4r1, tweaking v1 (#10) · c5da0377
    Chao Liu authored
    * Added bwd data v3r1: breaking down compute into a series of load balanced GEMM, and launch in a single kernel
    * Added bwd data v4r1: like v3r1, but launch GEMMs in multiple kernels
    * Tweaked v1r1  and v1r2 (atomic) on AMD GPU
    c5da0377
extract_asm-cuda.sh 125 Bytes