"README_origin.md" did not exist on "79595cd16302921eab2e95452aac95ef55993505"
  • turneram's avatar
    Use rocblas_gemm_ex for batched gemms with broadcasted B (#1354) · a10a8ef1
    turneram authored
    Improves performance for 4/6 GEMMs used by huggingface BERT models with batch_size>1 by using a non-batched rocBLAS call for GEMMs where the B input has a broadcasted batch dimension.
    The four verify tests added reflect the actual configurations used by bert-base-cased, with varied batch sizes.
    
    Also adds a matcher to simplify_reshapes to move multibroadcasts after concats.
    a10a8ef1
test_unbatched_gemm_1.cpp 2.76 KB