1. 06 Oct, 2022 2 commits
  2. 04 Oct, 2022 2 commits
  3. 03 Oct, 2022 2 commits
    • Umang Yadav's avatar
      Add output_alias and runs_on_offload_target flags for the custom ops (#1309) · c9ffb38d
      Umang Yadav authored
      Adds two methods for the custom_ops virtual class.
      
      bool runs_on_offload_target(), if the custom op runs directly on the gpu then it should be set to true. in this case, custom op expects its parameters to reside in GPU memory and writes output to the GPU memory. If it is set to false then, custom op expects it's parameter to reside on the host and puts back the result into the host memory.
      
      output_alias, if output of the custom op is aliasing the input buffer. i.e. interpreting the same input buffer with differnet shape and strides.
      
      Update as_vector() in C++ API to handle non-standard shapes. It required exposing element_index to space_index conversion method for the shape class.
      c9ffb38d
    • charlie's avatar
      Update comments · ed2acdc4
      charlie authored
      ed2acdc4
  4. 29 Sep, 2022 3 commits
  5. 28 Sep, 2022 1 commit
    • Umang Yadav's avatar
      Add compute_fp32 flag for quant_gemm tests (#1360) · 70e63960
      Umang Yadav authored
      test_gpu_pack_int8_args fails on gfx908 machine, because it doesn't set compute_fp32 flag correctly. This PR fixes the test such that it checks for the device-name, and rocblas-versions and sets this flag accordingly.
      70e63960
  6. 27 Sep, 2022 1 commit
  7. 26 Sep, 2022 3 commits
  8. 24 Sep, 2022 2 commits
  9. 23 Sep, 2022 1 commit
  10. 21 Sep, 2022 2 commits
  11. 19 Sep, 2022 4 commits
  12. 16 Sep, 2022 7 commits
  13. 15 Sep, 2022 2 commits
  14. 14 Sep, 2022 4 commits
  15. 13 Sep, 2022 1 commit
    • turneram's avatar
      Use rocblas_gemm_ex for batched gemms with broadcasted B (#1354) · a10a8ef1
      turneram authored
      Improves performance for 4/6 GEMMs used by huggingface BERT models with batch_size>1 by using a non-batched rocBLAS call for GEMMs where the B input has a broadcasted batch dimension.
      The four verify tests added reflect the actual configurations used by bert-base-cased, with varied batch sizes.
      
      Also adds a matcher to simplify_reshapes to move multibroadcasts after concats.
      a10a8ef1
  16. 09 Sep, 2022 1 commit
  17. 08 Sep, 2022 2 commits