1. 17 Nov, 2022 1 commit
    • Charlie Lin's avatar
      Dynamic ref contiguous (#1445) · 95d82a51
      Charlie Lin authored
      Extends the ref contiguous operator to handle dynamic shapes
      Updates the eliminate_contiguous pass to use the dyn_output struct
      95d82a51
  2. 15 Nov, 2022 1 commit
  3. 14 Nov, 2022 1 commit
  4. 13 Nov, 2022 1 commit
    • Charlie Lin's avatar
      Dyn ref multibroadcast; dyn binary (#1423) · d73c6d7c
      Charlie Lin authored
      Updated Multibroadcast op to have a two input version for dynamic shapes
      Current dynamic shape broadcasting logic
      dynamic_dimensions must be the same or one of them is {1, 1, 0} or {1, 1, 1}
      Works for dyn-dyn, dyn-static, and static-static shape combinations
      Changed common.cpp for multibroadcasting for binary ops with dynamic shapes
      Extended binary.hpp for dynamic shapes to test the new common.cpp stuff
      d73c6d7c
  5. 07 Nov, 2022 2 commits
  6. 06 Nov, 2022 1 commit
  7. 02 Nov, 2022 3 commits
  8. 01 Nov, 2022 2 commits
  9. 31 Oct, 2022 1 commit
  10. 28 Oct, 2022 1 commit
  11. 27 Oct, 2022 2 commits
  12. 26 Oct, 2022 2 commits
  13. 25 Oct, 2022 1 commit
  14. 24 Oct, 2022 1 commit
  15. 21 Oct, 2022 1 commit
  16. 19 Oct, 2022 2 commits
  17. 18 Oct, 2022 1 commit
  18. 17 Oct, 2022 1 commit
    • Umang Yadav's avatar
      memset fix (#1414) · 83784c52
      Umang Yadav authored
      hipMemset is causing random failure.
      hipMemsetAsync is doing the correct synchronization.
      83784c52
  19. 14 Oct, 2022 1 commit
  20. 13 Oct, 2022 2 commits
  21. 07 Oct, 2022 1 commit
  22. 04 Oct, 2022 2 commits
  23. 03 Oct, 2022 1 commit
    • Umang Yadav's avatar
      Add output_alias and runs_on_offload_target flags for the custom ops (#1309) · c9ffb38d
      Umang Yadav authored
      Adds two methods for the custom_ops virtual class.
      
      bool runs_on_offload_target(), if the custom op runs directly on the gpu then it should be set to true. in this case, custom op expects its parameters to reside in GPU memory and writes output to the GPU memory. If it is set to false then, custom op expects it's parameter to reside on the host and puts back the result into the host memory.
      
      output_alias, if output of the custom op is aliasing the input buffer. i.e. interpreting the same input buffer with differnet shape and strides.
      
      Update as_vector() in C++ API to handle non-standard shapes. It required exposing element_index to space_index conversion method for the shape class.
      c9ffb38d
  24. 29 Sep, 2022 2 commits
  25. 28 Sep, 2022 1 commit
    • Umang Yadav's avatar
      Add compute_fp32 flag for quant_gemm tests (#1360) · 70e63960
      Umang Yadav authored
      test_gpu_pack_int8_args fails on gfx908 machine, because it doesn't set compute_fp32 flag correctly. This PR fixes the test such that it checks for the device-name, and rocblas-versions and sets this flag accordingly.
      70e63960
  26. 27 Sep, 2022 1 commit
  27. 26 Sep, 2022 3 commits
  28. 24 Sep, 2022 1 commit
    • Chris Austen's avatar
      check concurrency on PR level with one running and one pending performance tests (#1401) · 94bc41dc
      Chris Austen authored
      Workflow has concurrency reintroduced with different set of rules. New expected behavior is to check concurrency on PR level with one running and one pending performance tests. In case of multiple commits in same PR, always the latest commit is queued after initiated performance test execution is completed. Any other PRs/commits are in pending/queued state
      94bc41dc