1. 05 Feb, 2022 1 commit
  2. 28 Jan, 2022 1 commit
  3. 27 Jan, 2022 1 commit
  4. 26 Jan, 2022 1 commit
    • Paul's avatar
      Updates · 1cc6c88c
      Paul authored
      1cc6c88c
  5. 21 Jan, 2022 1 commit
  6. 10 Jan, 2022 3 commits
  7. 07 Jan, 2022 2 commits
  8. 06 Jan, 2022 3 commits
  9. 11 Dec, 2021 5 commits
  10. 09 Dec, 2021 1 commit
    • Shucai Xiao's avatar
      Softmax perf optimization (#1014) · 2e337c7f
      Shucai Xiao authored
      Changed the number of threads in a block from 256 to 128
      Increased the max number of blocks in the kernel from 256 to 1M.
      For the case that the axis is the last dimension, we removed the computation of index since it is not required.
      
      With these change, we can get about 2x speedup compared to the develop branch for the softmax op used in the BertSquad model.
      2e337c7f
  11. 08 Dec, 2021 1 commit
  12. 07 Dec, 2021 1 commit
  13. 02 Dec, 2021 1 commit
  14. 01 Dec, 2021 4 commits
  15. 30 Nov, 2021 2 commits
  16. 24 Nov, 2021 3 commits
  17. 18 Nov, 2021 1 commit
  18. 16 Nov, 2021 4 commits
  19. 11 Nov, 2021 1 commit
    • Paul Fultz II's avatar
      Conditionally enable pointwise fusion (#992) · 157935ff
      Paul Fultz II authored
      This enables the pointwise fusions using the MIGRAPHX_ENABLE_POINTWISE_FUSION env variable. Its disabled by default since MIOpen fusions need to be refactored.
      
      This also adds a compile_ops pass to compile the pointwise modules. All tests except test_gpu_fast_math passes with MIGRAPHX_ENABLE_POINTWISE_FUSION=1 set.
      157935ff
  20. 09 Nov, 2021 2 commits
  21. 05 Nov, 2021 1 commit