1. 15 Feb, 2022 3 commits
  2. 09 Feb, 2022 5 commits
  3. 08 Feb, 2022 5 commits
  4. 04 Feb, 2022 2 commits
  5. 31 Jan, 2022 1 commit
  6. 28 Jan, 2022 2 commits
  7. 27 Jan, 2022 1 commit
  8. 21 Jan, 2022 1 commit
  9. 10 Jan, 2022 1 commit
  10. 09 Dec, 2021 1 commit
    • Shucai Xiao's avatar
      Softmax perf optimization (#1014) · 2e337c7f
      Shucai Xiao authored
      Changed the number of threads in a block from 256 to 128
      Increased the max number of blocks in the kernel from 256 to 1M.
      For the case that the axis is the last dimension, we removed the computation of index since it is not required.
      
      With these change, we can get about 2x speedup compared to the develop branch for the softmax op used in the BertSquad model.
      2e337c7f
  11. 08 Dec, 2021 1 commit
  12. 07 Dec, 2021 1 commit
  13. 02 Dec, 2021 1 commit
  14. 30 Nov, 2021 2 commits
  15. 24 Nov, 2021 1 commit
  16. 18 Nov, 2021 1 commit
  17. 11 Nov, 2021 1 commit
    • Paul Fultz II's avatar
      Conditionally enable pointwise fusion (#992) · 157935ff
      Paul Fultz II authored
      This enables the pointwise fusions using the MIGRAPHX_ENABLE_POINTWISE_FUSION env variable. Its disabled by default since MIOpen fusions need to be refactored.
      
      This also adds a compile_ops pass to compile the pointwise modules. All tests except test_gpu_fast_math passes with MIGRAPHX_ENABLE_POINTWISE_FUSION=1 set.
      157935ff
  18. 09 Nov, 2021 1 commit
  19. 05 Nov, 2021 1 commit
  20. 28 Oct, 2021 2 commits
  21. 20 Oct, 2021 1 commit
    • Shucai Xiao's avatar
      Roialign (#952) · d7653732
      Shucai Xiao authored
      Implementation of the roialign operator. For now, we have only the ref implementation. When we run a model on the GPU, we fall back the execution to use the ref implementation.
      d7653732
  22. 19 Oct, 2021 1 commit
  23. 08 Oct, 2021 2 commits
  24. 01 Oct, 2021 1 commit
    • turneram's avatar
      Add multinomial op (#954) · 0b7672d7
      turneram authored
      
      
      Add multinomial op to onnx parser with ref and GPU implementations.
      
      The onnx parser inserts a literal of shape {batch_size, sample_size} with random values in the range [0, 1) and inserts existing ops to compute the cumulative density function. The multinomial operator multiplies the random values by the sum of the CDF and returns the index of the first element of the CDF that is greater than the result, representing samples randomly drawn from [0, class_size) that follow the log-probability distribution.
      
      Resolves #821
      Co-authored-by: default avatarShucai Xiao <shucai@gmail.com>
      0b7672d7
  25. 27 Sep, 2021 1 commit