"src/vscode:/vscode.git/clone" did not exist on "0905762b14c846b4b910d6aa37755d75b9897a9b"
  1. 10 Mar, 2022 3 commits
  2. 08 Mar, 2022 5 commits
  3. 07 Mar, 2022 2 commits
  4. 04 Mar, 2022 5 commits
  5. 03 Mar, 2022 2 commits
  6. 02 Mar, 2022 1 commit
  7. 01 Mar, 2022 4 commits
  8. 28 Feb, 2022 8 commits
  9. 26 Feb, 2022 2 commits
  10. 08 Feb, 2022 3 commits
  11. 31 Jan, 2022 1 commit
  12. 09 Dec, 2021 1 commit
    • Shucai Xiao's avatar
      Softmax perf optimization (#1014) · 2e337c7f
      Shucai Xiao authored
      Changed the number of threads in a block from 256 to 128
      Increased the max number of blocks in the kernel from 256 to 1M.
      For the case that the axis is the last dimension, we removed the computation of index since it is not required.
      
      With these change, we can get about 2x speedup compared to the develop branch for the softmax op used in the BertSquad model.
      2e337c7f
  13. 08 Oct, 2021 1 commit
  14. 01 Oct, 2021 1 commit
    • turneram's avatar
      Add multinomial op (#954) · 0b7672d7
      turneram authored
      
      
      Add multinomial op to onnx parser with ref and GPU implementations.
      
      The onnx parser inserts a literal of shape {batch_size, sample_size} with random values in the range [0, 1) and inserts existing ops to compute the cumulative density function. The multinomial operator multiplies the random values by the sum of the CDF and returns the index of the first element of the CDF that is greater than the result, representing samples randomly drawn from [0, class_size) that follow the log-probability distribution.
      
      Resolves #821
      Co-authored-by: default avatarShucai Xiao <shucai@gmail.com>
      0b7672d7
  15. 27 Sep, 2021 1 commit