1. 17 Nov, 2021 1 commit
  2. 19 Oct, 2021 1 commit
  3. 08 Oct, 2021 1 commit
  4. 07 Oct, 2021 1 commit
  5. 04 Oct, 2021 1 commit
  6. 02 Oct, 2021 1 commit
  7. 24 Sep, 2021 1 commit
  8. 04 Sep, 2021 1 commit
    • Burc Eryilmaz's avatar
      fix CUBLAS guards (#1162) · 54b93919
      Burc Eryilmaz authored
      
      
      * support for fused dense layer with cublasLt, fusion in both fprop and bprop
      
      * fix typo causing syntax error
      
      * add fused GEMM+gelu+GEMM modue
      
      * fix typo for workspace size
      
      * update cublas check for 11600
      
      * add tests for fused dense layer
      
      * fix CUDA 10.x path
      
      * safer guard around CUBLAS constants, remove unreferenced variable
      
      * more guard changes
      
      * guard against cublas version instead of cuda
      Co-authored-by: default avatarSukru Eryilmaz <seryilmaz@computelab-dgx1v-32.nvidia.com>
      54b93919
  9. 01 Sep, 2021 2 commits
  10. 17 May, 2021 1 commit
  11. 19 Apr, 2021 1 commit
  12. 17 Apr, 2021 1 commit
  13. 15 Apr, 2021 1 commit
    • Sudhakar Singh's avatar
      Add unit tests for Fused NovoGrad (#1065) · 59d2f7ac
      Sudhakar Singh authored
      * Add unit tests for fused-novograd
      
      * Fix: tensors should reside on the same device
      
      * Fix: Cudastream should be called on the same device on which the tensors reside on. Found this during debugging fused novograd multi-device unit test
      
      * fixed issues mentioned in the comments
      59d2f7ac
  14. 25 Feb, 2021 1 commit
  15. 25 Jan, 2021 1 commit
    • Jeff Daily's avatar
      fix bugs in syncbn (#46) · 3f49dbf0
      Jeff Daily authored
      - incorrect use of __shfl_down
      - fix warp size assumptions
      - update unit tests to exit on failure
      3f49dbf0
  16. 21 Jan, 2021 1 commit
  17. 18 Jan, 2021 1 commit
  18. 15 Jan, 2021 1 commit
  19. 04 Nov, 2020 1 commit
  20. 19 Oct, 2020 1 commit
    • lly-zero-one's avatar
      Optimize the sync batchnorm by batching the communication (#980) · 8a1ed9e8
      lly-zero-one authored
      In this PR, we mainly tried to optimize the performance of Syncatchnorm and also fixed one potential issue in the welford_parallel kernel implementation.
      
      For performance improvement, we batched the mean/var/count all_gather communication together and sent it once in the forward path
      We also batch the all_reduce in backward path
      We add the contiguous call on the input of welford_parallel kernel.
      If there is any standard perf benchmark, I would be happy to run it.
      8a1ed9e8
  21. 05 Aug, 2020 2 commits
  22. 10 Jul, 2020 1 commit
  23. 06 Jul, 2020 1 commit
    • jjsjann123's avatar
      [sync BN] (#792) · 1ff54b8f
      jjsjann123 authored
      * [sync BN]
      
      support non-uniform batch size across process group.
      
      TODO: test should be added once cleaned up.
      
      * updating unit tests
      
      * new unit tests for different inputs
      
      * cleaning
      1ff54b8f
  24. 22 Jun, 2020 1 commit
  25. 15 Jun, 2020 1 commit
  26. 26 May, 2020 1 commit
  27. 23 May, 2020 1 commit
  28. 22 May, 2020 5 commits
  29. 21 May, 2020 2 commits
  30. 20 May, 2020 1 commit
  31. 14 May, 2020 1 commit
  32. 12 May, 2020 2 commits