1. 25 Feb, 2021 1 commit
  2. 25 Jan, 2021 1 commit
    • Jeff Daily's avatar
      fix bugs in syncbn (#46) · 3f49dbf0
      Jeff Daily authored
      - incorrect use of __shfl_down
      - fix warp size assumptions
      - update unit tests to exit on failure
      3f49dbf0
  3. 21 Jan, 2021 1 commit
  4. 18 Jan, 2021 1 commit
  5. 15 Jan, 2021 1 commit
  6. 04 Nov, 2020 1 commit
  7. 19 Oct, 2020 1 commit
    • lly-zero-one's avatar
      Optimize the sync batchnorm by batching the communication (#980) · 8a1ed9e8
      lly-zero-one authored
      In this PR, we mainly tried to optimize the performance of Syncatchnorm and also fixed one potential issue in the welford_parallel kernel implementation.
      
      For performance improvement, we batched the mean/var/count all_gather communication together and sent it once in the forward path
      We also batch the all_reduce in backward path
      We add the contiguous call on the input of welford_parallel kernel.
      If there is any standard perf benchmark, I would be happy to run it.
      8a1ed9e8
  8. 05 Aug, 2020 2 commits
  9. 10 Jul, 2020 1 commit
  10. 06 Jul, 2020 1 commit
    • jjsjann123's avatar
      [sync BN] (#792) · 1ff54b8f
      jjsjann123 authored
      * [sync BN]
      
      support non-uniform batch size across process group.
      
      TODO: test should be added once cleaned up.
      
      * updating unit tests
      
      * new unit tests for different inputs
      
      * cleaning
      1ff54b8f
  11. 22 Jun, 2020 1 commit
  12. 15 Jun, 2020 1 commit
  13. 26 May, 2020 1 commit
  14. 23 May, 2020 1 commit
  15. 22 May, 2020 5 commits
  16. 21 May, 2020 2 commits
  17. 20 May, 2020 1 commit
  18. 14 May, 2020 1 commit
  19. 12 May, 2020 2 commits
  20. 07 May, 2020 2 commits
    • Chaitanya Sri Krishna Lolla's avatar
      2d0f9cf2
    • Chaitanya Sri Krishna Lolla's avatar
      [Upstream] IFU 05072020 (#4) · e85a1d4b
      Chaitanya Sri Krishna Lolla authored
      
      
      * fix dropout scaling from p to 1/(1-p) (#816)
      Co-authored-by: default avatarSukru Eryilmaz <seryilmaz@computelab-dgx1v-32.nvidia.com>
      
      * Improvements to apex.mlp (#804)
      
      * update fused bias relu backward kernel
      
      * adding support for not require first layer dgrad
      
      * fix bug: wrong layer in requires grad
      
      * add infrastructure for optional bias and activation, currently only support no bias and no relu
      
      * make bias and relu optional separately
      
      * add sigmoid activation option
      
      * enable wider load/store for multi_tensor_apply kernels (#763)
      
      * modify MTA axpby for wider load/store
      
      * Make scale/axpby/l2/adam/lamb multi_tensor uses wider load
      
      * Changes to make xentropysoftmax load/store vectorized when possible: (#725)
      
      * Changes to make xentropysoftmax load/store vectorized when possible:
      Increase default ILP so that each thread handle 16 Bytes data in one step
      Make thread load/store longest vector possible
      Make unroll case handle adjacent data instead of strided...
      e85a1d4b
  21. 30 Apr, 2020 3 commits
  22. 28 Apr, 2020 2 commits
  23. 22 Apr, 2020 1 commit
  24. 10 Apr, 2020 1 commit
  25. 27 Feb, 2020 1 commit
  26. 04 Oct, 2019 1 commit
  27. 06 Sep, 2019 1 commit
    • mcarilli's avatar
      Fix for #456 (#477) · 325f5a0b
      mcarilli authored
      * Pushing for build tests
      
      * Contrib files
      
      * Removing deprecated checks
      325f5a0b
  28. 20 Aug, 2019 1 commit
  29. 17 Aug, 2019 1 commit