1. 12 Jul, 2019 1 commit
  2. 03 Jul, 2019 2 commits
  3. 28 Jun, 2019 1 commit
  4. 14 Jun, 2019 1 commit
  5. 11 Jun, 2019 1 commit
  6. 31 May, 2019 2 commits
  7. 26 Apr, 2019 1 commit
    • ptrblck's avatar
      Replace type().ScalarType() with scalar_type() (#272) · 855808f3
      ptrblck authored
      * change .type().ScalarType() to .scalar_type() + at::ScalarType::X to at::kX
      
      * revert scalar_type() to type() for AT_DISPATCH_FLOATING_TYPES_AND_HALF
      
      * revert scalar_type() to type() in AT_DISPATCH_FLOATING_TYPES
      
      * revert scalar_type() to type() for AT_DISPATCH_FLOATING_TYPES_AND_HALF in welford.cu
      
      * revert scalar_type() to type() in layer_norm_cuda_kernel.cu
      
      * revert at::kType  to at::ScalarType::Type
      
      * use DISPATCH_FLOAT_AND_HALF to get rid of warnings
      
      * add dispatch mechanisms for double+float and double+float+half
      855808f3
  8. 10 Apr, 2019 2 commits
  9. 09 Apr, 2019 1 commit
  10. 08 Apr, 2019 1 commit
  11. 04 Apr, 2019 1 commit
    • mcarilli's avatar
      WIP: Handle arbitrary combinations of optimizers/models/losses (#232) · 3f87614f
      mcarilli authored
      * Refactor to allow more flexible treatment of multiple optimizers/models/losses
      
      * Adding _process_optimizers.py
      
      * Created L0 tests (now passing).
      
      * fix: minor print typo (#234)
      
      * make L1 results easier to read
      
      * L0 multiple model/optimizer/loss test fleshed out
      
      * Adding test that master params remain synced across distributed processes
      
      * Docstring updates
      
      * Docstring updates
      3f87614f
  12. 21 Mar, 2019 2 commits
  13. 19 Mar, 2019 2 commits
  14. 15 Mar, 2019 1 commit
  15. 12 Mar, 2019 1 commit
  16. 10 Mar, 2019 2 commits
  17. 03 Mar, 2019 1 commit
  18. 28 Feb, 2019 1 commit
  19. 24 Feb, 2019 1 commit
  20. 22 Feb, 2019 1 commit
  21. 19 Feb, 2019 1 commit
  22. 13 Feb, 2019 1 commit
  23. 11 Feb, 2019 1 commit
  24. 08 Feb, 2019 1 commit
  25. 06 Feb, 2019 2 commits
  26. 05 Feb, 2019 1 commit
  27. 04 Feb, 2019 1 commit
  28. 01 Feb, 2019 1 commit
  29. 18 Jan, 2019 1 commit
  30. 15 Jan, 2019 1 commit
    • Jie's avatar
      [sync BN nhwc] · 443fa76e
      Jie authored
      Added kernel to support sync BN for channel last tensor
      443fa76e
  31. 06 Nov, 2018 1 commit
    • Jie's avatar
      [syncBN] · ee67e56a
      Jie authored
      adjusted kernel config for better perf.
      removed divergence in welford warp reduction.
      ee67e56a
  32. 30 Oct, 2018 1 commit
  33. 29 Oct, 2018 1 commit
    • mcarilli's avatar
      Merging in fused adam optimizer, additional DDP features tested in 18.10 (#60) · e0bc5d62
      mcarilli authored
      * test passes
      
      * notes
      
      * Using C++-side flatten and unflatten functions
      
      * Adding csrc
      
      * Persistent synchronization event so it doesn't need to be created and destroyed each time
      
      * Interop with parameter flattening in SSD
      
      * Added deterministic option to imagenet main.py
      
      * Adding options to split gradient averaging and allreduce in pure fp32
      
      * Fixing allreduce_maybe_retain call
      
      * Fixing allreduce_fallback
      
      * Also sync active_i_buckets from rank 0
      
      * Making retain_allreduce_buffers compatible with/orthogonal to delay_allreduce=True|False
      
      * Correcting syntax error, now all seems to work with SSD
      
      * Optional cpp extension build
      
      * Add mixed precision adam optimizer (#59)
      
      * Add FusedAdam Optimizer to Apex that places all the math into a cuda kernel.
      
      * Added fixes to fused_adam to get it to work with network.
      
      * wip work on python interface for adam with options
      
      * fix dispatch for halfs, add python options to handle optional half gradients and params
      
      * cleanup, get rid of grid-stride loop
      e0bc5d62