1. 15 May, 2020 2 commits
  2. 14 May, 2020 1 commit
  3. 13 May, 2020 1 commit
  4. 12 May, 2020 3 commits
  5. 08 May, 2020 1 commit
  6. 07 May, 2020 5 commits
  7. 06 May, 2020 3 commits
  8. 05 May, 2020 1 commit
  9. 04 May, 2020 1 commit
  10. 01 May, 2020 1 commit
    • Deyu Fu's avatar
      Changes to make xentropysoftmax load/store vectorized when possible: (#725) · cf50dc7c
      Deyu Fu authored
      * Changes to make xentropysoftmax load/store vectorized when possible:
      Increase default ILP so that each thread handle 16 Bytes data in one step
      Make thread load/store longest vector possible
      Make unroll case handle adjacent data instead of strided, so same order compare to vector case
      
      * Add shift for not aligned case. Remove less than 16 bytes aligned access
      cf50dc7c
  11. 30 Apr, 2020 5 commits
  12. 29 Apr, 2020 5 commits
  13. 28 Apr, 2020 1 commit
  14. 23 Apr, 2020 1 commit
  15. 22 Apr, 2020 2 commits
    • Deyu Fu's avatar
    • Vinicius Reis's avatar
      Fix LARC with mixed precision (#793) · 2ec84ebd
      Vinicius Reis authored
      The LARC optimizer wraps an underlying optimizer and then needs to be passed
      to amp.initialize for mixed precision. There were 3 different crashes happening
      in this situation, fix all of them and add a unit test.
      
      I don't know if the 'LARC' in sys.modules check ever worked. In my setup, the
      entry in sys.modules is 'apex.parallel.LARC'. Checking if the variable is
      defined seems more reliable though.
      2ec84ebd
  16. 20 Apr, 2020 3 commits
  17. 16 Apr, 2020 4 commits