1. 19 May, 2020 1 commit
  2. 18 May, 2020 1 commit
  3. 15 May, 2020 4 commits
  4. 14 May, 2020 1 commit
  5. 13 May, 2020 3 commits
  6. 12 May, 2020 5 commits
  7. 11 May, 2020 1 commit
  8. 09 May, 2020 1 commit
  9. 08 May, 2020 3 commits
  10. 07 May, 2020 5 commits
  11. 06 May, 2020 3 commits
  12. 05 May, 2020 1 commit
  13. 04 May, 2020 1 commit
  14. 01 May, 2020 1 commit
    • Deyu Fu's avatar
      Changes to make xentropysoftmax load/store vectorized when possible: (#725) · cf50dc7c
      Deyu Fu authored
      * Changes to make xentropysoftmax load/store vectorized when possible:
      Increase default ILP so that each thread handle 16 Bytes data in one step
      Make thread load/store longest vector possible
      Make unroll case handle adjacent data instead of strided, so same order compare to vector case
      
      * Add shift for not aligned case. Remove less than 16 bytes aligned access
      cf50dc7c
  15. 30 Apr, 2020 5 commits
  16. 29 Apr, 2020 4 commits