1. 31 Oct, 2023 2 commits
  2. 19 Oct, 2023 1 commit
  3. 13 Oct, 2023 1 commit
  4. 17 Jul, 2023 1 commit
  5. 02 Jul, 2023 1 commit
    • Paul Fultz II's avatar
      Improvement to ck integration (#1859) · 3c9df3b4
      Paul Fultz II authored
      Add a CI job to test CK
      Add MIGRAPHX_TUNE_CK env variable to only do tuning for CK
      Continue tuning even when there is invalid configs
      Fix a bug with parallel compilation not using all available threads
      Add additional test for gemms using half types
      Removed int32 as supported type since it doesnt pass our test suite
      3c9df3b4
  6. 15 Jun, 2023 1 commit
  7. 08 Jun, 2023 1 commit
  8. 23 May, 2023 1 commit
  9. 20 May, 2023 1 commit
  10. 19 May, 2023 1 commit
  11. 08 May, 2023 1 commit
  12. 24 Apr, 2023 2 commits
  13. 06 Apr, 2023 1 commit
  14. 10 Mar, 2023 1 commit
  15. 23 Feb, 2023 1 commit
  16. 16 Feb, 2023 2 commits
  17. 31 Jan, 2023 1 commit
    • Umang Yadav's avatar
      hipRTC fixes (#1531) · 91cc7242
      Umang Yadav authored
      Added CMakeFlag for hipRTC. MIGRAPHX_USE_HIPRTC.
      Added stages in Jenkins for hipRTC.
      Fixes for some of the pending issues from hipRTC.
      91cc7242
  18. 17 Jan, 2023 1 commit
  19. 11 Jan, 2023 1 commit
  20. 09 Jan, 2023 1 commit
  21. 07 Dec, 2022 1 commit
  22. 02 Nov, 2022 1 commit
  23. 27 Oct, 2022 1 commit
    • kahmed10's avatar
      Add JIT pad (#1411) · 0d841ded
      kahmed10 authored
      updated GPU pad to now use JIT version.
      added range functions for JIT kernels.
      0d841ded
  24. 04 Oct, 2022 1 commit
  25. 27 Sep, 2022 1 commit
  26. 21 Sep, 2022 1 commit
  27. 19 Sep, 2022 1 commit
    • Paul Fultz II's avatar
      Improve layernorm and reductions performance (#1348) · 97a1ed2d
      Paul Fultz II authored
      Compute mean and variance in same reduction
      Set block size to numbers divisible by 32 instead powers of 2
      Global is also set exactly instead of being divisible by block size
      More exact matching of global/local can help get rid of branching/loops
      Reduce vectors first before doing dpp_reduce
      Explicitly vectorize array operators since the compiler doesnt always vectorize them
      Still uses old for loop when its computing at compile-time since the reinterpret_cast nor the all the vector types is supported
      97a1ed2d
  28. 14 Sep, 2022 1 commit
  29. 06 Sep, 2022 1 commit
  30. 17 Aug, 2022 1 commit
  31. 16 Aug, 2022 1 commit
  32. 11 Jul, 2022 1 commit
  33. 05 Jul, 2022 1 commit
  34. 25 Jun, 2022 1 commit
  35. 22 Jun, 2022 1 commit
  36. 10 Jun, 2022 1 commit
  37. 24 May, 2022 1 commit