1. 23 Jun, 2023 1 commit
  2. 01 Jun, 2023 1 commit
  3. 25 May, 2023 1 commit
  4. 20 May, 2023 1 commit
  5. 04 May, 2023 1 commit
  6. 28 Apr, 2023 1 commit
  7. 27 Apr, 2023 2 commits
  8. 24 Apr, 2023 2 commits
  9. 07 Apr, 2023 1 commit
  10. 05 Apr, 2023 1 commit
  11. 31 Mar, 2023 1 commit
    • Charlie Lin's avatar
      Split single dynamic dimension compiler pass (#1580) · e9e3eacc
      Charlie Lin authored
      Adds a new GPU compiler pass split_single_dyn_dim that handles when one input parameter has a single non-fixed dynamic_dimension.
      commonly occurs for dynamic batch or BERT sequence length
      Splits the dynamic shape into several submodules will static input parameters to handle all of the cases in the dynamic_dimension range.
      Essentially does what I manually did for the select_module verify tests
      Adds a compile option split_single_dyn_dim that toggles the pass on/off. Defaults to false.
      Updates verify_program.hpp and run_verify.cpp to allow for the tests to change the compile_options
      e9e3eacc
  12. 29 Mar, 2023 1 commit
  13. 21 Mar, 2023 1 commit
  14. 18 Mar, 2023 1 commit
  15. 17 Mar, 2023 2 commits
  16. 10 Mar, 2023 2 commits
  17. 28 Feb, 2023 1 commit
    • Charlie Lin's avatar
      Select module op (#1569) · a63ee2e0
      Charlie Lin authored
      Creates the select_module operator that selects one of the submodules passed to it to run based on the submodule parameters.  The submodule is selected by having the exact same static shapes for the arguments to select_module as the parameters in the submodule
      a63ee2e0
  18. 23 Feb, 2023 1 commit
  19. 16 Feb, 2023 1 commit
  20. 17 Jan, 2023 1 commit
  21. 13 Jan, 2023 1 commit
  22. 11 Jan, 2023 1 commit
  23. 09 Jan, 2023 1 commit
  24. 02 Nov, 2022 1 commit
  25. 28 Oct, 2022 1 commit
  26. 27 Oct, 2022 1 commit
    • kahmed10's avatar
      Add JIT pad (#1411) · 0d841ded
      kahmed10 authored
      updated GPU pad to now use JIT version.
      added range functions for JIT kernels.
      0d841ded
  27. 26 Oct, 2022 1 commit
  28. 19 Oct, 2022 2 commits
  29. 13 Oct, 2022 2 commits
  30. 04 Oct, 2022 1 commit
  31. 27 Sep, 2022 1 commit
  32. 21 Sep, 2022 1 commit
  33. 19 Sep, 2022 1 commit
    • Paul Fultz II's avatar
      Improve layernorm and reductions performance (#1348) · 97a1ed2d
      Paul Fultz II authored
      Compute mean and variance in same reduction
      Set block size to numbers divisible by 32 instead powers of 2
      Global is also set exactly instead of being divisible by block size
      More exact matching of global/local can help get rid of branching/loops
      Reduce vectors first before doing dpp_reduce
      Explicitly vectorize array operators since the compiler doesnt always vectorize them
      Still uses old for loop when its computing at compile-time since the reinterpret_cast nor the all the vector types is supported
      97a1ed2d
  34. 14 Sep, 2022 1 commit