Commits · 671f24be50c30bec8263f2ce3d22e899ac4ef662 · gaoqiong / MIGraphX

11 May, 2022 1 commit

Prefuse layernorm for gpu (#1190) · 671f24be

Paul Fultz II authored May 11, 2022

Fuse layernorm and added triadd_layernorm fusion.  This is a prep performance booster

671f24be

10 May, 2022 1 commit
- Expose `add_literal` in C and Python API (#1173) · 5e5ed37a
  Umang Yadav authored May 10, 2022
```
Expose add_literal method in C/C++ api
```
  5e5ed37a
06 May, 2022 1 commit
- Add compile tests for gpu math functions (#1182) · 6a5cda96
  Paul Fultz II authored May 06, 2022
```
Add compile tests for gpu math functions
```
  6a5cda96
03 May, 2022 1 commit

Extend lifetimes in C++ API (#1139) · 4a5a23a4

Paul Fultz II authored May 02, 2022

Helps avoid dangling references. This also deprecates the constructors that didnt take a lifetime annotation since its ambiguous the lifetime.

4a5a23a4

29 Apr, 2022 1 commit
- Add GatherND operator (#1089) · 4ec35e5f
  turneram authored Apr 28, 2022
```
Add ref and gpu implementations for ONNX op GatherND

Resolves #1032
```
  4ec35e5f
26 Apr, 2022 1 commit
- Expose get_queue method for context in API (#1161) · 36656030
  Umang Yadav authored Apr 26, 2022
```
* expose get_queue method
```
  36656030
23 Apr, 2022 1 commit

ReverseSequence op (#1177) · 31906785

Charlie Lin authored Apr 22, 2022

Implements the ReverseSequence ONNX operator as a parser.

This parser can only handle a constant sequence_lens input. This is the same as what is handled for TensorRT as far as I can tell.
We could handle a variable sequence_lens input; that would require ref and GPU implementations of the operator.
The ONNX backend tests are disabled because this does not handle variable sequence_lens.

31906785

19 Apr, 2022 1 commit

Refactor Pooling and implement ONNX LpPool and GlobalLpPool (#1152) · 764273e4

Charlie Lin authored Apr 18, 2022

Refactored the reference implementation of pooling to something like what was done for roialign. Moved the reference implementation of pooling from targets/ref/lowering.cpp to pooling.hpp.
Removed cpu_pooling, instead using reference pooling in pooling.hpp
Added reference implementation of Lp Norm pooling and the global version
Added tests for the Lp Norm Pooling

764273e4

17 Apr, 2022 1 commit

Reduce with runtime compilation (#1150) · f9a5b81e

Paul Fultz II authored Apr 17, 2022

There is significant improvement on larger tensors with half almost 50% faster:

lens: [1024, 384, 768]
gpu::code_object[code_object=13832,symbol_name=kernel,global=39321600,local=256,]: 1.16685ms
gpu::reduce_sum[axes={2}]: 1.73126ms
Also for non-trivial layouts this can sometimes be over 2x faster:

lens: [64, 1024, 768, 4]
gpu::code_object[code_object=13832,symbol_name=kernel,global=39321600,local=256,]: 1.1706ms
gpu::reduce_sum[axes={1}]: 2.63375ms
Of course if the stride becomes larger this speed improvement diminishes due to poor memory access patterns. A lane_reduce instead of a block_reduce is needed for such type of kernels. I plan to address that in a future PR.

Finally, this also includes a MIGRAPHX_GPU_DUMP_ASM env variable which will print out the assembly when the kernel compiles.

f9a5b81e

14 Apr, 2022 1 commit

Half2 overloads (#1157) · 12007dba

bpickrel authored Apr 14, 2022

Issue 1127 Updates the math.hpp header file to perform overloads of various standard functions (ops) for the hip half2 type. The half2 type is two 16-bit floats packed into a 32-bit number and therefore the overloads act on vectors of sizes that are multiples of 2. They are invoked in runtime compilation any time one of the ops is called on a tensor declared with the data type shape::half_type.

Defined new template, made instances of the template for those math operations that the hip library contains, added verify tests for the sqrt operator for three cases:

tensor size not divisible by 2
tensor size divisible by 2 but not by 4
tensor size divisible by 4

12007dba

11 Apr, 2022 1 commit

scatter operator refactoring to include reduction (#1124) · 701c2014

bpickrel authored Apr 11, 2022

Change the "scatter" struct and op to a base/child set of three: scatter_none, scatter_add, scatter_mul to mirror Onnx' ScatterElements op. and its three reduction options. (Onnx Scatter op is deprecated and is equivalent to scatter_none.)

Provides both a reference op. and update to Onnx parsing. Tests updated and new test case added.

701c2014

08 Apr, 2022 1 commit
- Fix comparisons in migraphx::value class (#1146) · 1e0bbd78
  Paul Fultz II authored Apr 08, 2022
```
* Fix comparisons in migraphx::value class
```
  1e0bbd78
06 Apr, 2022 1 commit

Python Binding for the Manual Graph Buidling (#1143) · c4b6469a

Umang Yadav authored Apr 06, 2022

Adds following API binding and tests to python :

add_return
add_instruction
add_parameter
create_module.

c4b6469a

01 Apr, 2022 1 commit

Update developer overview, fix doc CMakeLists (#1140) · 0295965d

Charlie Lin authored Apr 01, 2022

* Fix and change doc CMakeLists
1. Fix include directory location with hange from #1088
2. Create a DoxygenWarningLog.txt file in <build_dir>/doc/doxygen
3. Move compiled html or pdf files to <build_dir>/doc/[pdf, html]

0295965d

31 Mar, 2022 1 commit
- Change the doc to mention only gpu or ref as targets (#1153) · c59f4079
  Umang Yadav authored Mar 31, 2022
```
Documentation update for valid targets
```
  c59f4079
29 Mar, 2022 2 commits

Python binding for shape : Fix constructor for the shape and enable tests (#1135) · b5c96d34
Umang Yadav authored Mar 29, 2022
```
Follow up to #1128
```
b5c96d34

Refactor runtime compiled kernels to use the same compile_ops pipeline (#1125) · 661046c6

Paul Fultz II authored Mar 29, 2022

This adds the infrastructure so we can compile everything in parallel, whereas before only pointwise kernels were compiled in parallel. This will also directly integrate with lowering and the gpu-driver. The kernels for pointwise and roialign are using this infrastructure. Scatternd is not since it does require standard shape.

This also makes it easier to add new runtime compiled kernels in the future.

661046c6

25 Mar, 2022 1 commit
- Improve handling of string literals in value class (#1141) · c73c0dae
  Paul Fultz II authored Mar 25, 2022
```
* Handle string literal in construction
* Improve get_default with vector
```
  c73c0dae
24 Mar, 2022 1 commit
- Add initial experimental custom op (#1109) · 251cdd74
  Paul Fultz II authored Mar 24, 2022
```
This creates a custom op which has name() and compute_shape() methods. 
```
  251cdd74
21 Mar, 2022 1 commit
- Lp normalization op (#1129) · 03225b57
  Charlie Lin authored Mar 21, 2022
```
* LpNormalization ONNX parser
```
  03225b57
18 Mar, 2022 2 commits

Complete GPU implementation of CumSum op (#1094) · 548783c8

turneram authored Mar 18, 2022

Add exclusive and reverse modes to gpu implementation of prefix_scan_sum, which completes support for ONNX op CumSum

548783c8

Make get_context experimental (#1137) · e521fa3f

Paul Fultz II authored Mar 18, 2022

The get_context may change in the future(when we support multi-targets) so make this experimental for now.

e521fa3f

15 Mar, 2022 1 commit

Expose APIs for the MIGraphX program (#1093) · 64e79a94

Umang Yadav authored Mar 15, 2022

API includes following
create_module,
get_main_module
add_instruction without module args
add_instruction with module args
add_parameter
add_return

64e79a94

11 Mar, 2022 1 commit

Improve print ins (#1096) · b3b44f5d

Shucai Xiao authored Mar 11, 2022

The module::debug_print(ins) is very slow, which makes the trave_eval==1/2 very slow. The reason is printing an ins involves search the whole module to get the instruction, the print it.  This change is to fix that by calling module::print() to get names of all instructions of a program, then print the instruction by getting its name from a hash map.

b3b44f5d

09 Mar, 2022 3 commits
- Celu ONNX parser and tests (#1114) · 5b37c53c
  Charlie Lin authored Mar 09, 2022
```
Add Celu ONNX operator
```
  5b37c53c
- Add python API to construct shape class (#1128) · 4467c158
  Paul Fultz II authored Mar 09, 2022
```
Add python API to construct shape class
```
  4467c158
- Expose context in C++ API (#1118) · 0e6bd17c
  kahmed10 authored Mar 09, 2022
```
Add a callable C++ API to migraphx
```
  0e6bd17c
08 Mar, 2022 1 commit
- Size ONNX op (#1122) · d71a7b6a
  Charlie Lin authored Mar 08, 2022
```
* Implement size ONNX operator and tests
```
  d71a7b6a
07 Mar, 2022 1 commit
- Use `add_common_op` for handling types and broadcast in Clip Onnx parsing (#1121) · a0ae2f79
  Umang Yadav authored Mar 07, 2022
```
add_common_op for parse_clip
Should fix #1119
```
  a0ae2f79
04 Mar, 2022 2 commits

EyeLike Operator (#1087) · 8f184d4a
Charlie Lin authored Mar 04, 2022
```
Adds EyeLike ONNX parser and unit tests.
```
8f184d4a

Mode as enum for pooling and roi_align (#1091) · a2e90b5d

bpickrel authored Mar 04, 2022

Changed the pooling values for two structures from strings to specialized enum classes. Many test and operator parsing changes to support this. Introduces one new source file, op_enums.cpp.

a2e90b5d

03 Mar, 2022 1 commit
- Add ScatterND operator (#1074) · 832f28c6
  turneram authored Mar 02, 2022
```
Add onnx parser and ref and gpu implementations of ONNX op ScatterND
```
  832f28c6
02 Mar, 2022 2 commits
- isnan operator (#1100) · bfedcd45
  Charlie Lin authored Mar 02, 2022
```
Implements the IsNaN operator, ref, gpu, and onnx parser.
```
  bfedcd45
- Clang format ver10 (#1106) · 9852aaef
  bpickrel authored Mar 02, 2022
```
Update the base version of clang-format from 5.0 to 10.0
```
  9852aaef
25 Feb, 2022 3 commits
- Add with_type to shape class (#1102) · 85b0563c
  Paul Fultz II authored Feb 25, 2022
```
Add with_type to shape class
```
  85b0563c
- Add reverse lookup of c++ class to c class (#1099) · 40c087bd
  Paul Fultz II authored Feb 25, 2022
```
Needed for custom_op so we can generically convert the C type back to the C++ type in the function pointer.
```
  40c087bd
- Add get_queue to context to get the current stream (#1097) · e5242676
  Paul Fultz II authored Feb 24, 2022
```
wrapped in a any_ptr class so the type can be checked at runtime for a mismatch.
```
  e5242676
24 Feb, 2022 1 commit

Some cmake fixes and updates (#1088) · cd0a4aa5

Paul Fultz II authored Feb 23, 2022

Make doc/CMakeLists.txt standalone
Switch to use rocm-cmake modules for document generation
Add CONFIGURE_DEPENDS to file(GLOB) so it will update without an explicit cmake run
Add STRINGS property for build type to make it easier to switch build types with ccmake
Various fixes and improvements

cd0a4aa5

23 Feb, 2022 1 commit

Keep std shape (#1059) · 98dfdf15

Shucai Xiao authored Feb 23, 2022

This PR is the resolve two problems in the issue#999, i.e., non_standard_shape input to reshape and reduce_mean.
Three fixes:

Any operator that has a standard shape requirement will add a contiguous input for its input.
Eliminate_contiguous, when computing whether a contiguous can be removed, we should use all the updated args, not just the one that is being checked.
In two optimization in the simplify_reshape, we remove the contiguous in the reshaper name list, since eliminate_contiguous will remove the contiguous if it can be removed.
the solution is add an attribute to the operator that requires standard input shape, then in the auto_contiguous pass, add a contiguous to every input of such operators.

98dfdf15

16 Feb, 2022 1 commit
- Support nonstandard shapes for the UnSqueeze Op (#1071) · 4480eb79
  Umang Yadav authored Feb 16, 2022
```
Support nonstandard shapes like slice, broadcast and transpose for the unsqueeze op
```
  4480eb79