Commits · 0fb17f71eb9f80514eb9dc5609072b45db03a4c8 · gaoqiong / MIGraphX

29 Sep, 2022 1 commit
- Fix context-free output_shape · 0fb17f71
  charlie authored Sep 29, 2022
  
  0fb17f71
28 Sep, 2022 5 commits
- Add convert handle dynamic shape check_shapes · ff195f97
  charlie authored Sep 28, 2022
  
  ff195f97
- Add onnx files · 69900d77
  charlie authored Sep 28, 2022
  
  69900d77
- Merge branch 'develop' of github.com:ROCmSoftwarePlatform/AMDMIGraphX into refactor_dynamic_compute · acad34c6
  charlie authored Sep 28, 2022
  
  acad34c6
- Unary ops changes and tests · 65e14286
  charlie authored Sep 28, 2022
  
  65e14286
- Add compute_fp32 flag for quant_gemm tests (#1360) · 70e63960
  Umang Yadav authored Sep 28, 2022
```
test_gpu_pack_int8_args fails on gfx908 machine, because it doesn't set compute_fp32 flag correctly. This PR fixes the test such that it checks for the device-name, and rocblas-versions and sets this flag accordingly.
```
  70e63960
27 Sep, 2022 5 commits
- Adding tests · 30243d2c
  charlie authored Sep 27, 2022
  
  30243d2c
- Dynamic unary function · a56e1601
  charlie authored Sep 27, 2022
  
  a56e1601
- Check dynamic() on shape when using dyn_output · 07c05efb
  charlie authored Sep 27, 2022
  
  07c05efb
- Remove stuff commented out · da5b6fef
  charlie authored Sep 27, 2022
  
  da5b6fef
- Add onnx mod operator gpu cpu (#1306) · 40118191
  Ted Themistokleous authored Sep 26, 2022
```
Implement operator for CPU and GPU implementations
```
  40118191
26 Sep, 2022 5 commits
- Rewrite ONNX parse batch norm (#1362) · c00f8202
  Charlie Lin authored Sep 26, 2022
```
Rewrites the BatchNormalization ONNX operator into other MIGX operators
- Added handling of 1D input tensor case (edge case in ONNX spec)
Removes the spatial and per_activation functionality (not in the ONNX spec)
- Did not remove the batch_norm_inference related code as the TensorFlow parser still uses it
- Can remove that code when the TF version is updated
```
  c00f8202
- Fixed using pack() correctly · f1c18355
  charlie authored Sep 26, 2022
  
  f1c18355
- Use larger vector size instead of preloading for broadcasted inputs (#1389) · 492c4a6c
  Paul Fultz II authored Sep 26, 2022
  
  492c4a6c
- Merge branch 'develop' of github.com:ROCmSoftwarePlatform/AMDMIGraphX into refactor_dynamic_compute · b76a9043
  charlie authored Sep 26, 2022
  
  b76a9043
- Upgrade cppcheck to 2.9 (#1400) · 66bbff1e
  Paul Fultz II authored Sep 26, 2022
```
Upgrade cppcheck to 2.9 
```
  66bbff1e
24 Sep, 2022 2 commits

check concurrency on PR level with one running and one pending performance tests (#1401) · 94bc41dc

Chris Austen authored Sep 24, 2022

Workflow has concurrency reintroduced with different set of rules. New expected behavior is to check concurrency on PR level with one running and one pending performance tests. In case of multiple commits in same PR, always the latest commit is queued after initiated performance test execution is completed. Any other PRs/commits are in pending/queued state

94bc41dc

update codecov version (#1402) · 1b575b5c
Chris Austen authored Sep 24, 2022
```
Codecov announced deprecating the bash uploader. Using updated uploader
```
1b575b5c

23 Sep, 2022 2 commits
- Still broken, figuring things out · 68c17b1b
  charlie authored Sep 23, 2022
  
  68c17b1b
- Remove unused device functions (#1394) · 8ea8473d
  Paul Fultz II authored Sep 23, 2022
```
* Remove device functions
* Update tests
```
  8ea8473d
22 Sep, 2022 2 commits
- Initial · f02f5d98
  charlie authored Sep 22, 2022
  
  f02f5d98
- expose to_shapes(const std::vector<argument>& args) · e0cb7b9a
  charlie authored Sep 22, 2022
  
  e0cb7b9a
21 Sep, 2022 2 commits

Parameterize epsilon for layernorm kernel (#1367) · d9578ba6

kahmed10 authored Sep 21, 2022

This PR allows for other values of epsilon to be matched when finding layernorm. Similarly, the calculation now uses the variable for epsilon.

d9578ba6

Multibroadcast find_mul_conv (#1384) · 9a70050b

Charlie Lin authored Sep 21, 2022

Change find_mul_conv to work with multibroadcast also. Checks the strides instead of the broadcast axis.

9a70050b

19 Sep, 2022 4 commits

Merge branch 'develop' into refactor_auto_pad_conv · 2b936b13
Charlie Lin authored Sep 19, 2022

2b936b13

Improve layernorm and reductions performance (#1348) · 97a1ed2d

Paul Fultz II authored Sep 19, 2022

Compute mean and variance in same reduction
Set block size to numbers divisible by 32 instead powers of 2
Global is also set exactly instead of being divisible by block size
More exact matching of global/local can help get rid of branching/loops
Reduce vectors first before doing dpp_reduce
Explicitly vectorize array operators since the compiler doesnt always vectorize them
Still uses old for loop when its computing at compile-time since the reinterpret_cast nor the all the vector types is supported

97a1ed2d

Fix MLIR test · ca360585
charlie authored Sep 19, 2022

ca360585
Disabled concurrency, queue added to perf-test.yml (#1386) · 34c08db7
Chris Austen authored Sep 19, 2022

34c08db7

16 Sep, 2022 7 commits
- Fix typo for add_sigmoid (#1385) · 10f37f49
  Umang Yadav authored Sep 16, 2022
```
* fix typo for add_sigmoid
```
  10f37f49
- Fix comments · b3c6b7eb
  charlie authored Sep 16, 2022
  
  b3c6b7eb
- Merge branch 'develop' of github.com:ROCmSoftwarePlatform/AMDMIGraphX into refactor_auto_pad_conv · 8378a397
  charlie authored Sep 16, 2022
  
  8378a397
- Naming fix · f60297db
  charlie authored Sep 16, 2022
  
  f60297db
- Fix normalize attribute, fix same_lower bug · ef738568
  charlie authored Sep 16, 2022
  
  ef738568
- Progress on changing padding_mode · 0afab294
  charlie authored Sep 16, 2022
```
Weird bug with ref padding shape
still need to change parse_convolution
```
  0afab294
- Update deprecated Pybind constructor (#1382) · 255fb11a
  Umang Yadav authored Sep 16, 2022
```
* remove deprecated constructor
```
  255fb11a
15 Sep, 2022 2 commits

[mlir] Replaced `find_library` with `find_package` to locate MLIR static library (#1373) · e1e36cdc

Lixun Zhang authored Sep 15, 2022

* Replaced `find_library` with `find_package` to locate MLIR static library
* Unified the include dir for headers and remove backward compatibility
* Embedded the external/include dir into the exported library

e1e36cdc

Initial · 376b18af
charlie authored Sep 15, 2022

376b18af

14 Sep, 2022 3 commits
- Reduce problem size of unbatched_gemm tests (#1383) · 333860ce
  turneram authored Sep 14, 2022
```
The verify tests from pr #1354 were still causing some codecov timeouts after merge. This PR further reduces the problem sizes to avoid these failures.
```
  333860ce
- Fix split_reshape for slice len of 1 (#1379) · 4b76dd0d
  Umang Yadav authored Sep 14, 2022
```
* fix slice_dim1 for case
```
  4b76dd0d
- Implement concat using jit compilation (#1356) · 7662d9c0
  Paul Fultz II authored Sep 14, 2022
```
* Implement concat using jit compilation
```
  7662d9c0