- 29 Apr, 2021 1 commit
-
-
Chao Liu authored
-
- 28 Apr, 2021 2 commits
- 26 Apr, 2021 1 commit
-
-
Chao Liu authored
-
- 24 Apr, 2021 1 commit
-
-
Chao Liu authored
-
- 23 Apr, 2021 6 commits
- 22 Apr, 2021 5 commits
- 21 Apr, 2021 6 commits
- 20 Apr, 2021 2 commits
- 19 Apr, 2021 1 commit
-
-
Chao Liu authored
-
- 17 Apr, 2021 3 commits
- 14 Apr, 2021 1 commit
-
-
Chao Liu authored
-
- 13 Apr, 2021 2 commits
-
-
Chao Liu authored
* overhaul vector_type, make int8x4_t real vector instead of aliasing from int32_t
-
Chao Liu authored
* initial implementation for magic number division and DynamicMerge_v2_magic_division that uses it * turn off DynamicMerge_v2_magic_division that use magic number division by default
-
- 07 Apr, 2021 1 commit
-
-
zjing14 authored
* Hybrid direct + implicit GEMM forward convolution NCHWc v5r1. Input tensor bypass LDS. Support fp32/fp16/int8
-
- 06 Apr, 2021 2 commits
- 25 Mar, 2021 1 commit
-
-
Chao Liu authored
* support dynamic tensor descriptor * use buffer load OOB feature for padding case * add navi support * add int8x4 inference kernel Co-authored-by:
Chao Liu <chao@ixt-rack-81.local.lan> Co-authored-by:
Jing Zhang <jizhan@amd.com>
-
- 06 Aug, 2020 1 commit
-
-
Chao Liu authored
* fix buffer_store bug * remove obsolete kernels * add bwd-data-v5r1-nhwc
-
- 29 Jul, 2020 1 commit
-
-
Chao Liu authored
* Use buffer load built-in OOB check. buffer size is limited to 2GB. * buffer APIs use combined wave and thread offset * use uint32_t for addr shift in buffer addressing
-
- 24 Jun, 2020 1 commit
-
-
Chao Liu authored
* tuning para, * testing on v100 * add fp16 * remove deprecated tensor descriptor * sync with miopen * update build script Co-authored-by:Jing Zhang <jizhan@amd.com>
-
- 18 Feb, 2020 1 commit
-
-
Chao Liu authored
* renaming
-
- 17 Feb, 2020 1 commit
-
-
Chao Liu authored
* update for miopen integration: cosmetic refactor
-