- 20 Jan, 2020 1 commit
-
-
Chao Liu authored
* Added bwd data v3r1: breaking down compute into a series of load balanced GEMM, and launch in a single kernel * Added bwd data v4r1: like v3r1, but launch GEMMs in multiple kernels * Tweaked v1r1 and v1r2 (atomic) on AMD GPU
-
- 03 Dec, 2019 1 commit
-
-
Chao Liu authored
* enabled atomic add in tensor copy * added gridwise GEMM * added backward data conv using GEMM + atomic * added backward data conv using GEMM, no atomic
-
- 13 Jun, 2019 1 commit
-
-
Chao Liu authored
-
- 12 Jun, 2019 2 commits
- 15 Feb, 2019 2 commits
- 14 Feb, 2019 1 commit
-
-
Chao Liu authored
-
- 22 Oct, 2018 1 commit
-
-
Chao Liu authored
-
- 09 Oct, 2018 1 commit
-
-
Chao Liu authored
-