Remove cudaStreamSync. call from transformer_engine.cpp (#1518)
* Remove cudaStreamSync. call Signed-off-by:Vasudevan Rengasamy <vrengasamy@nvidia.com> * Use cudaMemsetAsync instead of cudaMemcpyAsync Signed-off-by:
Vasudevan Rengasamy <vrengasamy@nvidia.com> * Update transformer_engine/common/transformer_engine.cpp Co-authored-by:
Tim Moon <4406448+timmoon10@users.noreply.github.com> Signed-off-by:
Kirthi Shankar Sivamani <ksivamani@nvidia.com> --------- Signed-off-by:
Vasudevan Rengasamy <vrengasamy@nvidia.com> Signed-off-by:
Kirthi Shankar Sivamani <ksivamani@nvidia.com> Co-authored-by:
Kirthi Shankar Sivamani <ksivamani@nvidia.com> Co-authored-by:
Tim Moon <4406448+timmoon10@users.noreply.github.com>
Showing
Please register or sign in to comment