Commits · 3571a927a342e1cb584f4cb6705a1396c4446f5e · OpenDAS / DeepEP

09 Jul, 2025 1 commit
- add DeepEP_multi_port_nobond ibgda support · 3571a927
  liuhe authored Jul 09, 2025
  
  3571a927
10 Jun, 2025 1 commit

Fully remove barrier FIFO designs (#200) · 8da2d7b3

Chenggang Zhao authored Jun 10, 2025

* Fully remove FIFO slots

* Fully remove FIFO buffers

* Minor fix styles

* Fix some typos

* Bugs fixed

* Cleanup `ibgda_poll_cq`

8da2d7b3

28 May, 2025 1 commit
- Use IBGDA only (#177) · 9fe9021f
  Shangyan Zhou authored May 28, 2025
  
  9fe9021f
23 May, 2025 1 commit
- Code cleanup and bug fixed · 92405ddf
  Chenggang Zhao authored May 23, 2025
  
  92405ddf
22 Apr, 2025 3 commits
- Use `put_nbi_warp`. · e255d57b
  Shangyan Zhou authored Apr 22, 2025
  
  e255d57b
- Several code lints · edbb1bc3
  Chenggang Zhao authored Apr 22, 2025
  
  edbb1bc3
- Refactor some code. · 20b2aaaf
  Shangyan Zhou authored Apr 22, 2025
  
  20b2aaaf
21 Apr, 2025 2 commits

Revert `ibgda_device.cuh` and remove some comments. · e2c57848
Shangyan Zhou authored Apr 21, 2025

e2c57848

In the Internode Normal Kernel, when using nvshmem ibrc for RDMA data... · 5ab80c28

moningchen authored Apr 21, 2025

In the Internode Normal Kernel, when using nvshmem ibrc for RDMA data transmission, a single QP is used for data transfer between two GPUs, which limits kernel performance in network card dual-port and RoCE network scenarios.

In our optimized Internode Normal Kernel, we implemented multiple QPs for data transmission between two GPUs, setting a different QP for each channel. Additionally, we modified the transmission method from IBRC to IBGAD.

Through these optimizations, the Internode Normal Kernel achieves optimal performance in both H800 and H20 environments, with RDMA transmission performance nearly reaching the physical network performance limit. Using the current default statistical method, in 4-node H800 and H20 environments, RDMA performance can reach 60GB/s+.

5ab80c28

14 Mar, 2025 2 commits
- Fix style. · 38cdaf39
  Shangyan Zhou authored Mar 14, 2025
  
  38cdaf39
- Low latency kernels use rdma atomic to support AR. · 2d0cf41d
  Shangyan Zhou authored Mar 14, 2025
  
  2d0cf41d
05 Mar, 2025 1 commit
- Fix AR bugs for normal kernels · 458cdcb2
  Chenggang Zhao authored Mar 05, 2025
  
  458cdcb2
25 Feb, 2025 1 commit
- Initial commit · ebfe47e4
  Chenggang Zhao authored Feb 24, 2025
  
  ebfe47e4