Commits · 4599cd69faae9279e2612a84dc04072cca688fb1 · tsoc / superbenchmark

02 Apr, 2026 6 commits
- Add dtk dockerfile for docker 18 · 4599cd69
  one authored Apr 02, 2026
  
  4599cd69
- Update docs · b8b080e2
  one authored Apr 02, 2026
  
  b8b080e2
- Re-implement kernel launch · 04564997
  one authored Apr 02, 2026
  
  04564997
- Fix runner test · 05cdf5d6
  one authored Apr 02, 2026
  
  05cdf5d6
- Use env file in docker instead of /tmp · c1bc12ce
  one authored Apr 02, 2026
  
  c1bc12ce
- Add topo mapping for dtk26.04 · c128dabb
  one authored Apr 02, 2026
  
  c128dabb
01 Apr, 2026 7 commits
- Update rocHPCG metrics · e514815d
  one authored Apr 01, 2026
  
  e514815d
- Add metric sorters for RCCL tests and rocHPCG · 05e137be
  one authored Apr 01, 2026
  
  05e137be
- Fix rocHPCG metric extraction · 742f203d
  one authored Apr 01, 2026
  
  742f203d
- Convert rochpcg script patch into shell script · b623c7e9
  one authored Apr 01, 2026
  
  b623c7e9
- Refactor environment variable handling in runner.py · a10c3e15
  one authored Apr 01, 2026
  
  a10c3e15
- Update dtk dockerfile to use venv · 325db60e
  one authored Apr 01, 2026
  
  325db60e
- Add gpu-hpcg metrics · 2056d7fa
  one authored Apr 01, 2026
  
  2056d7fa
31 Mar, 2026 1 commit
- Update dtk docker image · 4f69c7de
  one authored Mar 31, 2026
  
  4f69c7de
27 Mar, 2026 1 commit
- MicroBenchmark: rocHPCG · e4c2bd4c
  one authored Mar 27, 2026
  
  e4c2bd4c
25 Mar, 2026 1 commit
- Improve DTK gemm-flops · 211e63c7
  one authored Mar 25, 2026
  
  211e63c7
20 Mar, 2026 1 commit
- Fix paths, deps, envs in dockerfile · df0bde6c
  one authored Mar 20, 2026
  
  df0bde6c
19 Mar, 2026 3 commits

Migrate gpu-stream to BabelStream v5.0 · d4051602
one authored Mar 19, 2026

d4051602

Enhance DTK platform support and GPU detection · 1a57f2d6

one authored Mar 19, 2026

- Added Platform.DTK in the microbenchmark framework.
- Introduced new DTK hipblaslt benchmark class and corresponding tests.
- Updated Dockerfile to include hipblaslt-bench and its permissions.
- Registered DTK benchmarks in the benchmark registry for various performance tests.
- Enhanced GPU detection logic to recognize HYGON GPUs.

This update improves the benchmarking capabilities for DTK, ensuring compatibility and performance testing across platforms.

1a57f2d6

Update DTK dockerfile and microbenchmarks · c4f39919

one authored Mar 19, 2026

- Update rocm_commom.cmake for CMake>=3.24
- Prevent isolation build
- Add BabelStream as a submodule
- Update dockerignore

c4f39919

17 Mar, 2026 1 commit
- Add a dtk dockerfile · 0fdfe4c3
  one authored Mar 17, 2026
  
  0fdfe4c3
11 Mar, 2026 1 commit

Microbenchmark: upgrade Intel MLC to v3.12 in rocm5.0.x (#784) · 6b8e8104

Hongtao Zhang authored Mar 10, 2026



## Summary
- Upgrade Intel Memory Latency Checker from v3.11 to v3.12 in
rocm5.0.x.dockerfile
- Aligns with other dockerfiles that already use v3.12
Co-authored-by: Hongtao Zhang <hongtaozhang@microsoft.com>
Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com>

6b8e8104

04 Feb, 2026 1 commit

Submodule Update: update gpu-burn to newest version (#761) · 575859be

WenqingLan1 authored Feb 03, 2026



Updated 3rd party submodule gpu-burn to newest version for
implementation & doc support for cuda13.0.
Co-authored-by: guoshzhao <guzhao@microsoft.com>

575859be

28 Jan, 2026 1 commit

CI/CD - Fix Image build for cuda11.1.1 (#771) · 8b805d90

Hongtao Zhang authored Jan 28, 2026



**Description**

- When building the CUDA 11.1.1 image, pip (Python 3.8) cannot find a
pre-built wheel for the latest wandb release (v0.23.1). As a result, pip
attempts to build wandb from source. However, the build fails because
the image does not have Go installed, which is required for building
wandb from source. Then the error appears.

**Solution**

- For the CUDA 11.1.1 build, install the required build tools (e.g., Go,
Rust, and Cargo) needed for wandb.

---------
Co-authored-by: Hongtao Zhang <hongtaozhang@microsoft.com>
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>

8b805d90

21 Dec, 2025 1 commit

CI/CD - Fix Azure pipeline (#767) · c99380b4

Hongtao Zhang authored Dec 20, 2025



**Description**
Azure pipeline cpu-unit-test failed for "2025-12-10T03:47:59.0628597Z
ERROR: Could not install packages due to an OSError: [Errno 28] No space
left on device"

**Root Cause**
This happens because the matrix jobs (Python 3.7, 3.10, 3.12) run in
parallel and share the same VM's disk. Python 3.12 downloads
newer/larger packages (especially PyTorch and NVIDIA CUDA libraries
which are ~3GB+), and when multiple jobs run simultaneously, they
exhaust the disk space.

**Fix**
Disable the cache usage when installing SB
Co-authored-by: Hongtao Zhang <hongtaozhang@microsoft.com>

c99380b4

04 Dec, 2025 1 commit

Bug fix - update IB_DEVICES specification logic to fix ib-loopback test regression (#762) · e3fd943a

Henry Li authored Dec 03, 2025

**Description**

The ib-loopback test was regressed due to this recent
[change](https://github.com/microsoft/superbenchmark/commit/c65ae56713d6bfcc4a3be37d7fe24779590f9791).
When running ib-loopback using the standard
[config](https://github.com/microsoft/superbenchmark/blob/c65ae56713d6bfcc4a3be37d7fe24779590f9791/superbench/config/default.yaml#L69

),
the test would fail since it would pass numeric values like `0` into the
test command which would break since it is not a valid IB device name.

Example failure:

```
 [2025-11-25 22:08:38,100 vmssnc6ec000003:141056][micro_base.py:200][INFO] Execute command - round: 0, benchmark: ib-loopback, command: /usr/local/bin/run_perftest_loopback 47 45 /usr/local/b                                                                                                                                                        in/ib_write_bw -s 8388608 -F --iters=20000 -d 0 -p 45617 -x 0 --report_gbits.
[0]: IB device 0 not found
 Unable to find the Infiniband/RoCE device
IB device 0 not found
 Unable to find the Infiniband/RoCE device
[2025-11-25 22:08:39,113 vmssnc6ec000003:141056][micro_base.py:209][ERROR] Microbenchmark execution failed - round: 0, benchmark: ib-loopback, error message: IB device 0 not found
 Unable to find the Infiniband/RoCE device
IB device 0 not found
 Unable to find the Infiniband/RoCE device
```


**Major Revision**
- Major Revision A
- Major Revision B
- ...

**Minor Revision**
- Minor Revision A
- Minor Revision B
- ...

---------
Co-authored-by: Henry Li <lihl@microsoft.com>

e3fd943a

17 Nov, 2025 1 commit

Benchmarks: micro benchmarks - add --set_ib_devices option to auto-select IB... · c65ae567

Yuting Jiang authored Nov 17, 2025

Benchmarks: micro benchmarks - add --set_ib_devices option to auto-select IB device by MPI local rank in ib validation (#733)

**Description**
add --set_ib_devices option to auto-select IB device by MPI local rank 


**Major Revision**
- Add a new CLI flag --set_ib_devices to automatically select irregular
IB devices based on the MPI local rank.
- When enabled, the benchmark queries available IB devices via
network.get_ib_devices() and selects the device corresponding to
OMPI_COMM_WORLD_LOCAL_RANK.
- Fall back to existing --ib_dev behavior when the flag is not provided.

**Minor Revision**
- Add an env in network.get_ib_devices() to allow user to set the device
name

c65ae567

06 Nov, 2025 1 commit
- Fix pipelines - Update mlc version in dockerfiles from v3.11 to v3.12 (#752) · 25db1115
  WenqingLan1 authored Nov 06, 2025
```
Updated mlc wget link in dockerfiles.

---------
Co-authored-by: guoshzhao <guzhao@microsoft.com>
```
  25db1115
05 Nov, 2025 1 commit

CI/CD - Fix Azure test pipeline (#754) · 1b4377fc

Hongtao Zhang authored Nov 04, 2025

Python3.10 verification pipeline failed for conflict 'setuptools'
version as below.
<img width="1157" height="622" alt="image"
src="https://github.com/user-attachments/assets/ba0f6045-4b92-4fd8-b92f-1c474725534c

"
/>

Root Cause:
The problem is that modern pip (25.3) uses an isolated build environment
with the latest setuptools by default. The pipeline installs setuptools
65.7 in the user environment, but pip builds the package in an isolated
environment with newer setuptools, which conflicts with the version
check in [setup.py].

Solution:
Remove pip upgrade.

---------
Co-authored-by: Hongtao Zhang <hongtaozhang@microsoft.com>

1b4377fc

23 Oct, 2025 1 commit

Benchmarks: Micro benchmark - add ncu profile support in cublaslt-gemm (#740) · f6e65a98

Yuting Jiang authored Oct 23, 2025

**Description**
This PR adds NCU (NVIDIA Nsight Compute) profiling support to the
cublaslt-gemm micro benchmark, enabling detailed kernel analysis
including DRAM throughput, compute throughput, and launch arguments.

**Major Revision**
- Add --enable_ncu_profiling and --profiling_metrics for ncu profiling
- Modifies command execution to use NCU when profiling is enabled
- Updates result parsing to handle both standard and NCU profiled output
formats

f6e65a98

22 Oct, 2025 2 commits

Benchmarks: Micro benchmark - Support verification and parallel run for disk... · fe234262

Ziyue Yang authored Oct 22, 2025


Benchmarks: Micro benchmark - Support verification and parallel run for disk performance benchmark (#741)

**Description**
Adds verification and parallel run support for disk performance
benchmark.

**Major Revision**
- Adds `--verify` flag to support verify written data.
- Supports loading benchmark options from `PROC_RANK`, `BLOCK_DEVICES`
and `NUMA_NODES` environmental variables.

---------
Co-authored-by: guoshzhao <guzhao@microsoft.com>

fe234262

CI/CD - Fix python3.10 pipeline (#753) · 86a940c1

Hongtao Zhang authored Oct 21, 2025



**Description**
Python3.10 pipeline failed.

**Solution**
From log, 'bc' cmd is missing. Since our image tags are simple, the
solution is to remove 'bc' cmd directly.

---------
Co-authored-by: Hongtao Zhang <hongtaozhang@microsoft.com>

86a940c1

08 Oct, 2025 2 commits

Enhancement: Add nsys and pytorch profiler debug trace support (#744) · d804dbb6

Hongtao Zhang authored Oct 08, 2025



To improve benchmark debugging, the following debug methods were added:

pytorch profiler in model benchmark

- SB_ENABLE_PYTORCH_PROFILER: switch to enable/disable
- SB_TORCH_PROFILER_TRACE_DIR: log path
These 2 runtime variables need to be configured in SB config file.

nsys in SB runner

- SB_ENABLE_NSYS: switch to enable/disable 
- SB_NSYS_TRACE_DIR: log path
These 2 runtime variables need to be configured in runner's ENV

---------
Co-authored-by: Hongtao Zhang <hongtaozhang@microsoft.com>

d804dbb6

CI/CD - Fix image merge in GitHub Action. (#749) · b9864244
Yifan Xiong authored Oct 07, 2025
```
Fix image merge for release event in GitHub Action.
```
b9864244

01 Oct, 2025 1 commit

Dockerfile - add cuda13.0.dockerfile (#739) · 60189dd6

WenqingLan1 authored Oct 01, 2025



Add support for cuda13.0.
Add cuda13.0.dockerfile.
Add cuda13.0 image building task to github pipeline.
Update GPU STREAM to work with cuda13.0.
Fix data type conversion perf bug in GPU stream.
Update nvbandwidth submodule to be v0.8.
Update perftest submodule to be 4bee61f80d9e268fc97eaf40be00409e91d3a19e
(recent master).

---------
Co-authored-by: Ubuntu <dilipreddi@gmail.com>
Co-authored-by: guoshzhao <guzhao@microsoft.com>

60189dd6

30 Sep, 2025 1 commit

Benchmarks: Micro benchmark - Add simultanneously all-to-host / host-to-all... · 93e9d262

Yuting Jiang authored Sep 30, 2025

Benchmarks: Micro benchmark - Add simultanneously all-to-host / host-to-all bandwidth testcases to nvbandwidth (#736)

**Description**
Add simultanneously all-to-host / host-to-all bandwidth testcases to
nvbandwidth .

**Major Revision**
- nvbandwidth.patch: Add simultanneously all-to-host / host-to-all
bandwidth testcases to nvbandwidth
- upgrade nvbandwidth submodule into v0.8
- add patch into makefile build

93e9d262

29 Sep, 2025 2 commits
- Benchmark: Model benchmark - add option to exclude data copy time in model benchmarks (#734) · 76066b6d
  Yuting Jiang authored Sep 29, 2025
```
**Description**
add option to exclude data copy time in model benchmarks.

**Major Revision**
- add an option --no_copy
- move start time after data copy finish
```
  76066b6d
- Benchmarks: Micro benchmark - Add numa support for nvbandwidth (#742) · ad8e0143
  Yuting Jiang authored Sep 29, 2025
```
**Description**
Add numa support for nvbandwidth.
```
  ad8e0143
19 Sep, 2025 1 commit

Benchmarks: micro benchmarks - change cublasLtMatmulDescCreate scaleType from... · a7c4ed92

Yuting Jiang authored Sep 20, 2025

Benchmarks: micro benchmarks - change cublasLtMatmulDescCreate scaleType  from CUDA_R_32F to CUDA_R_16F in FP16 dist inference  (#732)

**Description**
change cublasLtMatmulDescCreate scaleType from CUDA_R_32F to CUDA_R_16F
in FP16 dist inference to fix cublaslt error.

a7c4ed92

12 Aug, 2025 1 commit

Release - SuperBench v0.12.0 (#729) · 0b4311cd

Hongtao Zhang authored Aug 12, 2025



**Description**

Cherry-pick bug fixes from v0.12.0 to main.

**Major Revisions**

* #725
* #727
* #728
Co-authored-by: Hongtao Zhang <hongtaozhang@microsoft.com>
Co-authored-by: Yifan Xiong <yixio@microsoft.com>
Co-authored-by: Guoshuai Zhao <guzhao@microsoft.com>

---------
Co-authored-by: Hongtao Zhang <hongtaozhang@microsoft.com>

0b4311cd