Commits · 045649973a9c6286398b4acce20e629b378e41c0 · tsoc / superbenchmark

02 Apr, 2026 1 commit
- Re-implement kernel launch · 04564997
  one authored Apr 02, 2026
  
  04564997
01 Apr, 2026 3 commits
- Update rocHPCG metrics · e514815d
  one authored Apr 01, 2026
  
  e514815d
- Fix rocHPCG metric extraction · 742f203d
  one authored Apr 01, 2026
  
  742f203d
- Add gpu-hpcg metrics · 2056d7fa
  one authored Apr 01, 2026
  
  2056d7fa
27 Mar, 2026 1 commit
- MicroBenchmark: rocHPCG · e4c2bd4c
  one authored Mar 27, 2026
  
  e4c2bd4c
25 Mar, 2026 1 commit
- Improve DTK gemm-flops · 211e63c7
  one authored Mar 25, 2026
  
  211e63c7
19 Mar, 2026 3 commits

Migrate gpu-stream to BabelStream v5.0 · d4051602
one authored Mar 19, 2026

d4051602

Enhance DTK platform support and GPU detection · 1a57f2d6

one authored Mar 19, 2026

- Added Platform.DTK in the microbenchmark framework.
- Introduced new DTK hipblaslt benchmark class and corresponding tests.
- Updated Dockerfile to include hipblaslt-bench and its permissions.
- Registered DTK benchmarks in the benchmark registry for various performance tests.
- Enhanced GPU detection logic to recognize HYGON GPUs.

This update improves the benchmarking capabilities for DTK, ensuring compatibility and performance testing across platforms.

1a57f2d6

Update DTK dockerfile and microbenchmarks · c4f39919

one authored Mar 19, 2026

- Update rocm_commom.cmake for CMake>=3.24
- Prevent isolation build
- Add BabelStream as a submodule
- Update dockerignore

c4f39919

17 Nov, 2025 1 commit

Benchmarks: micro benchmarks - add --set_ib_devices option to auto-select IB... · c65ae567

Yuting Jiang authored Nov 17, 2025

Benchmarks: micro benchmarks - add --set_ib_devices option to auto-select IB device by MPI local rank in ib validation (#733)

**Description**
add --set_ib_devices option to auto-select IB device by MPI local rank 


**Major Revision**
- Add a new CLI flag --set_ib_devices to automatically select irregular
IB devices based on the MPI local rank.
- When enabled, the benchmark queries available IB devices via
network.get_ib_devices() and selects the device corresponding to
OMPI_COMM_WORLD_LOCAL_RANK.
- Fall back to existing --ib_dev behavior when the flag is not provided.

**Minor Revision**
- Add an env in network.get_ib_devices() to allow user to set the device
name

c65ae567

23 Oct, 2025 1 commit

Benchmarks: Micro benchmark - add ncu profile support in cublaslt-gemm (#740) · f6e65a98

Yuting Jiang authored Oct 23, 2025

**Description**
This PR adds NCU (NVIDIA Nsight Compute) profiling support to the
cublaslt-gemm micro benchmark, enabling detailed kernel analysis
including DRAM throughput, compute throughput, and launch arguments.

**Major Revision**
- Add --enable_ncu_profiling and --profiling_metrics for ncu profiling
- Modifies command execution to use NCU when profiling is enabled
- Updates result parsing to handle both standard and NCU profiled output
formats

f6e65a98

22 Oct, 2025 1 commit

Benchmarks: Micro benchmark - Support verification and parallel run for disk... · fe234262

Ziyue Yang authored Oct 22, 2025


Benchmarks: Micro benchmark - Support verification and parallel run for disk performance benchmark (#741)

**Description**
Adds verification and parallel run support for disk performance
benchmark.

**Major Revision**
- Adds `--verify` flag to support verify written data.
- Supports loading benchmark options from `PROC_RANK`, `BLOCK_DEVICES`
and `NUMA_NODES` environmental variables.

---------
Co-authored-by: guoshzhao <guzhao@microsoft.com>

fe234262

01 Oct, 2025 1 commit

Dockerfile - add cuda13.0.dockerfile (#739) · 60189dd6

WenqingLan1 authored Oct 01, 2025



Add support for cuda13.0.
Add cuda13.0.dockerfile.
Add cuda13.0 image building task to github pipeline.
Update GPU STREAM to work with cuda13.0.
Fix data type conversion perf bug in GPU stream.
Update nvbandwidth submodule to be v0.8.
Update perftest submodule to be 4bee61f80d9e268fc97eaf40be00409e91d3a19e
(recent master).

---------
Co-authored-by: Ubuntu <dilipreddi@gmail.com>
Co-authored-by: guoshzhao <guzhao@microsoft.com>

60189dd6

29 Sep, 2025 1 commit
- Benchmarks: Micro benchmark - Add numa support for nvbandwidth (#742) · ad8e0143
  Yuting Jiang authored Sep 29, 2025
```
**Description**
Add numa support for nvbandwidth.
```
  ad8e0143
19 Sep, 2025 1 commit

Benchmarks: micro benchmarks - change cublasLtMatmulDescCreate scaleType from... · a7c4ed92

Yuting Jiang authored Sep 20, 2025

Benchmarks: micro benchmarks - change cublasLtMatmulDescCreate scaleType  from CUDA_R_32F to CUDA_R_16F in FP16 dist inference  (#732)

**Description**
change cublasLtMatmulDescCreate scaleType from CUDA_R_32F to CUDA_R_16F
in FP16 dist inference to fix cublaslt error.

a7c4ed92

30 Jun, 2025 1 commit

Benchmarks: Add Mixture of Experts Model (#679) · 44e35cda

pdr authored Jun 30, 2025



Added MoE model using MixtralConfig. 

1. Added 8x7b and 8x22b variants 
2. Requires high VRAM as all experts are loaded in memory. Thus,
disabled training due to memory constraint on test worker.

---------
Co-authored-by: Hongtao Zhang <garyworkzht@gmail.com>
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
Co-authored-by: Hongtao Zhang <hongtaozhang@microsoft.com>

44e35cda

24 Jun, 2025 1 commit

Benchmarks - Add FP4 GEMM FLOPS support for cublaslt_gemm benchmark (#711) · b795477e

guoshzhao authored Jun 24, 2025



**Description**
Add FP4 precision support for cublaslt_gemm benchmark.

**Major Revision**
- Add new type `fp4e2m1` and `__nv_fp4_e2m1`.
- For FP4 matmul, precision of MatrixC (add) should be FP16, precision
of MatricD (output) should be FP4, otherwise, it will not work.
- Add macro `CUDA_VERSION` to resolve the compatibility issue of
different CUDA versions.

---------
Co-authored-by: Ubuntu <aiperf@aiperf000000.hp5z1gqeinfufbj2u3jcty5fme.cdmx.internal.cloudapp.net>
Co-authored-by: AVA <39534996+avazr@users.noreply.github.com>
Co-authored-by: Guoshuai Zhao <microsoft@microsoft.com>

b795477e

20 Jun, 2025 2 commits

Benchmark - Support autotuning in cublaslt gemm (#706) · 60b13256

Babak Hejazi authored Jun 20, 2025

**Description**
Enable autotuning as an opt-in mode when benchmarking cublasLt via
`cublaslt_gemm`

The implementation is based on
https://github.com/NVIDIA/CUDALibrarySamples/blob/master/cuBLASLt/LtSgemmSimpleAutoTuning/sample_cublasLt_LtSgemmSimpleAutoTuning.cu

The behavior of original benchmark command remains unchanged, e.g.:
- `cublaslt_gemm -m 2048 -n 12288 -k 1536 -w10000 -i 1000 -t fp8e4m3`

The new opt-in options are `-a` (for autotune) and `-I` (for autotune
iterations, default is 50, same as the default for `-i`) and `-W` (for
autotune warmups, default=20, same as the default for `-w`), e.g.:
- `cublaslt_gemm -m 2048 -n 12288 -k 1536 -w 10000 -i 1000 -t fp8e4m3
-a`
- `cublaslt_gemm -m 2048 -n 12288 -k 1536 -w 10000 -i 1000 -t fp8e4m3 -a
-I 10 -W 10`

**Note:** This PR also changes the default `gemm_compute_type` for BF16
and FP16 to `CUBLAS_COMPUTE_32F`.

**Further observations:** 
1. The support matrix of the `cublaslt_gemm` could be further extended
in the future to support non-FP16 output as well for FP8 inputs.
2. Currently, the input matrices are initialized with values of 1.0 and
2.0 which makes them less demanding in terms of power. Another future
extension could be to enable another fill mode for, say, uniform random
numbers between -1 and 1.
3. cuBLAS workspace recommendations are listed under
https://docs.nvidia.com/cuda/cublas/#cublassetworkspace



Update (June 10, 2025): verified using higher level test driver with
these commands:

1. inline:
```
python3 -c "                                                                            
from superbench.benchmarks import BenchmarkRegistry, Platform
from superbench.common.utils import logger

parameters = (
    '--num_warmup 10 --num_steps 50 '
    '--shapes 512,512,512 1024,1024,1024 --in_types fp16 fp32 '
    '--enable_autotune --num_warmup_autotune 20 --num_steps_autotune 50'
)
context = BenchmarkRegistry.create_benchmark_context(
    'cublaslt-gemm', platform=Platform.CUDA, parameters=parameters
)
benchmark = BenchmarkRegistry.launch_benchmark(context)
logger.info('Result: {}'.format(benchmark.result))
"
```

2. newly added script: 
`python3 examples/benchmarks/cublaslt_function.py`

---------
Co-authored-by: Babak Hejazi <babakh@nvidia.com>

60b13256

Benchmark - Add Grace CPU support for CPU Stream (#719) · 0b8d1fd4

WenqingLan1 authored Jun 19, 2025



**Description**
Added support for Grace CPU neo2 architecture in CPU Stream. Now CPU
Stream supports dual socket benchmarking.

Example config for this arch support:
```yaml
    cpu-stream:numa0:
      timeout: *default_timeout
      modes:
      - name: local
        parallel: no
      parameters:
        cpu_arch: neo2
        numa_mem_nodes: 0
        cores: 0 1 2 3 4 5 6 7 8
    cpu-stream:numa1:
      timeout: *default_timeout
      modes:
      - name: local
        parallel: no
      parameters:
        cpu_arch: neo2
        numa_mem_nodes: 1
        cores: 64 65 66 67 68 69 70 71 72
    cpu-stream:numa-spread:
      timeout: *default_timeout
      modes:
      - name: local
        parallel: no
      parameters:
        cpu_arch: neo2
        numa_mem_nodes: 0 1
        cores: 0 1 2 3 4 5 6 7 8 64 65 66 67 68 69 70 71 72
```

---------
Co-authored-by: dpower4 <dilipreddi@gmail.com>

0b8d1fd4

18 Jun, 2025 1 commit

Benchmarks - Add GPU Stream Micro Benchmark (#697) · 4eddd50a

WenqingLan1 authored Jun 18, 2025

Added GPU Stream benchmark - measures the GPU memory bandwidth and
efficiency for double datatype through various memory operations
including copy, scale, add, and triad.
- added documentation for `gpu-stream` detailing its introduction,
metrics, and descriptions.
- added unit tests for `gpu-stream`. Example output is in
`superbenchmark/tests/data/gpu_stream.log`.

4eddd50a

14 Jun, 2025 1 commit

microbenchmark - CPU Stream Benchmark Revise (#712) · 991c0051

Hongtao Zhang authored Jun 14, 2025



In the current implementation, the CPU‑stream benchmark code renames the
binary before the microbench base class can verify its existence,
causing the default‐binary check to fail.

This PR adds a “default” binary—built with the standard compile
parameters—so that the base class can always find and validate it. Once
the default binary is in place, the CPU‑stream code will rename it as
needed and re‑check its presence before running the benchmark.

The PR also enable CPU stream in the default settings.

---------
Co-authored-by: Hongtao Zhang <hongtaozhang@microsoft.com>

991c0051

01 May, 2025 1 commit
- cuda arch flag for cublaslt (#701) · 3e090482
  pdr authored Apr 30, 2025
```
adding gb200 cuda arch flag for cublaslt compilation
```
  3e090482
21 Mar, 2025 1 commit

Dockerfile - Support cuda12.8 for Blackwell arch (#682) · 294f1f20

pdr authored Mar 20, 2025



**Description**
Updated docker for 12.8
Use cutlass latest relase 3.8 with ARCH 100(blackwell) support
add latest nccl-test release with ARCH 100(blackwell) 
Updated msccl to support build for sm_100
No breaking changes, so backward compatible tested with  cuda 12.4

---------
Co-authored-by: Hongtao Zhang <garyworkzht@gmail.com>

294f1f20

25 Feb, 2025 1 commit

Docs - Fix typos in documentation and code files (#686) · 71573f3c

Maxim Evtush authored Feb 25, 2025


Co-authored-by: Yifan Xiong <yifan.xiong@microsoft.com>
Co-authored-by: Hongtao Zhang <garyworkzht@gmail.com>

71573f3c

15 Feb, 2025 1 commit

Bugfix: Avoid Unintended nvbandwidth Function Calls in All Benchmarks (#685) · 41a484fa

Hongtao Zhang authored Feb 14, 2025



Root Cause:

1. '_get_all_test_cases()' was called in '_parser' while '_parser' was
defined in the base class.
2.  in '_get_all_test_cases()', cmd path was not included.

Fix:

1. Remove '_get_all_test_cases()' from '_parser'.
2. Construct path for cmd.

---------
Co-authored-by: hongtaozhang <hongtaozhang@microsoft.com>

41a484fa

05 Feb, 2025 2 commits

Bugfix - nvbandwidth benchmark need to handle N/A value (#675) · 45d06647

Hongtao Zhang authored Feb 05, 2025



**Description**

1. Fixed the bug that nvbandwidth benchmark need to handle 'N/A' values
in nvbandwidth cmd output.
2. Replaced the input format of test cases with a list.
3. Add nvbandwidth configuration example in default config files.

---------
Co-authored-by: hongtaozhang <hongtaozhang@microsoft.com>
Co-authored-by: Yifan Xiong <yifan.xiong@microsoft.com>

45d06647

Bug - Fix tensorrt-inference parsing (#674) · 7af7c0b7

Kirill Prosvirov authored Feb 05, 2025

**Description**
Today I was running a benchmark on my machine. And encountered a fancy
issue with tensorrt-inference.
I got code 33, which according to the source code is:
```
MICROBENCHMARK_RESULT_PARSING_FAILURE = 33
```
I dived deep into the code and found out the following problem. The
parser stumbled upon getting to the following line:
```
[11/28/2024-17:03:11] [I] Latency: min = 7.2793 ms, max = 10.1606 ms, mean = 7.41642 ms, median = 7.39551 ms, percentile(99%) = 8 ms
```
I ran it separately on the code and found out that the regular
expression was not suitable for the cases like this, when you encounter
an INT as a result in milliseconds.
That's why this pull request is created.
I came up with the closest possible regular expression to fix this issue
and not to introduce any other bug.

**Major Revision**
- 0.11.0

7af7c0b7

04 Feb, 2025 1 commit
- Microbenchmark - Add arch support for 10.0 in gemm-flops (#680) · 1d09b111
  Hongtao Zhang authored Feb 03, 2025
```
**Description**
Introduce architecture support for version 10.0 in gemm-flops.
```
  1d09b111
28 Nov, 2024 2 commits

Benchmarks - Add LLaMA-2 Models (#668) · 249e21c1

pdr authored Nov 27, 2024

Added llama benchmark - training and inference in accordance with the
existing pytorch models implementation like gpt2, lstm etc.

- added llama fp8 unit test for better code coverage, to reduce memory
required
- updated transformers version >= 4.28.0 for LLamaConfig
- set tokenizers version <= 0.20.3 to avoid 0.20.4 version
[issues](https://github.com/huggingface/tokenizers/issues/1691

) with
py3.8
- added llama2 to tensorrt
- llama2 tests not added to test_tensorrt_inference_performance.py due
to large memory requirement for worker gpu. tests validated separately
on gh200

---------
Co-authored-by: dpatlolla <dpatlolla@microsoft.com>

249e21c1

Bug Fix - Fix stderr message in gpu-copy benchmark (#673) · 4e6935ab
pdr authored Nov 27, 2024
```
Fix ordering of args in err messages.
```
4e6935ab

22 Nov, 2024 1 commit

Benchmarks: micro benchmarks - add nvbandwidth benchmark (#669) · 7cef624e

Hongtao Zhang authored Nov 21, 2024



**Description**

Add nvbandwidth benchmark.

---------
Co-authored-by: hongtaozhang <hongtaozhang@microsoft.com>

7cef624e

20 Nov, 2024 1 commit

Benchmarks: micro benchmarks - add general CPU bandwidth and latency benchmark (#662) · 9c35e80a

Hongtao Zhang authored Nov 20, 2024



**Description**
Add micro benchmark to measure general CPU bandwidth and latency without 'mlc'.

Test output:
```
{
"cpu-memory-bw-latency/return_code": 0,
"cpu-memory-bw-latency/mem_bandwidth_matrix_numa_0_1_bw": 5388.75021,
"cpu-memory-bw-latency/mem_bandwidth_matrix_numa_0_1_lat": 0.185571786,
"cpu-memory-bw-latency/mem_bandwidth_matrix_numa_1_0_bw": 4634.82028,
"cpu-memory-bw-latency/mem_bandwidth_matrix_numa_1_0_lat": 0.215758096,
}
```

---------
Co-authored-by: hongtaozhang <hongtaozhang@microsoft.com>

9c35e80a

06 Nov, 2024 1 commit

Dockerfile - Add support for arm64 build (#660) · 47949127

pdr authored Nov 06, 2024

Add support for arm64 build:

- Updated dockerfile for arm64 build
- extend cpu stream compilation for neoverse 
- handle onnxruntime-gpu installation
- third party builds filtering based on arch
- disable cuda decode perf build for non x86

47949127

05 Nov, 2024 1 commit

Bug Fix - Fix numa error on grace cpu in gpu-copy (#658) · 59d36f7f

pdr authored Nov 05, 2024

The current GPU Copy BW Performance fails on Nvidia Grace systems. This
is due to the memory only numa node and thus the numa_run_on_node fails
for such nodes and halts completely.

This fix checks for the presence of assigned CPU cores for the numa
node, on checking if it has no cpu cores assigned, it skips that
specific node during the args creation and continues.

59d36f7f

26 Jul, 2024 1 commit
- Benchmarks: Micro benchmarks - add support for NVIDIA L4/L40/L40s GPUs in gemm-flops (#634) · e304cf15
  Yuting Jiang authored Jul 26, 2024
```
**Description**
Add support GPU ARCH 8.9 for NVIDIA L4/L40/L40s GPUs in gemm-flops.
```
  e304cf15
02 Apr, 2024 1 commit
- Benchmarks: Revise Code - Add hipblasLt tuning to dist-inference cpp implementation (#616) · cc89ee59
  Ziyue Yang authored Apr 02, 2024
```
**Description**
Adds hipblasLt tuning to dist-inference cpp implementation.
```
  cc89ee59
08 Jan, 2024 1 commit

Release - SuperBench v0.10.0 (#607) · 2c88db90

Yifan Xiong authored Jan 07, 2024

**Description**

Cherry-pick bug fixes from v0.10.0 to main.

**Major Revisions**

* Benchmarks: Microbenchmark - Support different hipblasLt data types in dist_inference #590
* Benchmarks: Microbenchmark - Support in-place for NCCL/RCCL benchmark #591
* Bug Fix - Fix NUMA Domains Swap Issue in NDv4 Topology File #592
* Benchmarks: Microbenchmark - Add data type option for NCCL and RCCL tests #595
* Benchmarks: Bug Fix - Make metrics of dist-inference-cpp aligned with PyTorch version #596
* CI/CD - Add ndv5 topo file #597
* Benchmarks: Microbenchmark - Improve AMD GPU P2P performance with fine-grained GPU memory #593
* Benchmarks: Build Pipeline - fix nccl and nccl test version to 2.18.3 to resolve hang issue in cuda12.2 docker #599
* Dockerfile - Bug fix for rocm docker build and deploy #598
* Benchmarks: Microbenchmark - Adapt to hipblasLt data type changes #603
* Benchmarks: Micro benchmarks - Update hipblaslt metric unit to tflops #604
* Monitor - U...

2c88db90

11 Dec, 2023 1 commit
- Benchmark: Revision - Fix -O2 option passing in gpu_copy ROCm build (#589) · 2c2096ed
  Ziyue Yang authored Dec 11, 2023
```
**Description**
`add_compile_options` will not work for ROCm build, change it to setting
`CMAKE_CXX_FLAGS`.
```
  2c2096ed
10 Dec, 2023 1 commit
- Benchmarks: Microbenchmark - Add distributed inference benchmark cpp implementation (#586) · 719a427f
  Ziyue Yang authored Dec 11, 2023
```
**Description**
Add distributed inference benchmark cpp implementation.
```
  719a427f
09 Dec, 2023 1 commit

Dockerfile - Upgrade to rocm5.7 dockerfile (#587) · 1f5031bd

Yuting Jiang authored Dec 10, 2023



**Description**
upgrade to rocm5.7 dockerfile.

---------
Co-authored-by: yukirora <yuting.jiang@microsoft.com>

1f5031bd