Commits · 6501bd07b87b148d025e0b7c9cbc8f3bb34ef765 · OpenDAS / tilelang

02 Dec, 2025 4 commits

[Refactor] Update condition for benchmarking in example_gemv.py and simplify... · 6501bd07

Lei Wang authored Dec 02, 2025

[Refactor] Update condition for benchmarking in example_gemv.py and simplify cached library path handling in sparse.py (#1365)

6501bd07

[Debug] Always include line info in NVCC command for improved profiling and mapping (#1364) · d88594a3
Lei Wang authored Dec 02, 2025

d88594a3
[Bugfix] Remove debug print in PyStmtFunctionVisitor (#1363) · f951b924
Lei Wang authored Dec 02, 2025

f951b924

[CI] [pre-commit.ci] autoupdate (#1362) · e37f2eab

pre-commit-ci[bot] authored Dec 02, 2025

updates:
- [github.com/pre-commit/mirrors-clang-format: v21.1.2 → v21.1.6](https://github.com/pre-commit/mirrors-clang-format/compare/v21.1.2...v21.1.6)
- [github.com/astral-sh/ruff-pre-commit: v0.14.3 → v0.14.7](https://github.com/astral-sh/ruff-pre-commit/compare/v0.14.3...v0.14.7

)
Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>

e37f2eab

01 Dec, 2025 5 commits

[Enhancement] Implement dynamic unroll factor in CUDA code generation (#1360) · 388ee7ee

Lei Wang authored Dec 02, 2025

* [Enhancement] Implement dynamic unroll factor in CUDA code generation

This commit introduces support for specifying a dynamic unroll factor in the CUDA code generation. The `unroll_factor` map is added to store unroll factors for loop variables, allowing for more flexible and optimized loop unrolling. Additionally, the `unroll` function is integrated into the loop language, enabling users to define unroll factors directly in their code. This enhancement improves performance by allowing tailored unrolling strategies based on specific loop characteristics.

* lint fix

* [Bugfix] Correct initialization of non-zero counters in custom compress kernel and update TIR registration for gemm_sp_py to use the correct tile operation

388ee7ee

[Bugfix] Update TIR registration for GemmSPPy to use tile operation (#1361) · e547d247
Lei Wang authored Dec 01, 2025

e547d247

[Language] support `T.gemm_sp_v2` on sm80 and sm89 (#1056) · 283a9a00

botbw authored Dec 01, 2025

* [misc] add a cpp side wrapper for gemm_sp_py

* [misc] typing

* [IR] bind GemmSPWarpPolicy

* [chore] add wrapper code

* [IR] fix GemmSPWarpPolicy

* [codegen] apply ptxas instructions

* [intrinsic] add typical (unused) mma layout

* [template] add uint16 debug func

* [intrinsic] add b matrix layout

* [gemm_sp] enable fp16/bf16 on sm8x

* [layout] refactor fp16/bf16 layout

* [gemm_sp] enable int8

* [chore] update test case dtype

* [gemm_sp] enable fp32

* [layout] refactor layouts

* [intrinsic] enable ldmatrix for mat A

* [layout] enable ldsm for matrix b

* [layout] add ldmatrix for fp32 and fp8

* [chore] refine

* [chore] refactor

* [chore] add fp8 efactor

* [chore] refactor

* [chore] add remove negative zero util

* [example] add a custom compress kernel

* [chore] minor update

* [test] refactor gemm_sp test

* [refactor] make metadata layout func

* [example] add option for using cutlass layout

* [doc] add a gemm_sp doc

* [doc] minor polish

* [chore] remove unused

* [bugfix] fix non replicate b case

* [test] refactor

* [chore] add a check

* [bugfix] fix util bug

* [wip] init a new test case for v2

* [chore] minor refactor

* [chore] minor update

* [bugfix] enable 16bit rs

* [language] enable rs

* [language] enable gemm_sp_sr

* [language] enable gemm_sp_rr

* [test] enable more tests

* [tvm] update ffi binding

* [chore] remove print

* [chore] fix benchmark script

* [lint] precommit lint

* [chore] apply feedback

* [test] use arch 8.0

* [chore] rollback ::ordered_metadata for backward compatibility

* [bugfix] fix captialized

* [example] keep gemm_sp on hopper

* [test] fix no fp8 normal kernel

* [test] reduce matmul size to satisfy accum error

* [test] use cal_diff for assertion

* [bugfix] expand float8 type

* [lib] add make_int4 for short type

* [language] add transpose E

* [bugfix] fix wrong var

* [format] format

* [chore] refactor binding

* [chore] fix wrong passing var

283a9a00

[Analysis] Enhance NestedLoopChecker with tile op cases (#1358) · b10ef75f
Chaofan Lin authored Dec 01, 2025
```
* [Analysis] Enhance NestedLoopChecker with tile op cases

* fix tileop issue
```
b10ef75f

[Refactor] Update Fragment Indexing in ParallelOpNode's InferLayout Method (#1359) · 1b42c87b

Lei Wang authored Dec 01, 2025

This commit refines the Fragment creation process in the InferLayout method of ParallelOpNode. It removes the unnecessary forward_index array and utilizes default fragment indexing for consistency with other operations. Additionally, it binds the thread range to enhance comparability across different operations.

1b42c87b

30 Nov, 2025 1 commit

[Bugfix] Fix the jit_kernel issue (#1357) · c6a19fb2

Leon Lu authored Nov 30, 2025



* [Bugfix] Fix the jit_kernel issue

* Update README.md

---------
Co-authored-by: Lei Wang <34334180+LeiWang1999@users.noreply.github.com>

c6a19fb2

28 Nov, 2025 3 commits

[Bugfix] Disable floordiv optimization due to integer overflow risk (#1355) · a4ea7da9

LJC00118 authored Nov 28, 2025

* disable overflow-prone floordiv optimization in lower_intrin.cc

* disable overflow-prone floordiv optimization in lower_intrin.cc

a4ea7da9

[Enhancement] Improve error handling and assertion messages across runtime and... · 17cfeb76

Lei Wang authored Nov 28, 2025

[Enhancement] Improve error handling and assertion messages across runtime and argument binding (#1356)

This commit enhances the error handling mechanisms in the runtime by introducing CPU-safe runtime helpers and refining assertion messages in the CodeGenCHost and ArgBinder. It includes structured packed error messages for various conditions, improving clarity in diagnostics. Additionally, the CMake configuration is updated to always include necessary runtime helpers, ensuring consistent error reporting. The changes aim to provide clearer feedback during runtime errors and improve the overall robustness of the argument binding process.

17cfeb76

[Refactor] Simplify index sign state handling in LegalizeNegativeIndex (#1354) · 36a2b2f3

Lei Wang authored Nov 28, 2025

This commit refines the logic for determining the sign state of indices in the LegalizeNegativeIndex transformation. It prioritizes vector patterns, specifically Ramp and Broadcast nodes, to avoid compile-time lane queries. The handling of scalar indices is also streamlined, ensuring clearer diagnostics when non-negativity cannot be proven. These changes enhance the robustness and clarity of index handling in the transformation pass.

36a2b2f3

27 Nov, 2025 2 commits

[Refactor] Improve assertion handling in CodeGenCHost and ArgBinder (#1352) · 1e92d11c

Lei Wang authored Nov 28, 2025

* [Refactor] Improve assertion handling in CodeGenCHost and ArgBinder

This commit refines the assertion message generation in CodeGenCHost by optimizing the handling of equality checks and reducing buffer size for error messages. Additionally, it enhances the ArgBinder by introducing a nullable guard mechanism for assertions, allowing for more precise error handling when binding arguments. The changes improve the clarity and efficiency of assertion handling across the codebase.

* [Enhancement] Update matmul kernel and optimize argument binding

This commit enhances the matmul kernel by introducing additional tensor parameters and refining the pipeline stages for improved performance. It also updates the argument binding mechanism to include a flag indicating whether buffers are used, enhancing the efficiency of buffer management. Furthermore, the optimization phase in the engine is improved by adding a simplification step, ensuring better performance and clarity in the generated code.

* lint fix

* [Enhancement] Add tensor checks documentation and improve argument binding assertions

This commit introduces a new documentation page for host-side tensor checks, detailing the automatic validations performed by TileLang on kernel arguments. It enhances the ArgBinder by adding assertions for non-null pointers when arguments are used, improving error handling. Additionally, the optimization phase in the engine is updated to include a simplification step, ensuring better performance and clarity in the generated code.

* [Enhancement] Update .gitignore and refine matmul kernel for improved performance

This commit adds host checks logs to the .gitignore file to prevent unnecessary log files from being tracked. Additionally, it refines the matmul kernel by adjusting pipeline stages, updating tensor parameters, and enhancing argument handling for better performance. The changes also include improved error messages in the argument binding process, ensuring clearer diagnostics for users.

* lint fix

* lint fix

* [Refactor] Simplify tensor_null_test function and remove ptr_null_test

This commit refactors the tensor_null_test function by adding a with_bias parameter and removing the ptr_null_test function, which was previously unused. The run_test function is updated to reflect these changes, streamlining the testing process for tensor operations.

* lint fix

* fix

1e92d11c

Add sparse fine-tuning kernel for deepseek sparse attention to example (#1296) · b8240b7a
Yuxuan Hu authored Nov 27, 2025
```
* [EXAMPLE] add example for dsa sparse finetuning

* [Refactor]
```
b8240b7a

26 Nov, 2025 7 commits

[Enhancement] Add support for k_pack in gemm_mfma (#1344) · 6bae64f6
Gongen-Ali authored Nov 26, 2025
```
* add support for k_pack

* support benchmark on ROCm

* fix format
```
6bae64f6
[Fix] Fix missing `not` rewrite in frontend (#1348) · 4f844000
Kuris authored Nov 26, 2025

4f844000

[Refactor] Enhance CopyNode's IterVar Creation and Range Handling (#1346) · 17718bec

Lei Wang authored Nov 26, 2025

* [Refactor] Enhance CopyNode's IterVar Creation and Range Handling

This commit refines the `MakeIterVars` method in `CopyNode` to select base ranges based on memory scope levels, ensuring that the chosen ranges are not smaller than the original source ranges. Additionally, it updates the Python `copy` function to clarify range handling, including broadcasting logic and extent alignment. These changes improve the robustness and clarity of the copy operation's implementation.

* test fix

17718bec

[Enhancement] add more dtype and fix mma.ws for fp16 for tcgen05 (#1327) · f0c721a4

Yunqian Fan authored Nov 26, 2025

* feat: add fp8 variants; add placeholder for fp6/fp4 in meta

support ld with pack for fp32 dtype

add dump

add tempalte expand

remove unused dtype and change to rebased apis

* fix: when atom-m!=128, enable_ws

* fix: typo in tcgen05 meta; dispatch in gemm sm100

f0c721a4

[Refactor] Phaseout vmap for Tile Operators (#1334) · f5d9da46

Lei Wang authored Nov 26, 2025



* Refactor GEMM and Reduce operations by moving NormalizeToBufferRegion and MakeAccessPtrFromRegion to utils.{h,cc} for better code organization and reuse.

* lint fix

* Refactor region handling by removing the RegionOp and updating NormalizeToBufferRegion to only accept BufferLoad and BufferRegion. This change improves code organization and simplifies the handling of memory regions across various operations.

* fix

* Refactor memory region handling by introducing `tl.region` calls across various operations, including GEMM and fill functions. This change enhances the consistency of region management and improves code organization by utilizing utility functions for buffer region conversions.

* fix

* fix

* test fix

* lint fix

* Refactor GEMM operations to improve memory region handling by replacing `mbarPtr_` with `mbarRegion_` and updating related logic in both C++ and Python implementations. This change enhances the clarity and consistency of buffer region management.

* fix

* lint fix

* fix

* fix

* test fix

* lint fix

* lint fix

* minor fix

* fix

---------
Co-authored-by: Zhiwen Mo <zm125@ic.ac.uk>

f5d9da46

[Feat] Extend LegalizeNegativeIndex to support buffer store stmts (#1339) · fac04006

ConvolutedDog authored Nov 26, 2025

This commit enhances the LegalizeNegativeIndex transformation pass to handle
both buffer load and store operations with negative indices and adds some
test cases.

fac04006

Add unit tests for T.assume (#1341) · f810f976

LJC00118 authored Nov 26, 2025



* Add test for T.assume

* Add unit test for T.assume

* Add unit test for T.assume

* Add unit tests for T.assume

* Remove debug print for kernel source

Remove print statement for kernel source in tests.

* Update test_tilelang_language_assume.py

---------
Co-authored-by: Lei Wang <34334180+LeiWang1999@users.noreply.github.com>

f810f976

25 Nov, 2025 5 commits
- [Language][UX] Semantic check for parallel fragment access (#1338) · e2b10c58
  Chaofan Lin authored Nov 25, 2025
  
  e2b10c58
- [Fix] Fix bug copying from or to local buffer (#1304) (#1324) · 2ae4f1b7
  Kuris authored Nov 25, 2025
```
* [Fix] fix copy from or to local buffer (#1304)

* fix lint error

* minor fix testing script
```
  2ae4f1b7
- [Refactor] Moving `NormalizeToBufferRegion` and `MakeAccessPtrFromRegion` to utils (#1333) · 2f34840f
  Lei Wang authored Nov 25, 2025
```
* Refactor GEMM and Reduce operations by moving NormalizeToBufferRegion and MakeAccessPtrFromRegion to utils.{h,cc} for better code organization and reuse.

* lint fix
```
  2f34840f
- [Refactor] Disable strided buffer load inside tvm (#1301) (#1332) · 71b73e18
  Kuris authored Nov 25, 2025
  
  71b73e18
- [Fix] fix wrong uint narrowing bug in tvm in #1310 (#1320) · b0206854
  Kuris authored Nov 25, 2025
  
  b0206854
24 Nov, 2025 8 commits
- [BugFix] Use BufferRegion in tl.cumsum to infer buffer shape (#1321) · 9dda774a
  Chaofan Lin authored Nov 25, 2025
```
* [BugFix] Use BufferRegion in tl.cumsum to infer buffer shape

* remove debug lines

* remove rubbish

* Fix decorator syntax for atomic_different_memory_orders_program

---------
Co-authored-by: Lei Wang <34334180+LeiWang1999@users.noreply.github.com>
```
  9dda774a
- [Enhancement] Support more dtype in `T.print` (#1329) · c30df2a1
  Wenhao Xie authored Nov 25, 2025
```
* [Enhancement] Support more dtype in `T.print`

* upd

* upd
```
  c30df2a1
- [Feat] Support warp reduce (#1316) · caa6dd3f
  Tong WU authored Nov 24, 2025
```
* [Feat] Support warp reduce

* lint

* add test

* lint
```
  caa6dd3f
- [Release] Allow developer with write permission to trigger wheel release (#1322) · 6c2162a9
  Yichen Yan authored Nov 24, 2025
  
  6c2162a9
- [Installation] Fix building using customized TVM path (#1326) · 01d207fa
  Chaofan Lin authored Nov 24, 2025
  
  01d207fa
- [CI]: Bump pypa/cibuildwheel from 3.2 to 3.3 (#1318) · 2a70fd3f
  dependabot[bot] authored Nov 24, 2025
```
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
```
  2a70fd3f
- [CI]: Bump actions/checkout from 5 to 6 (#1319) · fddcbbd6
  dependabot[bot] authored Nov 24, 2025
```
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
```
  fddcbbd6
- Revert "[WIP] support more dtypes for tcgen05 (#1229)" (#1323) · ca98cc39
  Lei Wang authored Nov 24, 2025
```
This reverts commit 0d101c11

.
Co-authored-by: Zhiwen Mo <zm125@ic.ac.uk>
```
  ca98cc39
23 Nov, 2025 1 commit

[Refactor] Backup Analyzer to get the appropriate arith informations (#1311) · 9f7bac4c

Lei Wang authored Nov 23, 2025

* [Refactor] Update Vectorization Functions to Accept Analyzer Parameter

- Modified `VectorizeLoop` and related functions to accept an `arith::Analyzer` parameter, enhancing their capability to perform analysis during vectorization.
- Updated multiple instances in `copy.cc`, `fill.cc`, `parallel.cc`, and layout inference files to utilize the new analyzer parameter for improved performance and correctness.
- Ensured consistency across vectorization logic by integrating the analyzer into existing workflows, facilitating better optimization opportunities.

* [Fix] Corrected PostOrderVisit call in loop_vectorize.cc

- Updated the PostOrderVisit function to analyze the body of the loop node instead of the node itself, ensuring proper handling of nested loops during vectorization analysis.

* fix

* lint fix

* fix

9f7bac4c

22 Nov, 2025 2 commits

[Bugfix] Fix autotune cache (#1315) · 721baedb
Lei Wang authored Nov 22, 2025

721baedb

Improve memory access safety and `T.assume` handling (#1292) · 470eb74c

LJC00118 authored Nov 22, 2025



* Improve memory access safety and T.assume handling

* Improve memory access safety and T.assume handling

* bugfix

* lint fix

* bugfix

* bugfix

* refactor legalize safe memory access pass

---------
Co-authored-by: Lei Wang <leiwang1999@outlook.com>

470eb74c

21 Nov, 2025 2 commits

[WIP] support more dtypes for tcgen05 (#1229) · 0d101c11

Yunqian Fan authored Nov 21, 2025

support ld with pack for fp32 dtype

add dump

add tempalte expand

remove unused dtype and change to rebased apis

0d101c11

[Fix] Fix frame scope error in T.macro (#1308) · bf90a5f5

Kuris authored Nov 21, 2025



* [Fix] Fix #1307 by adding macro inside function

* fix lint error

* add comments and fix lint error

* Remove debug print from enter_frame method

Removed debug print statement from enter_frame method.

---------
Co-authored-by: Lei Wang <34334180+LeiWang1999@users.noreply.github.com>

bf90a5f5