Commits · f0c721a467ed0e535b160e3f7e76709faa77cf57 · OpenDAS / tilelang

26 Nov, 2025 4 commits

[Enhancement] add more dtype and fix mma.ws for fp16 for tcgen05 (#1327) · f0c721a4

Yunqian Fan authored Nov 26, 2025

* feat: add fp8 variants; add placeholder for fp6/fp4 in meta

support ld with pack for fp32 dtype

add dump

add tempalte expand

remove unused dtype and change to rebased apis

* fix: when atom-m!=128, enable_ws

* fix: typo in tcgen05 meta; dispatch in gemm sm100

f0c721a4

[Refactor] Phaseout vmap for Tile Operators (#1334) · f5d9da46

Lei Wang authored Nov 26, 2025



* Refactor GEMM and Reduce operations by moving NormalizeToBufferRegion and MakeAccessPtrFromRegion to utils.{h,cc} for better code organization and reuse.

* lint fix

* Refactor region handling by removing the RegionOp and updating NormalizeToBufferRegion to only accept BufferLoad and BufferRegion. This change improves code organization and simplifies the handling of memory regions across various operations.

* fix

* Refactor memory region handling by introducing `tl.region` calls across various operations, including GEMM and fill functions. This change enhances the consistency of region management and improves code organization by utilizing utility functions for buffer region conversions.

* fix

* fix

* test fix

* lint fix

* Refactor GEMM operations to improve memory region handling by replacing `mbarPtr_` with `mbarRegion_` and updating related logic in both C++ and Python implementations. This change enhances the clarity and consistency of buffer region management.

* fix

* lint fix

* fix

* fix

* test fix

* lint fix

* lint fix

* minor fix

* fix

---------
Co-authored-by: Zhiwen Mo <zm125@ic.ac.uk>

f5d9da46

[Feat] Extend LegalizeNegativeIndex to support buffer store stmts (#1339) · fac04006

ConvolutedDog authored Nov 26, 2025

This commit enhances the LegalizeNegativeIndex transformation pass to handle
both buffer load and store operations with negative indices and adds some
test cases.

fac04006

Add unit tests for T.assume (#1341) · f810f976

LJC00118 authored Nov 26, 2025



* Add test for T.assume

* Add unit test for T.assume

* Add unit test for T.assume

* Add unit tests for T.assume

* Remove debug print for kernel source

Remove print statement for kernel source in tests.

* Update test_tilelang_language_assume.py

---------
Co-authored-by: Lei Wang <34334180+LeiWang1999@users.noreply.github.com>

f810f976

25 Nov, 2025 5 commits
- [Language][UX] Semantic check for parallel fragment access (#1338) · e2b10c58
  Chaofan Lin authored Nov 25, 2025
  
  e2b10c58
- [Fix] Fix bug copying from or to local buffer (#1304) (#1324) · 2ae4f1b7
  Kuris authored Nov 25, 2025
```
* [Fix] fix copy from or to local buffer (#1304)

* fix lint error

* minor fix testing script
```
  2ae4f1b7
- [Refactor] Moving `NormalizeToBufferRegion` and `MakeAccessPtrFromRegion` to utils (#1333) · 2f34840f
  Lei Wang authored Nov 25, 2025
```
* Refactor GEMM and Reduce operations by moving NormalizeToBufferRegion and MakeAccessPtrFromRegion to utils.{h,cc} for better code organization and reuse.

* lint fix
```
  2f34840f
- [Refactor] Disable strided buffer load inside tvm (#1301) (#1332) · 71b73e18
  Kuris authored Nov 25, 2025
  
  71b73e18
- [Fix] fix wrong uint narrowing bug in tvm in #1310 (#1320) · b0206854
  Kuris authored Nov 25, 2025
  
  b0206854
24 Nov, 2025 8 commits
- [BugFix] Use BufferRegion in tl.cumsum to infer buffer shape (#1321) · 9dda774a
  Chaofan Lin authored Nov 25, 2025
```
* [BugFix] Use BufferRegion in tl.cumsum to infer buffer shape

* remove debug lines

* remove rubbish

* Fix decorator syntax for atomic_different_memory_orders_program

---------
Co-authored-by: Lei Wang <34334180+LeiWang1999@users.noreply.github.com>
```
  9dda774a
- [Enhancement] Support more dtype in `T.print` (#1329) · c30df2a1
  Wenhao Xie authored Nov 25, 2025
```
* [Enhancement] Support more dtype in `T.print`

* upd

* upd
```
  c30df2a1
- [Feat] Support warp reduce (#1316) · caa6dd3f
  Tong WU authored Nov 24, 2025
```
* [Feat] Support warp reduce

* lint

* add test

* lint
```
  caa6dd3f
- [Release] Allow developer with write permission to trigger wheel release (#1322) · 6c2162a9
  Yichen Yan authored Nov 24, 2025
  
  6c2162a9
- [Installation] Fix building using customized TVM path (#1326) · 01d207fa
  Chaofan Lin authored Nov 24, 2025
  
  01d207fa
- [CI]: Bump pypa/cibuildwheel from 3.2 to 3.3 (#1318) · 2a70fd3f
  dependabot[bot] authored Nov 24, 2025
```
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
```
  2a70fd3f
- [CI]: Bump actions/checkout from 5 to 6 (#1319) · fddcbbd6
  dependabot[bot] authored Nov 24, 2025
```
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
```
  fddcbbd6
- Revert "[WIP] support more dtypes for tcgen05 (#1229)" (#1323) · ca98cc39
  Lei Wang authored Nov 24, 2025
```
This reverts commit 0d101c11

.
Co-authored-by: Zhiwen Mo <zm125@ic.ac.uk>
```
  ca98cc39
23 Nov, 2025 1 commit

[Refactor] Backup Analyzer to get the appropriate arith informations (#1311) · 9f7bac4c

Lei Wang authored Nov 23, 2025

* [Refactor] Update Vectorization Functions to Accept Analyzer Parameter

- Modified `VectorizeLoop` and related functions to accept an `arith::Analyzer` parameter, enhancing their capability to perform analysis during vectorization.
- Updated multiple instances in `copy.cc`, `fill.cc`, `parallel.cc`, and layout inference files to utilize the new analyzer parameter for improved performance and correctness.
- Ensured consistency across vectorization logic by integrating the analyzer into existing workflows, facilitating better optimization opportunities.

* [Fix] Corrected PostOrderVisit call in loop_vectorize.cc

- Updated the PostOrderVisit function to analyze the body of the loop node instead of the node itself, ensuring proper handling of nested loops during vectorization analysis.

* fix

* lint fix

* fix

9f7bac4c

22 Nov, 2025 2 commits

[Bugfix] Fix autotune cache (#1315) · 721baedb
Lei Wang authored Nov 22, 2025

721baedb

Improve memory access safety and `T.assume` handling (#1292) · 470eb74c

LJC00118 authored Nov 22, 2025



* Improve memory access safety and T.assume handling

* Improve memory access safety and T.assume handling

* bugfix

* lint fix

* bugfix

* bugfix

* refactor legalize safe memory access pass

---------
Co-authored-by: Lei Wang <leiwang1999@outlook.com>

470eb74c

21 Nov, 2025 4 commits

[WIP] support more dtypes for tcgen05 (#1229) · 0d101c11

Yunqian Fan authored Nov 21, 2025

support ld with pack for fp32 dtype

add dump

add tempalte expand

remove unused dtype and change to rebased apis

0d101c11

[Fix] Fix frame scope error in T.macro (#1308) · bf90a5f5

Kuris authored Nov 21, 2025



* [Fix] Fix #1307 by adding macro inside function

* fix lint error

* add comments and fix lint error

* Remove debug print from enter_frame method

Removed debug print statement from enter_frame method.

---------
Co-authored-by: Lei Wang <34334180+LeiWang1999@users.noreply.github.com>

bf90a5f5

[Bugfix] Fallback to the old AtomicAdd implementation for legacy architectures (#1306) · 17bbc0ca
Lei Wang authored Nov 21, 2025

17bbc0ca

[Fix] Remove unused let_bindings_ in CodeGenC to fix #1300 (#1305) · 2426090f

Kuris authored Nov 21, 2025

* [Feat] add missing support of uint32x2

* [Feat] Add `T.Ref` annotation and tests

* fix lint error

* minor update for error message on twice decl

* Remove unused let_bindings_ in CodeGenC to fix #1300

2426090f

20 Nov, 2025 4 commits
- [Enhancement] Shared Memory Size Can be Dynamic (#1294) · d4b6d094
  Lei Wang authored Nov 20, 2025
```
* bugfix

* lint fix

* test

* lint fix

* increate procs

* recover
```
  d4b6d094
- [Feat] add support for passing reference in T.Var annotation (#1291) · dd7fdb8e
  Kuris authored Nov 20, 2025
  
  dd7fdb8e
- [Feat] Add support for using `T.Tensor(n * 2 + 1)` in function annotation (#1285) · bccb6485
  Kuris authored Nov 20, 2025
```
* [Feature] Add support for A: T.Tensor(n + 1) and A: T.Tensor(2*n)

* issue fix

* fix

* fix

* decreate nproc for debugging

---------
Co-authored-by: Lei Wang <leiwang1999@outlook.com>
```
  bccb6485
- [Compatibility] Support CUDA 11.3 (#1290) · bef7e52e
  Lei Wang authored Nov 20, 2025
  
  bef7e52e
19 Nov, 2025 5 commits

[Language][UX] Nested loop checker in pre-lowering stage (#1288) · 9e67b861
Chaofan Lin authored Nov 20, 2025
```
* [Language][UX] Nested loop checker in pre-lowering stage

* rename

* comment

* address comments
```
9e67b861
Fix the bug in issue #1266 (#1284) · 49f35393
liu yuhao authored Nov 19, 2025
```
Co-authored-by: cheeryBloosm <liu_yu_hao@126.com>
```
49f35393

[Enhancement] Enhance CUDA compilation by integrating pass context configuration (#1283) · 551ac60d

Lei Wang authored Nov 19, 2025

- Updated the `tilelang_callback_cuda_compile` function to accept a `pass_config` parameter, allowing for more flexible compilation options.
- Introduced handling for fast math and PTXAS options based on the provided pass configuration.
- Modified the CUDA build process in `rt_mod_cuda.cc` to utilize the current pass context, improving the integration of compilation settings.
- Refactored NVCC command construction to use a dedicated function for better clarity and maintainability.

551ac60d

[Fix] Fix memory leak bug (#1281) · cd681e63

Kuris authored Nov 19, 2025

* add typing stub for tir.ir

* remove idents

* minor update

* [Refactor] add numpy conversion for dtype

* fix lint error

* remove unused np.float_ in dtype conversion

* fix type in np.int_

* fix typo

* minor fix

* remove debug files

* fix memory leak bug

* fix lint error

* add comments

* fix lint error

* remove duplicated, because tilelang doesn't dependent deprecated

cd681e63

[Bugfix] Supply missing `T.print` for bool type (#1279) · 4c8b9ada
Lei Wang authored Nov 19, 2025
```
* fix for bool dtype

* lint fix

* fix

* ci fix
```
4c8b9ada

18 Nov, 2025 7 commits

[FFI] Use tvm ffi as the default execution backend (#1259) · 74da3696

Lei Wang authored Nov 18, 2025

* [Refactor] Update FFI type handling and simplify argument management

* Refactored FFI type definitions in runtime and code generation files to use `TVMFFIAny` instead of `TVMValue`, enhancing type clarity.
* Updated function registration in `runtime.cc` to utilize canonical names for better consistency.
* Simplified argument handling in the `simplify` transformation, ensuring unused buffer parameters are removed only when simplification is enabled.
* Adjusted autotuner and profiler parameters to standardize the execution backend to `tvm_ffi`, improving clarity in backend selection.
* Removed obsolete `adapt_torch2tvm` function from tensor utilities to streamline the codebase and reduce complexity.

* [Update] Sync TVM submodule and enhance kernel source handling

* Updated the TVM submodule to commit cdc2aced, ensuring compatibility with recent changes.
* Added functionality to print kernel source in `example_blocksparse_gemm.py` for better debugging.
* Commented out the main execution call in test files to prevent unintended execution during testing.
* Introduced `tilelang.disable_cache()` in various test files to streamline testing and avoid cache-related issues.
* Refactored kernel source retrieval methods to improve clarity and consistency across different execution backends.

* [Refactor] Clean up imports and improve code formatting

* Removed unused import of `tilelang.testing` in `test_example_blocksparse_gemm.py` to streamline the code.
* Reformatted several lines in `arg_binder.cc`, `make_packed_api.cc`, `tvm_ffi.py`, and `adapter.py` for improved readability and consistency.
* Updated comments and spacing in `tvm_ffi.py` to enhance clarity without altering functionality.

* Update execution backend options and improve resolution logic

- Changed default execution backend from "cython" to "auto" in multiple locations to allow automatic selection based on the target.
- Expanded the list of supported execution backends to include "torch" and "nvrtc" across various classes and functions.
- Enhanced backend resolution logic in `KernelCache` and `AutoTuner` to ensure appropriate backend selection based on the target.
- Updated documentation to reflect changes in execution backend options and their defaults.

* lint fix

* fix

* Enhance argument handling in CUDA and HIP runtime modules

- Updated `ExtractFuncInfo` in `rt_mod_cuda.cc` and `rt_mod_hip.cc` to map boolean argument types to int32, ensuring compatibility with device runtime.
- Refactored `BindDLTensor` in `arg_binder.cc` to improve null handling and validation checks for DLTensor parameters, utilizing expression-level guards to prevent dereferencing null pointers.
- Enhanced error checking for buffer shape, strides, and data fields, ensuring robust handling of optional inputs and maintaining consistency across various checks.

* lint fix

* minor fix

* fix

* recover check

* Refactor argument binding and validation in `arg_binder.cc`

- Improved null handling and validation checks in `BindDLTensor`, ensuring safe dereferencing of pointers.
- Enhanced consistency checks for buffer shape, strides, and data fields, utilizing expression-level guards.
- Updated `MakePackedAPI` to maintain code clarity and consistency in argument handling.
- Minor adjustments in test files to streamline kernel execution and improve readability.

* lint fix

* stride fix

* minor fix

* fix

* lint fix

* Add CUDA stream access policy window helpers and integrate with L2 persistent cache management

- Introduced functions to set and reset the CUDA stream access policy window, allowing for better control over L2 cache usage.
- Updated runtime files to include new FFI packed functions for managing stream attributes.
- Modified lower_hopper_intrin to incorporate prologue and epilogue statements for L2 cache setup and teardown.
- Enhanced tests to verify the inclusion of new FFI calls in the generated kernel source.

* check with symbolic

* support null ptr

* Update CMakeLists and lower.py for code generation and subproject status

- Added `codegen_c_host.cc` to the list of source files in CMakeLists.txt for improved code generation support.
- Updated the function call in `lower.py` to use `target.build.tilelang_c` for C target host code generation, enhancing compatibility.
- Marked the TVM subproject as dirty to indicate local modifications.

* lint fix

* Update comments for clarity in quickstart.py

74da3696

[Language] Add shape check in `T.view/reshape` (#1277) · 921b96a3
Chaofan Lin authored Nov 18, 2025
```
* [Language] Add shape check in T.view/reshape

* address comments
```
921b96a3
[Bugfix] Minor fix for some cases (#1278) · 1b0efb65
Lei Wang authored Nov 18, 2025

1b0efb65

Bug fix for Gated Delta Net benchmark script (#1267) · 0f980f15

Jay Zhuang authored Nov 18, 2025



* fix argument order for fla chunk_gated_delta_rule_fwd_h

* explicit import assert_similar from utils

* rename utils module to avoid name clash

* set store_final_state and save_new_value to True

* fix

---------
Co-authored-by: LeiWang1999 <leiwang1999@outlook.com>

0f980f15

Fix various issues under `int64_t` static and dynamic shape. (#1218) · 49c85715

Elevator14B authored Nov 18, 2025



* Fix various issues under int64_t static and dynamic shape.

* Resolve reviewed issues.

* Add unit test.

* fix

---------
Co-authored-by: LeiWang1999 <leiwang1999@outlook.com>

49c85715

[BugFix] Adding extra parameters into autotune hashkey (#1274) · e805f8e5
Chaofan Lin authored Nov 18, 2025
```
* [BugFix] Adding extra parameters into autotune hashkey

* lint

* None check

* check serializable
```
e805f8e5
[Minor] Remove from __future__ import annotations for python 3.8 (#1273) · b1922518
Yichen Yan authored Nov 18, 2025

b1922518