Release v0.3.1 (#1430)

e79f6cd7 · Lianmin Zheng · GitHub · 9ba1f097 · e79f6cd7 · e79f6cd7
Unverified Commit e79f6cd7 authored Sep 15, 2024 by Lianmin Zheng Committed by GitHub Sep 15, 2024
Showing with 5 additions and 5 deletions

README.md README.md +2 -2

docs/en/install.md docs/en/install.md +1 -1

python/pyproject.toml python/pyproject.toml +1 -1

python/sglang/version.py python/sglang/version.py +1 -1

No files found.
--- a/README.md
+++ b/README.md
@@ -60,7 +60,7 @@ pip install flashinfer -i https://flashinfer.ai/whl/cu121/torch2.4/
 ### Method 2: From source
 ```
 # Use the last release branch
-git clone -b v0.3.0 https://github.com/sgl-project/sglang.git
+git clone -b v0.3.1 https://github.com/sgl-project/sglang.git
 cd sglang
 pip install --upgrade pip
@@ -139,7 +139,7 @@ sky status --endpoint 30000 sglang
 ### Common Notes
- [FlashInfer](https://github.com/flashinfer-ai/flashinfer) is the default attention kernel backend. It only supports sm75 and above. If you encounter any FlashInfer-related issues on sm75+ devices (e.g., T4, A10, A100, L4, L40S, H100), please disable it by adding `--disable-flashinfer --disable-flashinfer-sampling` and open an issue on GitHub.
+- [FlashInfer](https://github.com/flashinfer-ai/flashinfer) is the default attention kernel backend. It only supports sm75 and above. If you encounter any FlashInfer-related issues on sm75+ devices (e.g., T4, A10, A100, L4, L40S, H100), please switch to other kernels by adding `--attention-backend triton --sampling-backend pytorch` and open an issue on GitHub.
 - If you only need to use the OpenAI backend, you can avoid installing other dependencies by using `pip install "sglang[openai]"`.
 ## Backend: SGLang Runtime (SRT)

--- a/docs/en/install.md
+++ b/docs/en/install.md
@@ -92,5 +92,5 @@ sky status --endpoint 30000 sglang
 </details>
 ### Common Notes
- [FlashInfer](https://github.com/flashinfer-ai/flashinfer) is the default attention kernel backend. It only supports sm75 and above. If you encounter any FlashInfer-related issues on sm75+ devices (e.g., T4, A10, A100, L4, L40S, H100), please disable it by adding `--disable-flashinfer --disable-flashinfer-sampling` and open an issue on GitHub.
+- [FlashInfer](https://github.com/flashinfer-ai/flashinfer) is the default attention kernel backend. It only supports sm75 and above. If you encounter any FlashInfer-related issues on sm75+ devices (e.g., T4, A10, A100, L4, L40S, H100), please switch to other kernels by adding `--attention-backend triton --sampling-backend pytorch` and open an issue on GitHub.
 - If you only need to use the OpenAI backend, you can avoid installing other dependencies by using `pip install "sglang[openai]"`.
--- a/python/pyproject.toml
+++ b/python/pyproject.toml
@@ -4,7 +4,7 @@ build-backend = "setuptools.build_meta"
 [project]
 name = "sglang"
-version = "0.3.0"
+version = "0.3.1"
 description = "SGLang is yet another fast serving framework for large language models and vision language models."
 readme = "README.md"
 requires-python = ">=3.8"

--- a/python/sglang/version.py
+++ b/python/sglang/version.py
-__version__ = "0.3.0"
+__version__ = "0.3.1"