Improve docs and fix the broken links (#1875)

d1b31b06 · Lianmin Zheng · GitHub · d59a4782 · d1b31b06 · d1b31b06
Unverified Commit d1b31b06 authored Nov 01, 2024 by Lianmin Zheng Committed by GitHub Nov 01, 2024
8 changed files
--- a/docs/backend/backend.md
+++ b/docs/backend/backend.md
@@ -84,7 +84,8 @@ python -m sglang.launch_server --model-path meta-llama/Meta-Llama-3-8B-Instruct
 - To enable torchao quantization, add `--torchao-config int4wo-128`. It supports various quantization strategies.
 - To enable fp8 weight quantization, add `--quantization fp8` on a fp16 checkpoint or directly load a fp8 checkpoint without specifying any arguments.
 - To enable fp8 kv cache quantization, add `--kv-cache-dtype fp8_e5m2`.
- If the model does not have a chat template in the Hugging Face tokenizer, you can specify a [custom chat template](https://sgl-project.github.io/custom_chat_template.html).
+- If the model does not have a chat template in the Hugging Face tokenizer, you can specify a [custom chat template](https://sgl-project.github.io/references/custom_chat_template.html).
+
 - To run tensor parallelism on multiple nodes, add `--nnodes 2`. If you have two nodes with two GPUs on each node and want to run TP=4, let `sgl-dev-0` be the hostname of the first node and `50000` be an available port, you can use the following commands. If you meet deadlock, please try to add `--disable-cuda-graph`
 ```
 # Node 0

--- a/docs/backend/openai_api.ipynb
+++ b/docs/backend/openai_api.ipynb
--- a/docs/backend/vision_language_model.ipynb
+++ b/docs/backend/vision_language_model.ipynb
--- a/docs/index.rst
+++ b/docs/index.rst
@@ -23,8 +23,8 @@ The core features include:
   :maxdepth: 1
   :caption: Backend Tutorial

-   backend/openai_api.ipynb
-   backend/vision_language_model.ipynb
+   backend/openai_api_completions.ipynb
+   backend/openai_api_vision.ipynb
   backend/backend.md


@@ -46,5 +46,5 @@ The core features include:
   references/choices_methods.md
   references/benchmark_and_profiling.md
   references/troubleshooting.md
-   references/embedding_model.ipynb
+   references/custom_chat_template.md
   references/learn_more.md
--- a/docs/references/custom_chat_template.md
+++ b/docs/references/custom_chat_template.md
+.. _custom-chat-template:
+
 # Custom Chat Template in SGLang Runtime

 **NOTE**: There are two chat template systems in SGLang project. This document is about setting a custom chat template for the OpenAI-compatible API server (defined at [conversation.py](https://github.com/sgl-project/sglang/blob/main/python/sglang/srt/conversation.py)). It is NOT related to the chat template used in the SGLang language frontend (defined at [chat_template.py](https://github.com/sgl-project/sglang/blob/main/python/sglang/lang/chat_template.py)).

--- a/docs/references/sampling_params.md
+++ b/docs/references/sampling_params.md
+.. _sampling-parameters:
+
 # Sampling Parameters in SGLang Runtime
 This doc describes the sampling parameters of the SGLang Runtime.
 It is the low-level endpoint of the runtime.

--- a/docs/starts/send_request.ipynb
+++ b/docs/starts/send_request.ipynb
--- a/test/srt/test_vision_openai_server.py
+++ b/test/srt/test_vision_openai_server.py
@@ -132,7 +132,7 @@ class TestOpenAIVisionServer(unittest.TestCase):
        assert response.usage.completion_tokens > 0
        assert response.usage.total_tokens > 0

-    def test_mult_images_chat_completion(self):
+    def test_multi_images_chat_completion(self):
        client = openai.Client(api_key=self.api_key, base_url=self.base_url)

        response = client.chat.completions.create(