[doc][faq] add warning to download models for every nodes (#5783)

c2462129 · youkaichao · GitHub · edd5fe5f · c2462129
Unverified Commit c2462129 authored Jun 24, 2024 by youkaichao Committed by GitHub Jun 24, 2024
Hide whitespace changes
Inline Side-by-side

Showing with 4 additions and 1 deletion

docs/source/serving/distributed_serving.rst docs/source/serving/distributed_serving.rst +4 -1

No files found.
--- a/docs/source/serving/distributed_serving.rst
+++ b/docs/source/serving/distributed_serving.rst
@@ -35,4 +35,7 @@ To scale vLLM beyond a single machine, install and start a `Ray runtime <https:/
    $ # On worker nodes
    $ ray start --address=<ray-head-address>

-After that, you can run inference and serving on multiple machines by launching the vLLM process on the head node by setting :code:`tensor_parallel_size` to the number of GPUs to be the total number of GPUs across all machines.
\ No newline at end of file
+After that, you can run inference and serving on multiple machines by launching the vLLM process on the head node by setting :code:`tensor_parallel_size` to the number of GPUs to be the total number of GPUs across all machines.
+
+.. warning::
+    Please make sure you downloaded the model to all the nodes, or the model is downloaded to some distributed file system that is accessible by all nodes.