[Docs] Improve docstring formatting for `FusedMoEParallelConfig.make` (#21117)

Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>

[Docs] Improve docstring formatting for `FusedMoEParallelConfig.make` (#21117)
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
fe8a2c54 · Harry Mellor · GitHub · 4ef00b5c · fe8a2c54
Unverified Commit fe8a2c54 authored Jul 17, 2025 by Harry Mellor Committed by GitHub Jul 17, 2025
Show whitespace changes
Inline Side-by-side

Showing with 34 additions and 28 deletions

vllm/model_executor/layers/fused_moe/config.py vllm/model_executor/layers/fused_moe/config.py +34 -28

No files found.
--- a/vllm/model_executor/layers/fused_moe/config.py
+++ b/vllm/model_executor/layers/fused_moe/config.py
@@ -192,40 +192,43 @@ class FusedMoEParallelConfig:
    def make(tp_size_: int, dp_size_: int,
             vllm_parallel_config: ParallelConfig) -> "FusedMoEParallelConfig":
        """
-        Determine MoE parallel configuration. Based on the input tp_size_,
+        Determine MoE parallel configuration. Based on the input `tp_size_`,
-        dp_size_, ep_size_ and vllm's parallel config, determine what
+        `dp_size_` and vllm's parallel config, determine what
        level's of parallelism to use in the fused moe layer.
        Args:
-            tp_size_ (int): tp_size passed into the FusedMoE constructor.
+            tp_size_ (int): `tp_size` passed into the FusedMoE constructor.
-            dp_size_ (int): dp_size passed into the FusedMoE constructor.
+            dp_size_ (int): `dp_size` passed into the FusedMoE constructor.
-            ep_size_ (int): ep_size passed into the FusedMoE constructor.
+            vllm_parallel_config (ParallelConfig): vLLM's parallel config
-            vllm_parallel_config (ParallelConfig): vllm's parallel config
+                object which contains the `enable_expert_parallel` flag.
-            object.
        Examples:
-        When there is no parallelism requested, i.e. tp_size_ = dp_size_ = 1,
+            When there is no parallelism requested,
-        we simply return the sizes unaltered and the ranks set to 0.
+            i.e. `tp_size_` = `dp_size_` = 1, we simply return the sizes
+            unaltered and the ranks set to 0.
-        Expert Parallelism is considered only when either dp_size_ or tp_size_
+            Expert Parallelism is considered only when either `dp_size_` or
-        is non trivial.
+            `tp_size_` is non trivial.
            When TP = 2, DP = 1 and EP = False, the configuration on different
-        devices,
+            devices:
            - device 0 : TP = {2, 0} DP = {1, 0} EP = {1, 0} //
                legend : {size, rank}
            - device 1 : TP = {2, 1} DP = {1, 0} EP = {1, 0}
            - Comment : Tensors are sharded across 2 devices.
            When TP = 1, DP = 2 and EP = False, the configuration on different
-        devices,
+                devices:
            - device 0 : TP = {2, 0} DP = {2, 0} EP = {1, 0}
            - device 1 : TP = {2, 1} DP = {2, 1} EP = {1, 0}
            - Comment: There are 2 engine instances and the tensors are sharded
                across 2 decvices.
            When TP = 2, DP = 2 and EP = False, the configuration on different
-        devices,
+                devices:
            - device 0: TP = {4, 0} DP = {2, 0} EP = {1, 0}
            - device 1: TP = {4, 1} DP = {2, 0} EP = {1, 0}
            - device 2: TP = {4, 2} DP = {2, 1} EP = {1, 0}
@@ -234,20 +237,23 @@ class FusedMoEParallelConfig:
                across 4 devices.
            When, TP = 2, DP = 1 and EP = True, the configuration on different
-        devices,
+                devices:
            - device 0: TP = {1, 0} DP = {1, 0} EP = {2, 0}
            - device 1: TP = {1, 0} DP = {1, 0} EP = {2, 1}
            - Comment: The experts are split between the 2 devices.
            When, TP = 1, DP = 2 and EP = True, the configuration on different
-        devices,
+                devices:
            - device 0: TP = {1, 0} DP = {2, 0} EP = {2, 0}
            - device 1: TP = {1, 0} DP = {2, 1} EP = {2, 1}
            - Comment: There are 2 engine instances and the experts are split
                between the 2 devices.
            When TP = 2, DP = 2 and EP = True, the configuration on different
-        devices,
+                devices:
            - device 0: TP = {1, 0} DP = {2, 0} EP = {4, 0}
            - device 1: TP = {1, 0} DP = {2, 0} EP = {4, 1}
            - device 2: TP = {1, 0} DP = {2, 1} EP = {4, 2}