README_ORIGIN.md 9.35 KB
Newer Older
zhuwenwen's avatar
zhuwenwen committed
1
2
3
4
5
6
7
8
9
10
11
12
<p align="center">
  <picture>
    <source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/vllm-project/vllm/main/docs/source/assets/logos/vllm-logo-text-dark.png">
    <img alt="vLLM" src="https://raw.githubusercontent.com/vllm-project/vllm/main/docs/source/assets/logos/vllm-logo-text-light.png" width=55%>
  </picture>
</p>

<h3 align="center">
Easy, fast, and cheap LLM serving for everyone
</h3>

<p align="center">
zhuwenwen's avatar
zhuwenwen committed
13
| <a href="https://docs.vllm.ai"><b>Documentation</b></a> | <a href="https://vllm.ai"><b>Blog</b></a> | <a href="https://arxiv.org/abs/2309.06180"><b>Paper</b></a> | <a href="https://x.com/vllm_project"><b>Twitter/X</b></a> | <a href="https://slack.vllm.ai"><b>Developer Slack</b></a> |
zhuwenwen's avatar
zhuwenwen committed
14
15
</p>

zhuwenwen's avatar
zhuwenwen committed
16
17
---

zhuwenwen's avatar
zhuwenwen committed
18
19
20
21
We are excited to invite you to our Menlo Park meetup with Meta, evening of Thursday, February 27! Meta engineers will discuss the improvements on top of vLLM, and vLLM contributors will share updates from the v0.7.x series of releases. [Register Now](https://lu.ma/h7g3kuj9)

---

zhuwenwen's avatar
zhuwenwen committed
22
*Latest News* 🔥
zhuwenwen's avatar
zhuwenwen committed
23

zhuwenwen's avatar
zhuwenwen committed
24
- [2025/01] We are excited to announce the alpha release of vLLM V1: A major architectural upgrade with 1.7x speedup! Clean code, optimized execution loop, zero-overhead prefix caching, enhanced multimodal support, and more. Please check out our blog post [here](https://blog.vllm.ai/2025/01/27/v1-alpha-release.html).
zhuwenwen's avatar
zhuwenwen committed
25
- [2025/01] We hosted [the eighth vLLM meetup](https://lu.ma/zep56hui) with Google Cloud! Please find the meetup slides from vLLM team [here](https://docs.google.com/presentation/d/1epVkt4Zu8Jz_S5OhEHPc798emsYh2BwYfRuDDVEF7u4/edit?usp=sharing), and Google Cloud team [here](https://drive.google.com/file/d/1h24pHewANyRL11xy5dXUbvRC9F9Kkjix/view?usp=sharing).
zhuwenwen's avatar
zhuwenwen committed
26
27
28
29
- [2024/12] vLLM joins [pytorch ecosystem](https://pytorch.org/blog/vllm-joins-pytorch)! Easy, Fast, and Cheap LLM Serving for Everyone!
- [2024/11] We hosted [the seventh vLLM meetup](https://lu.ma/h0qvrajz) with Snowflake! Please find the meetup slides from vLLM team [here](https://docs.google.com/presentation/d/1e3CxQBV3JsfGp30SwyvS3eM_tW-ghOhJ9PAJGK6KR54/edit?usp=sharing), and Snowflake team [here](https://docs.google.com/presentation/d/1qF3RkDAbOULwz9WK5TOltt2fE9t6uIc_hVNLFAaQX6A/edit?usp=sharing).
- [2024/10] We have just created a developer slack ([slack.vllm.ai](https://slack.vllm.ai)) focusing on coordinating contributions and discussing features. Please feel free to join us there!
- [2024/10] Ray Summit 2024 held a special track for vLLM! Please find the opening talk slides from the vLLM team [here](https://docs.google.com/presentation/d/1B_KQxpHBTRa_mDF-tR6i8rWdOU5QoTZNcEg2MKZxEHM/edit?usp=sharing). Learn more from the [talks](https://www.youtube.com/playlist?list=PLzTswPQNepXl6AQwifuwUImLPFRVpksjR) from other vLLM contributors and users!
zhuwenwen's avatar
zhuwenwen committed
30
- [2024/09] We hosted [the sixth vLLM meetup](https://lu.ma/87q3nvnh) with NVIDIA! Please find the meetup slides [here](https://docs.google.com/presentation/d/1wrLGwytQfaOTd5wCGSPNhoaW3nq0E-9wqyP7ny93xRs/edit?usp=sharing).
31
- [2024/07] We hosted [the fifth vLLM meetup](https://lu.ma/lp0gyjqr) with AWS! Please find the meetup slides [here](https://docs.google.com/presentation/d/1RgUD8aCfcHocghoP3zmXzck9vX3RCI9yfUAB2Bbcl4Y/edit?usp=sharing).
32
33
- [2024/07] In partnership with Meta, vLLM officially supports Llama 3.1 with FP8 quantization and pipeline parallelism! Please check out our blog post [here](https://blog.vllm.ai/2024/07/23/llama31.html).
- [2024/06] We hosted [the fourth vLLM meetup](https://lu.ma/agivllm) with Cloudflare and BentoML! Please find the meetup slides [here](https://docs.google.com/presentation/d/1iJ8o7V2bQEi0BFEljLTwc5G1S10_Rhv3beed5oB0NJ4/edit?usp=sharing).
zhuwenwen's avatar
zhuwenwen committed
34
- [2024/04] We hosted [the third vLLM meetup](https://robloxandvllmmeetup2024.splashthat.com/) with Roblox! Please find the meetup slides [here](https://docs.google.com/presentation/d/1A--47JAK4BJ39t954HyTkvtfwn0fkqtsL8NGFuslReM/edit?usp=sharing).
35
36
- [2024/01] We hosted [the second vLLM meetup](https://lu.ma/ygxbpzhl) with IBM! Please find the meetup slides [here](https://docs.google.com/presentation/d/12mI2sKABnUw5RBWXDYY-HtHth4iMSNcEoQ10jDQbxgA/edit?usp=sharing).
- [2023/10] We hosted [the first vLLM meetup](https://lu.ma/first-vllm-meetup) with a16z! Please find the meetup slides [here](https://docs.google.com/presentation/d/1QL-XPFXiFpDBh86DbEegFXBXFXjix4v032GhShbKf3s/edit?usp=sharing).
zhuwenwen's avatar
zhuwenwen committed
37
38
39
40
- [2023/08] We would like to express our sincere gratitude to [Andreessen Horowitz](https://a16z.com/2023/08/30/supporting-the-open-source-ai-community/) (a16z) for providing a generous grant to support the open-source development and research of vLLM.
- [2023/06] We officially released vLLM! FastChat-vLLM integration has powered [LMSYS Vicuna and Chatbot Arena](https://chat.lmsys.org) since mid-April. Check out our [blog post](https://vllm.ai).

---
zhuwenwen's avatar
zhuwenwen committed
41

zhuwenwen's avatar
zhuwenwen committed
42
## About
zhuwenwen's avatar
zhuwenwen committed
43

zhuwenwen's avatar
zhuwenwen committed
44
45
vLLM is a fast and easy-to-use library for LLM inference and serving.

zhuwenwen's avatar
zhuwenwen committed
46
Originally developed in the [Sky Computing Lab](https://sky.cs.berkeley.edu) at UC Berkeley, vLLM has evolved into a community-driven project with contributions from both academia and industry.
zhuwenwen's avatar
zhuwenwen committed
47

zhuwenwen's avatar
zhuwenwen committed
48
49
50
vLLM is fast with:

- State-of-the-art serving throughput
zhuwenwen's avatar
zhuwenwen committed
51
- Efficient management of attention key and value memory with [**PagedAttention**](https://blog.vllm.ai/2023/06/20/vllm.html)
zhuwenwen's avatar
zhuwenwen committed
52
53
- Continuous batching of incoming requests
- Fast model execution with CUDA/HIP graph
54
55
56
57
- Quantizations: [GPTQ](https://arxiv.org/abs/2210.17323), [AWQ](https://arxiv.org/abs/2306.00978), INT4, INT8, and FP8.
- Optimized CUDA kernels, including integration with FlashAttention and FlashInfer.
- Speculative decoding
- Chunked prefill
zhuwenwen's avatar
zhuwenwen committed
58

zhuwenwen's avatar
zhuwenwen committed
59
**Performance benchmark**: We include a performance benchmark at the end of [our blog post](https://blog.vllm.ai/2024/09/05/perf-update.html). It compares the performance of vLLM against other LLM serving engines ([TensorRT-LLM](https://github.com/NVIDIA/TensorRT-LLM), [SGLang](https://github.com/sgl-project/sglang) and [LMDeploy](https://github.com/InternLM/lmdeploy)). The implementation is under [nightly-benchmarks folder](.buildkite/nightly-benchmarks/) and you can [reproduce](https://github.com/vllm-project/vllm/issues/8176) this benchmark using our one-click runnable script.
60

zhuwenwen's avatar
zhuwenwen committed
61
62
63
64
vLLM is flexible and easy to use with:

- Seamless integration with popular Hugging Face models
- High-throughput serving with various decoding algorithms, including *parallel sampling*, *beam search*, and more
65
- Tensor parallelism and pipeline parallelism support for distributed inference
zhuwenwen's avatar
zhuwenwen committed
66
67
- Streaming outputs
- OpenAI-compatible API server
68
69
70
- Support NVIDIA GPUs, AMD CPUs and GPUs, Intel CPUs and GPUs, PowerPC CPUs, TPU, and AWS Neuron.
- Prefix caching support
- Multi-lora support
zhuwenwen's avatar
zhuwenwen committed
71

zhuwenwen's avatar
zhuwenwen committed
72
73
vLLM seamlessly supports most popular open-source models on HuggingFace, including:
- Transformer-like LLMs (e.g., Llama)
74
- Mixture-of-Expert LLMs (e.g., Mixtral, Deepseek-V2 and V3)
75
- Embedding Models (e.g. E5-Mistral)
zhuwenwen's avatar
zhuwenwen committed
76
77
78
79
80
- Multi-modal LLMs (e.g., LLaVA)

Find the full list of supported models [here](https://docs.vllm.ai/en/latest/models/supported_models.html).

## Getting Started
zhuwenwen's avatar
zhuwenwen committed
81

zhuwenwen's avatar
zhuwenwen committed
82
Install vLLM with `pip` or [from source](https://docs.vllm.ai/en/latest/getting_started/installation/gpu/index.html#build-wheel-from-source):
zhuwenwen's avatar
zhuwenwen committed
83
84
85
86
87

```bash
pip install vllm
```

zhuwenwen's avatar
zhuwenwen committed
88
89
90
91
Visit our [documentation](https://docs.vllm.ai/en/latest/) to learn more.
- [Installation](https://docs.vllm.ai/en/latest/getting_started/installation/index.html)
- [Quickstart](https://docs.vllm.ai/en/latest/getting_started/quickstart.html)
- [List of Supported Models](https://docs.vllm.ai/en/latest/models/supported_models.html)
zhuwenwen's avatar
zhuwenwen committed
92
93
94
95
96
97

## Contributing

We welcome and value any contributions and collaborations.
Please check out [CONTRIBUTING.md](./CONTRIBUTING.md) for how to get involved.

zhuwenwen's avatar
zhuwenwen committed
98
99
100
101
102
103
## Sponsors

vLLM is a community project. Our compute resources for development and testing are supported by the following organizations. Thank you for your support!

<!-- Note: Please sort them in alphabetical order. -->
<!-- Note: Please keep these consistent with docs/source/community/sponsors.md -->
zhuwenwen's avatar
zhuwenwen committed
104
Cash Donations:
zhuwenwen's avatar
zhuwenwen committed
105
- a16z
zhuwenwen's avatar
zhuwenwen committed
106
107
108
109
110
111
- Dropbox
- Sequoia Capital
- Skywork AI
- ZhenFund

Compute Resources:
zhuwenwen's avatar
zhuwenwen committed
112
113
114
115
116
117
- AMD
- Anyscale
- AWS
- Crusoe Cloud
- Databricks
- DeepInfra
118
- Google Cloud
zhuwenwen's avatar
zhuwenwen committed
119
- Lambda Lab
zhuwenwen's avatar
zhuwenwen committed
120
- Nebius
zhuwenwen's avatar
zhuwenwen committed
121
- Novita AI
zhuwenwen's avatar
zhuwenwen committed
122
123
124
125
126
127
128
- NVIDIA
- Replicate
- Roblox
- RunPod
- Trainy
- UC Berkeley
- UC San Diego
zhuwenwen's avatar
zhuwenwen committed
129
130

Slack Sponsor: Anyscale
zhuwenwen's avatar
zhuwenwen committed
131
132
133

We also have an official fundraising venue through [OpenCollective](https://opencollective.com/vllm). We plan to use the fund to support the development, maintenance, and adoption of vLLM.

zhuwenwen's avatar
zhuwenwen committed
134
135
136
## Citation

If you use vLLM for your research, please cite our [paper](https://arxiv.org/abs/2309.06180):
zhuwenwen's avatar
zhuwenwen committed
137

zhuwenwen's avatar
zhuwenwen committed
138
139
140
141
142
143
144
```bibtex
@inproceedings{kwon2023efficient,
  title={Efficient Memory Management for Large Language Model Serving with PagedAttention},
  author={Woosuk Kwon and Zhuohan Li and Siyuan Zhuang and Ying Sheng and Lianmin Zheng and Cody Hao Yu and Joseph E. Gonzalez and Hao Zhang and Ion Stoica},
  booktitle={Proceedings of the ACM SIGOPS 29th Symposium on Operating Systems Principles},
  year={2023}
}
zhuwenwen's avatar
zhuwenwen committed
145
146
147
148
```

## Contact Us

zhuwenwen's avatar
zhuwenwen committed
149
150
151
152
- For technical questions and feature requests, please use Github issues or discussions.
- For discussing with fellow users and coordinating contributions and development, please use Slack.
- For security disclosures, please use Github's security advisory feature.
- For collaborations and partnerships, please contact us at vllm-questions AT lists.berkeley.edu.
zhuwenwen's avatar
zhuwenwen committed
153
154
155

## Media Kit

zhuwenwen's avatar
zhuwenwen committed
156
- If you wish to use vLLM's logo, please refer to [our media kit repo](https://github.com/vllm-project/media-kit).