- 20 Aug, 2024 1 commit
-
-
Antoni Baum authored
-
- 01 Aug, 2024 1 commit
-
-
Woosuk Kwon authored
-
- 27 Jul, 2024 2 commits
-
-
Woosuk Kwon authored
-
Woosuk Kwon authored
-
- 16 Jul, 2024 1 commit
-
-
Michael Goin authored
-
- 12 Jul, 2024 1 commit
-
-
Woosuk Kwon authored
-
- 08 Jul, 2024 1 commit
-
-
afeldman-nm authored
[Kernel] Correctly invoke prefill & decode kernels for cross-attention (towards eventual encoder/decoder model support) (#4888) Co-authored-by:Woosuk Kwon <woosuk.kwon@berkeley.edu>
-
- 28 Jun, 2024 1 commit
-
-
Woosuk Kwon authored
-
- 26 Jun, 2024 3 commits
-
-
Woosuk Kwon authored
-
Woosuk Kwon authored
-
Stephanie Wang authored
Signed-off-by:
Stephanie Wang <swang@cs.berkeley.edu> Signed-off-by:
Stephanie <swang@anyscale.com> Co-authored-by:
Stephanie <swang@anyscale.com>
-
- 14 Jun, 2024 1 commit
-
-
Woosuk Kwon authored
-
- 13 Jun, 2024 1 commit
-
-
youkaichao authored
[Core][Distributed] add coordinator to reduce code duplication in tp and pp (#5293)
-
- 12 Jun, 2024 1 commit
-
-
Woosuk Kwon authored
-