- 13 Dec, 2024 1 commit
-
-
Sungjae Lee authored
[Core] support LoRA and prompt adapter in content-based hashing for Block Manager v2 prefix caching (#8240)
-
- 11 Dec, 2024 1 commit
-
-
Cyrus Leung authored
Signed-off-by:DarkLight1337 <tlleungac@connect.ust.hk>
-
- 26 Nov, 2024 1 commit
-
-
Murali Andoorveedu authored
Signed-off-by:
andoorve <37849411+andoorve@users.noreply.github.com> Signed-off-by:
Sourashis Roy <sroy@roblox.com> Co-authored-by:
Sourashis Roy <sroy@roblox.com>
-
- 23 Nov, 2024 1 commit
-
-
Ricky Xu authored
Signed-off-by:rickyx <rickyx@anyscale.com>
-
- 21 Nov, 2024 1 commit
-
-
Pavani Majety authored
Signed-off-by:Pavani Majety <pmajety@nvidia.com>
-
- 06 Nov, 2024 1 commit
-
-
Aaron Pham authored
Signed-off-by:Aaron Pham <contact@aarnphm.xyz>
-
- 05 Nov, 2024 1 commit
-
-
Cyrus Leung authored
Signed-off-by:DarkLight1337 <tlleungac@connect.ust.hk>
-
- 24 Oct, 2024 1 commit
-
-
youkaichao authored
Co-authored-by:Zhuohan Li <zhuohan123@gmail.com>
-
- 18 Oct, 2024 2 commits
-
-
Cody Yu authored
-
Cyrus Leung authored
-
- 17 Oct, 2024 1 commit
-
-
Kuntai Du authored
Removing the block manager v1. This is the initial piece of prefix-caching-centric design. In order to achieve prefix-caching-centric design, we need to simplify the code path so that we only use v2 block manager (which has much higher performance on prefix caching).
-
- 10 Oct, 2024 1 commit
-
-
sroy745 authored
[Core] Add an environment variable which needs to be set explicitly to allow BlockSpaceManagerV1 (#9149)
-
- 07 Oct, 2024 1 commit
-
-
youkaichao authored
-
- 06 Oct, 2024 1 commit
-
-
Varun Sundar Rabindranath authored
Co-authored-by:Varun Sundar Rabindranath <varun@neuralmagic.com>
-
- 29 Sep, 2024 1 commit
-
-
sroy745 authored
-
- 25 Sep, 2024 2 commits
- 24 Sep, 2024 1 commit
-
-
sroy745 authored
-
- 28 Aug, 2024 1 commit
-
-
Cody Yu authored
-
- 27 Aug, 2024 1 commit
-
-
Megha Agarwal authored
Co-authored-by:Alexander Matveev <alexm@neuralmagic.com>
-
- 26 Aug, 2024 1 commit
-
-
Cody Yu authored
-
- 19 Aug, 2024 2 commits
-
-
Cody Yu authored
-
SangBin Cho authored
-
- 09 Aug, 2024 1 commit
-
-
Cade Daniel authored
-
- 08 Aug, 2024 1 commit
-
-
Zach Zheng authored
-
- 06 Aug, 2024 1 commit
-
-
afeldman-nm authored
[Core] Subclass ModelRunner to support cross-attention & encoder sequences (towards eventual encoder/decoder model support) (#4942) Co-authored-by:
Andrew Feldman <afeld2012@gmail.com> Co-authored-by:
Nick Hill <nickhill@us.ibm.com>
-
- 01 Aug, 2024 1 commit
-
-
youkaichao authored
-
- 22 Jul, 2024 1 commit
-
-
Jiaxin Shan authored
Co-authored-by:Antoni Baum <antoni.baum@protonmail.com>
-
- 19 Jul, 2024 1 commit
-
-
Antoni Baum authored
-
- 02 Jul, 2024 1 commit
-
-
Alexander Matveev authored
-
- 15 Jun, 2024 2 commits
-
-
Cyrus Leung authored
-
leiwen83 authored
Signed-off-by:
Lei Wen <wenlei03@qiyi.com> Co-authored-by:
Lei Wen <wenlei03@qiyi.com>
-
- 12 Jun, 2024 1 commit
-
-
SangBin Cho authored
-
- 03 Jun, 2024 1 commit
-
-
Kaiyang Chen authored
-
- 29 May, 2024 2 commits
-
-
Cyrus Leung authored
-
afeldman-nm authored
[Core] Cross-attention KV caching and memory-management (towards eventual encoder/decoder model support) (#4837)
-
- 28 May, 2024 2 commits
-
-
Cyrus Leung authored
Co-authored-by:Roger Wang <ywang@roblox.com>
-
Michał Moskal authored
Co-authored-by:Ruth Evans <ruthevans@Ruths-MacBook-Pro.local>
-
- 24 May, 2024 1 commit
-
-
leiwen83 authored
Co-authored-by:Lei Wen <wenlei03@qiyi.com>
-
- 13 May, 2024 1 commit
-
-
SangBin Cho authored
Co-authored-by:Robert Shaw <114415538+robertgshaw2-neuralmagic@users.noreply.github.com>
-