- 06 Sep, 2024 1 commit
-
-
Dipika Sikka authored
-
- 28 Aug, 2024 2 commits
-
-
Mor Zusman authored
-
rasmith authored
[Kernel] [Triton] [AMD] Adding Triton implementations awq_dequantize and awq_gemm to support AWQ (#7386)
-
- 27 Aug, 2024 1 commit
-
-
Dipika Sikka authored
Co-authored-by:ElizaWszola <eliza@neuralmagic.com>
-
- 22 Aug, 2024 1 commit
-
-
Michael Goin authored
-
- 21 Aug, 2024 1 commit
-
-
Dipika Sikka authored
Co-authored-by:ElizaWszola <eliza@neuralmagic.com>
-
- 20 Aug, 2024 1 commit
-
-
Lucas Wilkinson authored
-
- 16 Aug, 2024 2 commits
-
-
bnellnm authored
-
Charlie Fu authored
-
- 13 Aug, 2024 1 commit
-
-
Woosuk Kwon authored
-
- 06 Aug, 2024 1 commit
-
-
Luka Govedič authored
Co-authored-by:Tyler Michael Smith <tyler@neuralmagic.com>
-
- 05 Aug, 2024 1 commit
-
-
Isotr0py authored
Co-authored-by:Michael Goin <michael@neuralmagic.com>
-
- 02 Aug, 2024 1 commit
-
-
Lucas Wilkinson authored
-
- 01 Aug, 2024 1 commit
-
-
Jee Jee Li authored
-
- 31 Jul, 2024 3 commits
-
-
HandH1998 authored
-
Cyrus Leung authored
-
Cyrus Leung authored
-
- 30 Jul, 2024 1 commit
-
-
Tyler Michael Smith authored
-
- 27 Jul, 2024 1 commit
-
-
Alexander Matveev authored
-
- 24 Jul, 2024 1 commit
-
-
Antoni Baum authored
-
- 21 Jul, 2024 1 commit
-
-
Alexander Matveev authored
-
- 20 Jul, 2024 2 commits
-
-
Robert Shaw authored
-
Varun Sundar Rabindranath authored
Co-authored-by:Varun Sundar Rabindranth <varun@neuralmagic.com>
-
- 19 Jul, 2024 1 commit
-
-
Robert Shaw authored
-
- 18 Jul, 2024 1 commit
-
-
Varun Sundar Rabindranath authored
Co-authored-by:Varun Sundar Rabindranath <varun@neuralmagic.com>
-
- 17 Jul, 2024 1 commit
-
-
Alexander Matveev authored
-
- 16 Jul, 2024 1 commit
-
-
Michael Goin authored
-
- 03 Jul, 2024 1 commit
-
-
Michael Goin authored
-
- 26 Jun, 2024 1 commit
-
-
Luka Govedič authored
Co-authored-by:
Chih-Chieh-Yang <7364402+cyang49@users.noreply.github.com> Co-authored-by:
Lucas Wilkinson <lwilkinson@neuralmagic.com>
-
- 20 Jun, 2024 2 commits
-
-
Tyler Michael Smith authored
-
Roger Wang authored
-
- 17 Jun, 2024 1 commit
-
-
Kunshang Ji authored
Co-authored-by:
Jiang Li <jiang1.li@intel.com> Co-authored-by:
Abhilash Majumder <abhilash.majumder@intel.com> Co-authored-by:
Abhilash Majumder <30946547+abhilash1910@users.noreply.github.com>
-
- 13 Jun, 2024 1 commit
-
-
Tyler Michael Smith authored
Co-authored-by:
Michael Goin <michael@neuralmagic.com> Co-authored-by:
youkaichao <youkaichao@gmail.com> Co-authored-by:
zifeitong <zifei.tong@parasail.io> Co-authored-by:
Robert Shaw <114415538+robertgshaw2-neuralmagic@users.noreply.github.com>
-
- 12 Jun, 2024 1 commit
-
-
youkaichao authored
-
- 09 Jun, 2024 1 commit
-
-
bnellnm authored
-
- 07 Jun, 2024 3 commits
-
-
Dipika Sikka authored
Co-authored-by:
Varun Sundar Rabindranath <varunsundar08@gmail.com> Co-authored-by:
Varun Sundar Rabindranath <varun@neuralmagic.com>
-
Tyler Michael Smith authored
Switching from torch._scaled_mm to vLLM's cutlass fp8 kernels when supported as we are seeing 5-15% improvement in e2e performance on neuralmagic/Meta-Llama-3-8B-Instruct-FP8 see https://docs.google.com/spreadsheets/d/1GiAnmzyGHgZ6zL_LDSTm35Bdrt4A8AaFEurDlISYYA4/ for some quick e2e benchmarks and #5144 for comparisons across different GEMM sizes.
-
Jie Fu (傅杰) authored
-
- 03 Jun, 2024 1 commit
-
-
Tyler Michael Smith authored
-
- 25 May, 2024 1 commit
-
-
Eric Xihui Lin authored
Co-authored-by:
beagleski <yunanzhang@microsoft.com> Co-authored-by:
bapatra <bapatra@microsoft.com> Co-authored-by:
Barun Patra <codedecde@users.noreply.github.com> Co-authored-by:
Michael Goin <michael@neuralmagic.com>
-