Commits · d5cb0be2cd16e6c5eefd4d266a38357fde83a660 · ModelZoo / Qwen_lmdeploy

24 Aug, 2023 1 commit
- [Fix] Fix llama2 70b & qwen quantization error (#273) · d5cb0be2
  pppppM authored Aug 24, 2023
```
* fix llama2 70b

* fix qwen quantization

* remove pdb

* add faq
```
  d5cb0be2
22 Aug, 2023 1 commit

AllentDan authored Aug 22, 2023

* add restful api

* refine

* add simple doc

* lint

* add uvicorn requirement

* more args

* add llama2

* docstring

* update doc

* save

* refine

* lint

* better decode

* add v1/embedding

* add GenerateRequest

* add llama2 chat template

* correct profiling

* update documents

* add length judge

* add faq

* update doc and rename req_que to req_queue

* fix md link, use get_logger, fix sequence_end bug

* use another doc link for go to avoid lint error

* add api_client.py

* update doc

* update doc

* update function interface

* update FAQ

* resolve comments

d5c10e7a

21 Aug, 2023 1 commit

Pass chat template args including meta_prompt to model (#225) · 7785142d

AllentDan authored Aug 21, 2023

* pass args like meta_prompt to model

* update chatbot

* update

* rollback

* update llama2 and qwen

* refine

7785142d

18 Aug, 2023 2 commits

Support TP for w4a16 (#262) · 89f3d322
Li Zhang authored Aug 18, 2023

89f3d322

[Feature] Support Qwen-7B, dynamic NTK scaling and logN scaling in turbomind (#230) · 4a60b45d

Li Zhang authored Aug 18, 2023

* qwen support

* dynamic ntk & logn attn

* fix ntk & add chat template

* fix ntk scaling & stop words

* fix lint

* add tiktoken to requirements.txt

* fix tokenizer, set model format automatically

* update model.py

* update readme

* fix lint

4a60b45d

16 Aug, 2023 2 commits
- Adjust dependency of gradio server (#236) · 0d21f366
  AllentDan authored Aug 16, 2023
```
* import if lib directory exists

* only modify app.py
```
  0d21f366
- remove chat template (#252) · f06db80d
  Lyu Han authored Aug 16, 2023
  
  f06db80d
15 Aug, 2023 1 commit
- Fix wrong RPATH using the absolute path instead of relative one (#239) · 271a19fe
  Chen Xin authored Aug 15, 2023
  
  271a19fe
14 Aug, 2023 4 commits
- Bump version to v0.0.4 (#231) · 8cdcb2a9
  Lyu Han authored Aug 14, 2023
  
  8cdcb2a9
- Fix TIS client got-no-space-result side effect brought by PR #197 (#222) · 68296844
  Lyu Han authored Aug 14, 2023
```
* rollback

* rollback chatbot.py
```
  68296844
- feat(quantization): kv cache use asymmetric (#218) · 902a3e16
  tpoisonooo authored Aug 14, 2023
```
* feat(quantization): kv cache use asymmetric
```
  902a3e16
- [Feature] Blazing fast W4A16 inference (#202) · c3290cad
  Li Zhang authored Aug 14, 2023
```
* add w4a16

* fix `deploy.py`

* add doc

* add w4a16 kernels

* fuse w1/w3 & bugfixes

* fix typo

* python

* guard sm75/80 features

* add missing header

* refactor

* qkvo bias

* update cost model

* fix lint

* update `deploy.py`
```
  c3290cad
11 Aug, 2023 1 commit

[Feature] Support AWQ (#108) · d3dbe179

pppppM authored Aug 11, 2023

* support kv cache offload

* add dataloader docstring

* complete gitignore

* refactor collect mod fn

* add calibration

* fix lint

* add observers and quantizers

* fix lints

* add global available mixin

* fix lints

* split batch inference

* support smoothquant and awq

* update export kv scales

* fix lints

* fix some bugs

* update weight only usage

* update usage

* auto mapping and support smooth internlm

* trust remote code

* fix num head key error

* fix bias error

* align shape and pack order with llm-awq

* modified according to LZHgrla's comments.

* update gitignore

* fix kv qparams export error

* update usage

* decouple calibrate and awq

* update docstrings

* update api name

* update readme

* update readme

* update readme

* update readme

* update kv_qparams and readme

* fix typos

d3dbe179

07 Aug, 2023 5 commits

[Refactor] Support multi-session chat (#178) · 4bd0b487

WRH authored Aug 07, 2023

* add some dist utils

* add model utils

* add termio and basicstreamer

* typo

* fix world size

* refactor chat and tested llama1

* add internlm adapter and support stoping criteria

* concat with id for internlm

* update docstring

* update and support llama2

* typo

* move docs to docs

* update docstring of session manager

* update docstring

* update docs

* fix accel none in model

* fix and add test for tensor broadcast

* fix session using typing to check type

* add docstrings and comprehensive condition test

* unit test for dist

* fix session

* split unittests of utils

* typo

* update control flow of accel

* move test model

* remove main in unittest

* remove some log

* remove some comments

4bd0b487

bump version to v0.0.3 (#205) · c80f3e49
lvhan028 authored Aug 07, 2023

c80f3e49
Add non-stream inference api for chatbot (#200) · 3de0dbb6
lvhan028 authored Aug 07, 2023
```
* add non-stream inference api for chatbot

* update according to reviewer's comments
```
3de0dbb6
[Feature] Add script to split HuggingFace model to the smallest sharded checkpoints (#199) · b7e7e668
LZHgrla authored Aug 07, 2023
```
* add get_small_sharded_hf.py

* fix pre-commit
```
b7e7e668
Improve postprocessing in TIS serving by applying Incremental de-tokenizing (#197) · 0ed1e4d4
lvhan028 authored Aug 07, 2023
```
* change to incremental decoding

* update
```
0ed1e4d4

04 Aug, 2023 1 commit

Support serving with gradio without communicating to TIS (#162) · 18c386d9

AllentDan authored Aug 04, 2023



* use local model for webui

* local model for app.py

* lint

* remove print

* add seed

* comments

* fixed seesion_id

* support turbomind batch inference

* update app.py

* lint and docstring

* move webui to serve/gradio

* update doc

* update doc

* update docstring and rmeove print conversition

* log

* Update docs/zh_cn/build.md
Co-authored-by: Chen Xin <xinchen.tju@gmail.com>

* Update docs/en/build.md
Co-authored-by: Chen Xin <xinchen.tju@gmail.com>

* use latest gradio

* fix

* replace partial with InterFace

* use host ip instead of coolie

---------
Co-authored-by: Chen Xin <xinchen.tju@gmail.com>

18c386d9

03 Aug, 2023 1 commit
- Move lmdeploy/turbomind/utils.py to lmdeploy/utils.py (#191) · 7a2128be
  lvhan028 authored Aug 03, 2023
  
  7a2128be
31 Jul, 2023 1 commit

Support Runtime tensor parallelism (#158) · 4767b04d

q.yao authored Jul 31, 2023

* works on interlm and vicuna

* support GQA

* remove comment

* update readme, add logger, default tp=1

* remove log

4767b04d

28 Jul, 2023 1 commit

bump version to v0.0.2 (#177) · 7e0b75bb

lvhan028 authored Jul 28, 2023

* bump version to v0.0.2

* fix command

* update installation and inference section

7e0b75bb

27 Jul, 2023 1 commit
- add model_name param for chatbot (#174) · 7bc8d171
  MaxMatthew authored Jul 27, 2023
  
  7bc8d171
26 Jul, 2023 1 commit
- Add triton_models to whl package (#163) · e7bc11b4
  Chen Xin authored Jul 26, 2023
```
* defer symlink

* fix lint
```
  e7bc11b4
25 Jul, 2023 2 commits
- support fmha gqa (#160) · 5ed6bb59
  q.yao authored Jul 25, 2023
```
Co-authored-by: grimoire <yaoqian@pjlab.org.cn>
```
  5ed6bb59
- fix getting package root path error in python3.9 (#157) · 5203c850
  lvhan028 authored Jul 25, 2023
  
  5203c850
24 Jul, 2023 1 commit
- [Feature] decode-only forward pass (#153) · 0cc9d095
  Li Zhang authored Jul 24, 2023
```
* decode only forward pass

* fix lint

* batch embedding
```
  0cc9d095
23 Jul, 2023 1 commit

Refactor the chat template of supported models using factory pattern (#144) · 7b470f07

lvhan028 authored Jul 23, 2023

* refactor model.py and support baichuan-7b

* remove model_name

* remove hard session_len

* export tokenizer.py to target dir

* remove model_name from client

* remove model_name

* update

* correct throughput equation

* fix session.response

* update serving.md

* update readme

* update according to review comments

* update

* update

* update

* update

7b470f07

22 Jul, 2023 1 commit

add profile throughput benchmark (#146) · 2067862d

q.yao authored Jul 22, 2023



* add profile throughput benchmark

* add output only throughput

* update req/min

* update benckmark readme

* fix lint

---------
Co-authored-by: grimoire <yaoqian@pjlab.org.cn>

2067862d

21 Jul, 2023 3 commits

remove slicing reponse and add resume api (#154) · b728064e

MaxMatthew authored Jul 21, 2023

* Fix lmdeploy.serve.turbomind bug
* add __init__.py for turbomind
* add resume function
* fix the assignment for session.response

* Fix code style

b728064e

[Feature] Support Llama-2 with GQA (#147) · f07b697b

Li Zhang authored Jul 21, 2023

* add GQA for llama2

* fix model conversion

* fix lint & remove dev log

* update news

* minor

* fix allocation size

* fix split_dim for w_qkv.bias

f07b697b

[Fix] Support DeepSpeed on autoTP and kernel injection (#138) · 2a475478

Kevin Wang authored Jul 21, 2023



* [Fix] fix issue 127

* 优化防止接口更改

* 如果没有deepspeed用python启动需要手动加载到GPU上

* rollback the changes about max_out_tokens and delelte torch > 2.0 if statement

* support kernel injection with customized deepspeed

* spelling error

* Update chat.py

---------
Co-authored-by: wangruohui <12756472+wangruohui@users.noreply.github.com>

2a475478

20 Jul, 2023 3 commits

add llama2 chat template (#140) · 406f8c9f

q.yao authored Jul 20, 2023



* add llama2 template

* update readme and fix lint

* update readme

* add bos

* add bos

* remove bos

* Update model.py

---------
Co-authored-by: grimoire <yaoqian@pjlab.org.cn>

406f8c9f

return carriage cause overwriting at the same line (#143) · 8ba2d7c5
WRH authored Jul 20, 2023

8ba2d7c5

[Fix] Fix bug for issues #141 (#145) · cde17e73

humu789 authored Jul 20, 2023

* fix get_dataset error

* fix lint

* add datasets to requirements.txt

* update some msci

cde17e73

19 Jul, 2023 2 commits
- fix the offset during streaming chat (#142) · 289ffa3c
  lvhan028 authored Jul 19, 2023
  
  289ffa3c
- Fix tensor-parallel inference of internlm with bias (#135) · 79595cd1
  q.yao authored Jul 19, 2023
```
* remove copy

* repetition_penalty=1

* add repetition_penalty to chat args

* update readme

* update readme
```
  79595cd1
18 Jul, 2023 3 commits

update doc and requirements.txt (#119) · 4970d798

AllentDan authored Jul 18, 2023



* update requirements

* update transformers version

* lint

* comments

* lint

* update requirements

* remove setup_requires

---------
Co-authored-by: dongchunyu <dongchunyu@pjlab.org.cn>

4970d798

print info copy-paste error (#133) · 8664946d
Kevin Wang authored Jul 18, 2023

8664946d

Tensor Parallel python api (#82) · 7cbfe2ea

q.yao authored Jul 18, 2023

* wip

* profile disable tp

* fix profile

* lint

* fix dlpack

* remove comment

* add tp flag

* add session len check

* add eos

* remove tp and session len inputs

* warp tokenizer

* multithread load weight

* update profile

* refactor tokenizer

* remove pre/post process

* remove mpi4py requirement

* remove

* remove bind

* remove mpi requirement

* check backend_tokenizer

7cbfe2ea