Commits · cfec5bede59ebff0a62129abb5e1437b851635c5 · guobj / Qwen_lmdeploy

"app/vscode:/vscode.git/clone" did not exist on "5208cf09b13a5630ce967e0bccd1821ac8d5309a"

11 Sep, 2023 2 commits

[Fix] Update puyu model (#399) · cfec5bed
liukuikun authored Sep 11, 2023

cfec5bed

Lyu Han authored Sep 11, 2023

* tmp

* add demo for codellama inference

* update

* update

* update

* update codellama.md

* export rope_theta

* update

* update doc

* fix client.py

* define SamplingParam

* rollback 'end'

* rotary_emb_base to rotary_embedding_base

* change to baichuan2-7b

65c662f9

08 Sep, 2023 1 commit

Support baichuan2-chat chat template (#378) · 55764e0b

WRH authored Sep 08, 2023



* support baichuan2-chat

* update args from generation config

* update deploy.py

* update readme

* tested with tp

* step-1 when last id is eos

* add news

---------
Co-authored-by: chenxin <chenxin@pjlab.org.cn>

55764e0b

01 Sep, 2023 1 commit

Decode generated token_ids incrementally (#309) · 9bfe03c6

AllentDan authored Sep 01, 2023

* add incremental decoding for turbomind

* update TIS

* fix triton post processing

* update doc

* fix typo

* SentencePieceTokenizer incremental decode, add qwen message prompt

* docstring

* update bot

9bfe03c6

22 Aug, 2023 1 commit

Add Restful API (#223) · d5c10e7a

AllentDan authored Aug 22, 2023

* add restful api

* refine

* add simple doc

* lint

* add uvicorn requirement

* more args

* add llama2

* docstring

* update doc

* save

* refine

* lint

* better decode

* add v1/embedding

* add GenerateRequest

* add llama2 chat template

* correct profiling

* update documents

* add length judge

* add faq

* update doc and rename req_que to req_queue

* fix md link, use get_logger, fix sequence_end bug

* use another doc link for go to avoid lint error

* add api_client.py

* update doc

* update doc

* update function interface

* update FAQ

* resolve comments

d5c10e7a

21 Aug, 2023 1 commit

Pass chat template args including meta_prompt to model (#225) · 7785142d

AllentDan authored Aug 21, 2023

* pass args like meta_prompt to model

* update chatbot

* update

* rollback

* update llama2 and qwen

* refine

7785142d

18 Aug, 2023 1 commit

[Feature] Support Qwen-7B, dynamic NTK scaling and logN scaling in turbomind (#230) · 4a60b45d

Li Zhang authored Aug 18, 2023

* qwen support

* dynamic ntk & logn attn

* fix ntk & add chat template

* fix ntk scaling & stop words

* fix lint

* add tiktoken to requirements.txt

* fix tokenizer, set model format automatically

* update model.py

* update readme

* fix lint

4a60b45d

16 Aug, 2023 1 commit
- remove chat template (#252) · f06db80d
  Lyu Han authored Aug 16, 2023
  
  f06db80d
23 Jul, 2023 1 commit

Refactor the chat template of supported models using factory pattern (#144) · 7b470f07

lvhan028 authored Jul 23, 2023

* refactor model.py and support baichuan-7b

* remove model_name

* remove hard session_len

* export tokenizer.py to target dir

* remove model_name from client

* remove model_name

* update

* correct throughput equation

* fix session.response

* update serving.md

* update readme

* update according to review comments

* update

* update

* update

* update

7b470f07

20 Jul, 2023 1 commit

add llama2 chat template (#140) · 406f8c9f

q.yao authored Jul 20, 2023



* add llama2 template

* update readme and fix lint

* update readme

* add bos

* add bos

* remove bos

* Update model.py

---------
Co-authored-by: grimoire <yaoqian@pjlab.org.cn>

406f8c9f

19 Jul, 2023 2 commits
- fix the offset during streaming chat (#142) · 289ffa3c
  lvhan028 authored Jul 19, 2023
  
  289ffa3c
- Fix tensor-parallel inference of internlm with bias (#135) · 79595cd1
  q.yao authored Jul 19, 2023
```
* remove copy

* repetition_penalty=1

* add repetition_penalty to chat args

* update readme

* update readme
```
  79595cd1
14 Jul, 2023 2 commits
- add puyu model for internal use (#105) · 4cfb118f
  lvhan028 authored Jul 14, 2023
```
* add puyu model for internal use

* get/set session

* update

* add docstring
```
  4cfb118f
- miss <bos> of InternLM chat template (#112) · 86e70591
  lvhan028 authored Jul 14, 2023
  
  86e70591
12 Jul, 2023 1 commit
- add docstring for turbomind (#97) · 955c019c
  lvhan028 authored Jul 12, 2023
```
* add docstring

* update

* update

* fix according to review results
```
  955c019c
05 Jul, 2023 1 commit

update internlm‘s chat template (#54) · 3de27ead

lvhan028 authored Jul 05, 2023

* update internlm model

* update

* update

* update

* update

* update temperature, topk and top_p

* update

* update

* loosen log level

3de27ead

30 Jun, 2023 1 commit
- rename llmdeploy to lmdeploy (#30) · 46f4738c
  lvhan028 authored Jun 30, 2023
```
* change llmdeploy to lmdeploy

* update logo

* update readme
```
  46f4738c
25 Jun, 2023 1 commit

Add profile (#15) · 23c05372

lvhan028 authored Jun 25, 2023

* remove constraints on model name

* remove duplicate model converter

* add profile

* get eos and bos from server

* update stop_words

* update sequence_length when the last generated token is eos_id

* fix

* fix

* check-in models

* valicate model_name

* make stop_words as property

* debug profiling

* better stats

* fix assistant reponse

* update profile serving

* update

* update

23c05372