- 15 Dec, 2023 1 commit
-
-
Chen Xin authored
* support image_embs input * add some checks * update interactive/config.pbtxt && TurbomindModelConfig * update docstring * refactor * support convert embeddings to bf16 * update interactive/config.pbtxt * embeddings -> input_embeddings * use input_embedding_ranges * remove embedding_begins/ends
-
- 14 Dec, 2023 1 commit
-
-
AllentDan authored
-
- 13 Dec, 2023 3 commits
-
-
AllentDan authored
* add encode for opencompass * doc * remove **kwargs
-
RunningLeon authored
* update dockerfile and ci * remove unused env * ignore link check of reddit website urls
-
AllentDan authored
* add api.py * update serve function * add model_name arg and provide examples * docstring * remove service_available * type hint
-
- 12 Dec, 2023 3 commits
- 11 Dec, 2023 4 commits
-
-
AllentDan authored
* FIFO pipe for api_server * asyncio sleep 0 * remove unwanted import * rename symbols * speed benchmark up by disable preprocess for string input * replace Queue with set * comment
-
Li Zhang authored
* disable attention mask when not needed * fix for sm<80 and float data type
-
Li Zhang authored
-
Li Zhang authored
* simplify block manager * fix lint
-
- 07 Dec, 2023 1 commit
-
-
Li Zhang authored
-
- 06 Dec, 2023 3 commits
-
-
Lyu Han authored
-
Lyu Han authored
* update test scripts for models with different sizes * update * only test after tunning gemm * chmod +x * fix typo * benchmark on a100 * fix typo * fix typo * per-token latency percentile in profile_throughput * fix * fix * rename * make the script accept parameters * minor fix * indent * reformat table * change to 3000 * minor fix
-
Lyu Han authored
-
- 05 Dec, 2023 2 commits
- 04 Dec, 2023 4 commits
-
-
Chen Xin authored
* update cuda12.1 build check ci * use matrix
-
Li Zhang authored
* Unify prefill and decode passes * dynamic split-fuse * refactor * correct input count calculation * remove unused * lint * lint * fix msvc build * fix msvc build * fix msvc build * fix msvc build * fix msvc build * fix msvc build * fix msvc build * fix msvc build * fix msvc build
-
Lyu Han authored
* minor fix in the profile scripts and docs * miss arguments * typo * fix lint * update
-
AllentDan authored
-
- 02 Dec, 2023 1 commit
-
-
Li Zhang authored
-
- 29 Nov, 2023 7 commits
-
-
Lyu Han authored
* user guide of benchmark generation * update benchmark generation guide * update profiling throughput guide * update profiling api_server guide * rename file names * update profile tis user guide * update * fix according to review comments * update * update according to review comments * updaste * add an example * update
-
Lyu Han authored
-
Chen Xin authored
-
Lyu Han authored
* update profile scripts * add top_p, top_k and temperature as input arguments * fix input_ids * update profile_throughput * update profile_restful_api * update profile_serving * update * update * add progress bar * remove TODO comments * update * remove useless profile_* argument * remove log level * change concurrency default value to 64 * update restful_api.md * update according to review comments * fix docstring
-
tpoisonooo authored
* feat(build): enable ninja and lld * fix(.github): add ninja installation * fix(CI): remove dimsize=256 * fix(CI): add option for generate.sh * fix(docs): update
-
q.yao authored
* fix * fix lint
-
RunningLeon authored
* add triton server test and workflow yml * update * revert changes in dockerfile * update prompts
-
- 28 Nov, 2023 1 commit
-
-
q.yao authored
-
- 27 Nov, 2023 2 commits
- 24 Nov, 2023 1 commit
-
-
Lyu Han authored
-
- 23 Nov, 2023 3 commits
- 22 Nov, 2023 1 commit
-
-
Chen Xin authored
* turbomind support export model params * fix overflow * support turbomind.from_pretrained * fix tp * support AutoModel * support load kv qparams * update auto_awq * udpate docstring * export lmdeploy version * update doc * remove download_hf_repo * LmdeployForCausalLM -> LmdeployForCausalLM * refactor turbomind.py * update comment * add bfloat16 convert back * support gradio run_locl load hf * support resuful api server load hf * add docs * support loading previous quantized model * adapt pr 690 * udpate docs * not export turbomind config when quantize a model * check model_name when can not get it from config.json * update readme * remove model_name in auto_awq * update * update * udpate * fix build * absolute import
-
- 21 Nov, 2023 1 commit
-
-
Zaida Zhou authored
-
- 20 Nov, 2023 1 commit
-
-
Lyu Han authored
* update * update config guide * update guide * upate user guide according to review comments
-