- 06 Mar, 2024 1 commit
-
-
zhouxiang authored
-
- 04 Mar, 2024 1 commit
-
-
zhouxiang authored
-
- 20 Feb, 2024 1 commit
-
-
zhouxiang authored
-
- 15 Jan, 2024 1 commit
-
-
xiabo authored
-
- 12 Jan, 2024 4 commits
- 20 Dec, 2023 2 commits
- 18 Dec, 2023 8 commits
-
-
Lyu Han authored
-
Chen Xin authored
-
AllentDan authored
* launch gradio server directly with hf model * end session * end session * fix api_server backend for gradio * fix out of boundary index * remove log
-
q.yao authored
-
maxchiron authored
-
maxchiron authored
fix ".generate" to ".chat_interactive_v1"
-
pppppM authored
-
AllentDan authored
* await before get * add comments * recover stop back
-
- 15 Dec, 2023 4 commits
-
-
Yam(长琴) authored
* fix: awq should save bin files * fix: doc * Update kv_int8.md * fix lint --------- Co-authored-by:pppppM <gjf_mail@126.com>
-
q.yao authored
-
q.yao authored
* Add bf16 template sp * prepare merge * add enable bf * add bf16 decode attention support * fix python lint * fix yapf * fix c format * c format11 * fix cast * fix on sm<80 * fix linux bf162 cast * fix type cast * fix lint * support from hf pretrained * fix pybind * fix converter * add trust remote code * fix comment * fix convert qwen * fix lint * fix baichuan * update weight map
-
Chen Xin authored
* support image_embs input * add some checks * update interactive/config.pbtxt && TurbomindModelConfig * update docstring * refactor * support convert embeddings to bf16 * update interactive/config.pbtxt * embeddings -> input_embeddings * use input_embedding_ranges * remove embedding_begins/ends
-
- 14 Dec, 2023 1 commit
-
-
AllentDan authored
-
- 13 Dec, 2023 3 commits
-
-
AllentDan authored
* add encode for opencompass * doc * remove **kwargs
-
RunningLeon authored
* update dockerfile and ci * remove unused env * ignore link check of reddit website urls
-
AllentDan authored
* add api.py * update serve function * add model_name arg and provide examples * docstring * remove service_available * type hint
-
- 12 Dec, 2023 3 commits
- 11 Dec, 2023 4 commits
-
-
AllentDan authored
* FIFO pipe for api_server * asyncio sleep 0 * remove unwanted import * rename symbols * speed benchmark up by disable preprocess for string input * replace Queue with set * comment
-
Li Zhang authored
* disable attention mask when not needed * fix for sm<80 and float data type
-
Li Zhang authored
-
Li Zhang authored
* simplify block manager * fix lint
-
- 07 Dec, 2023 1 commit
-
-
Li Zhang authored
-
- 06 Dec, 2023 3 commits
-
-
Lyu Han authored
-
Lyu Han authored
* update test scripts for models with different sizes * update * only test after tunning gemm * chmod +x * fix typo * benchmark on a100 * fix typo * fix typo * per-token latency percentile in profile_throughput * fix * fix * rename * make the script accept parameters * minor fix * indent * reformat table * change to 3000 * minor fix
-
Lyu Han authored
-
- 05 Dec, 2023 2 commits
- 04 Dec, 2023 1 commit
-
-
Chen Xin authored
* update cuda12.1 build check ci * use matrix
-