Commits · 459a88c8063d8ef7c4cd720a4e9524adf5b5c367 · OpenDAS / ColossalAI

30 Oct, 2023 2 commits

[Kernels]Updated Triton kernels into 2.1.0 and adding flash-decoding for llama... · 459a88c8

Cuiqing Li authored Oct 30, 2023


[Kernels]Updated Triton kernels into 2.1.0 and adding flash-decoding for llama token attention  (#4965)

* adding flash-decoding

* clean

* adding kernel

* adding flash-decoding

* add integration

* add

* adding kernel

* adding kernel

* adding triton 2.1.0 features for inference

* update bloom triton kernel

* remove useless vllm kernels

* clean codes

* fix

* adding files

* fix readme

* update llama flash-decoding

---------
Co-authored-by: cuiqing.li <lixx336@gmail.com>

459a88c8

[Inference] Dynamic Batching Inference, online and offline (#4953) · cf579ff4

Jianghai authored Oct 30, 2023



* [inference] Dynamic Batching for Single and Multiple GPUs (#4831)

* finish batch manager

* 1

* first

* fix

* fix dynamic batching

* llama infer

* finish test

* support different lengths generating

* del prints

* del prints

* fix

* fix bug

---------

Co-authored-by: CjhHa1 <cjh18671720497outlook.com>

* [inference] Async dynamic batching  (#4894)

* finish input and output logic

* add generate

* test forward

* 1

* [inference]Re push async dynamic batching (#4901)

* adapt to ray server

* finish async

* finish test

* del test

---------
Co-authored-by: yuehuayingxueluo <867460659@qq.com>

* Revert "[inference]Re push async dynamic batching (#4901)" (#4905)

This reverts commit fbf3c09e673794ed18c91d4bab1a7dfea052e95a.

* Revert "[inference] Async dynamic batching  (#4894)"

This reverts commit fced14025043e29ce816b315f440601188f7f79f.

* Revert "[inference] Async dynamic batching  (#4894)" (#4909)

This reverts commit fced14025043e29ce816b315f440601188f7f79f.

* Add Ray Distributed Environment Init Scripts

* support DynamicBatchManager base function

* revert _set_tokenizer version

* add driver async generate

* add async test

* fix bugs in test_ray_dist.py

* add get_tokenizer.py

* fix code style

* fix bugs about No module named 'pydantic' in ci test

* fix bugs in ci test

* fix bugs in ci test

* fix bugs in ci test

* [infer]Add Ray Distributed Environment Init Scripts (#4911)

* Revert "[inference] Async dynamic batching  (#4894)"

This reverts commit fced14025043e29ce816b315f440601188f7f79f.

* Add Ray Distributed Environment Init Scripts

* support DynamicBatchManager base function

* revert _set_tokenizer version

* add driver async generate

* add async test

* fix bugs in test_ray_dist.py

* add get_tokenizer.py

* fix code style

* fix bugs about No module named 'pydantic' in ci test

* fix bugs in ci test

* fix bugs in ci test

* fix bugs in ci test

* support dynamic batch for bloom model and is_running function

* [Inference]Test for new Async engine (#4935)

* infer engine

* infer engine

* test engine

* test engine

* new manager

* change step

* add

* test

* fix

* fix

* finish test

* finish test

* finish test

* finish test

* add license

---------
Co-authored-by: yuehuayingxueluo <867460659@qq.com>

* add assertion for config (#4947)

* [Inference] Finish dynamic batching offline test (#4948)

* test

* fix test

* fix quant

* add default

* fix

* fix some bugs

* fix some bugs

* fix

* fix bug

* fix bugs

* reset param

---------
Co-authored-by: yuehuayingxueluo <867460659@qq.com>
Co-authored-by: Cuiqing Li <lixx3527@gmail.com>
Co-authored-by: CjhHa1 <cjh18671720497outlook.com>

cf579ff4

27 Oct, 2023 2 commits

updated c++17 compiler flags (#4983) · 4e4a10c9
アマデウス authored Oct 27, 2023

4e4a10c9

[Pipeline inference] Combine kvcache with pipeline inference (#4938) · 1db67276

Bin Jia authored Oct 27, 2023

* merge kvcache with pipeline inference and refactor the code structure

* support ppsize > 2

* refactor pipeline code

* do pre-commit

* modify benchmark

* fix bench mark

* polish code

* add docstring and update readme

* refactor the code

* fix some logic bug of ppinfer

* polish readme

* fix typo

* skip infer test

1db67276

24 Oct, 2023 1 commit

[Inference]ADD Bench Chatglm2 script (#4963) · c6cd629e

Jianghai authored Oct 24, 2023

* add bench chatglm

* fix bug and make utils

---------

Co-authored-by: CjhHa1 <cjh18671720497outlook.com>

c6cd629e

20 Oct, 2023 2 commits

[inference] add reference and fix some bugs (#4937) · 785802e8

Xu Kai authored Oct 20, 2023



* add reference and fix some bugs

* update gptq init

---------
Co-authored-by: Xu Kai <xukai16@foxamil.com>

785802e8

[test] merge old components to test to model zoo (#4945) · b8e770c8

Hongxin Liu authored Oct 20, 2023

* [test] add custom models in model zoo

* [test] update legacy test

* [test] update model zoo

* [test] update gemini test

* [test] remove components to test

b8e770c8

19 Oct, 2023 1 commit

[Refactor] Integrated some lightllm kernels into token-attention (#4946) · 3a41e830

Cuiqing Li authored Oct 19, 2023



* add some req for inference

* clean codes

* add codes

* add some lightllm deps

* clean codes

* hello

* delete rms files

* add some comments

* add comments

* add doc

* add lightllm deps

* add lightllm cahtglm2 kernels

* add lightllm cahtglm2 kernels

* replace rotary embedding with lightllm kernel

* add some commnets

* add some comments

* add some comments

* add

* replace fwd kernel att1

* fix a arg

* add

* add

* fix token attention

* add some comments

* clean codes

* modify comments

* fix readme

* fix bug

* fix bug

---------
Co-authored-by: cuiqing.li <lixx336@gmail.com>
Co-authored-by: CjhHa1 <cjh18671720497@outlook.com>

3a41e830

18 Oct, 2023 4 commits
- [nfc] fix some typo with colossalai/ docs/ etc. (#4920) · 11009103
  digger yu authored Oct 18, 2023
  
  11009103
- [format] applied code formatting on changed files in pull request 4820 (#4886) · 486d06a2
  github-actions[bot] authored Oct 18, 2023
```
Co-authored-by: github-actions <github-actions@github.com>
```
  486d06a2
- [test] add no master test for low level zero plugin (#4934) · c7aa319b
  Zhongkai Zhao authored Oct 18, 2023
  
  c7aa319b
- [hotfix] fix torch 2.0 compatibility (#4936) · 1f5d2e80
  Hongxin Liu authored Oct 18, 2023
```
* [hotfix] fix launch

* [test] fix test gemini optim

* [shardformer] fix vit
```
  1f5d2e80
17 Oct, 2023 2 commits

[gemini] support gradient accumulation (#4869) · 21ba89ca

Baizhou Zhang authored Oct 17, 2023

* add test

* fix no_sync bug in low level zero plugin

* fix test

* add argument for grad accum

* add grad accum in backward hook for gemini

* finish implementation, rewrite tests

* fix test

* skip stuck model in low level zero test

* update doc

* optimize communication & fix gradient checkpoint

* modify doc

* cleaning codes

* update cpu adam fp16 case

21ba89ca

[format] applied code formatting on changed files in pull request 4908 (#4918) · a41cf88e
github-actions[bot] authored Oct 17, 2023
```
Co-authored-by: github-actions <github-actions@github.com>
```
a41cf88e

16 Oct, 2023 3 commits

[kernel] support pure fp16 for cpu adam and update gemini optim tests (#4921) · 4f68b3f1

Hongxin Liu authored Oct 16, 2023

* [kernel] support pure fp16 for cpu adam (#4896)

* [kernel] fix cpu adam kernel for pure fp16 and update tests (#4919)

* [kernel] fix cpu adam

* [test] update gemini optim test

4f68b3f1

Update flash_attention_patch.py · 7768afba

Zian(Andy) Zheng authored Oct 13, 2023

To be compatible with the new change in the Transformers library, where a new argument 'padding_mask' was added to forward function of attention layer.
https://github.com/huggingface/transformers/pull/25598

7768afba

[inference] Add smmoothquant for llama (#4904) · 611a5a80

Xu Kai authored Oct 16, 2023

* [inference] add int8 rotary embedding kernel for smoothquant (#4843)

* [inference] add smoothquant llama attention (#4850)

* add smoothquant llama attention

* remove uselss code

* remove useless code

* fix import error

* rename file name

* [inference] add silu linear fusion for smoothquant llama mlp  (#4853)

* add silu linear

* update skip condition

* catch smoothquant cuda lib exception

* prcocess exception for tests

* [inference] add llama mlp for smoothquant (#4854)

* add llama mlp for smoothquant

* fix down out scale

* remove duplicate lines

* add llama mlp check

* delete useless code

* [inference] add smoothquant llama (#4861)

* add smoothquant llama

* fix attention accuracy

* fix accuracy

* add kv cache and save pretrained

* refactor example

* delete smooth

* refactor code

* [inference] add smooth function and delete useless code for smoothquant (#4895)

* add smooth function and delete useless code

* update datasets

* remove duplicate import

* delete useless file

* refactor codes (#4902)

* rafactor code

* add license

* add torch-int and smoothquant license

611a5a80

13 Oct, 2023 2 commits

[feature] support no master weights option for low level zero plugin (#4816) · a0684e7b

Zhongkai Zhao authored Oct 13, 2023

* [feature] support no master weights for low level zero plugin

* [feature] support no master weights for low level zero plugin, remove data copy when no master weights

* remove data copy and typecasting when no master weights

* not load weights to cpu when using no master weights

* fix grad: use fp16 grad when no master weights

* only do not update working param when no master weights

* fix: only do not update working param when no master weights

* fix: passing params in dict format in hybrid plugin

* fix: remove extra params (tp_process_group) in hybrid_parallel_plugin

a0684e7b

[inference] add llama2 support (#4898) · 77a93283
Xu Kai authored Oct 13, 2023
```
* add llama2 support

* fix multi group bug
```
77a93283

12 Oct, 2023 4 commits

[hotfix] fix lr scheduler bug in torch 2.0 (#4864) · 39f2582e
Baizhou Zhang authored Oct 12, 2023

39f2582e

[feature] Add clip_grad_norm for hybrid_parallel_plugin (#4837) · 83b52c56

littsk authored Oct 12, 2023

* Add clip_grad_norm for hibrid_parallel_plugin

* polish code

* add unittests

* Move tp to a higher-level optimizer interface.

* bug fix

* polish code

83b52c56

[gemini] support amp o3 for gemini (#4872) · df635641

Hongxin Liu authored Oct 12, 2023

* [gemini] support no reuse fp16 chunk

* [gemini] support no master weight for optim

* [gemini] support no master weight for gemini ddp

* [test] update gemini tests

* [test] update gemini tests

* [plugin] update gemini plugin

* [test] fix gemini checkpointio test

* [test] fix gemini checkpoint io

df635641

Merge pull request #4889 from ppt0011/main · c1fab951
ppt0011 authored Oct 12, 2023
```
[doc] add reminder for issue encountered with hybrid adam
```
c1fab951

11 Oct, 2023 4 commits

[hotfix] fix bug in sequence parallel test (#4887) · ffd9a3cb
littsk authored Oct 11, 2023

ffd9a3cb
[doc] add reminder for issue encountered with hybrid adam · 1dcaf249
ppt0011 authored Oct 11, 2023

1dcaf249
fix test llama (#4884) · fdec650b
Xu Kai authored Oct 11, 2023

fdec650b

[Pipeline Inference] Sync pipeline inference branch to main (#4820) · 08a9f76b

Bin Jia authored Oct 11, 2023

* [pipeline inference] pipeline inference (#4492)

* add pp stage manager as circle stage

* fix a bug when create process group

* add ppinfer basic framework

* add micro batch manager and support kvcache-pp gpt2 fwd

* add generate schedule

* use mb size to control mb number

* support generate with kv cache

* add output, remove unused code

* add test

* reuse shardformer to build model

* refactor some code and use the same attribute name of hf

* fix review and add test for generation

* remove unused file

* fix CI

* add cache clear

* fix code error

* fix typo

* [Pipeline inference] Modify to tieweight (#4599)

* add pp stage manager as circle stage

* fix a bug when create process group

* add ppinfer basic framework

* add micro batch manager and support kvcache-pp gpt2 fwd

* add generate schedule

* use mb size to control mb number

* support generate with kv cache

* add output, remove unused code

* add test

* reuse shardformer to build model

* refactor some code and use the same attribute name of hf

* fix review and add test for generation

* remove unused file

* modify the way of saving newtokens

* modify to tieweight

* modify test

* remove unused file

* solve review

* add docstring

* [Pipeline inference] support llama pipeline inference (#4647)

* support llama pipeline inference

* remove tie weight operation

* [pipeline inference] Fix the blocking of communication when ppsize is 2 (#4708)

* add benchmark verbose

* fix export tokens

* fix benchmark verbose

* add P2POp style to do p2p communication

* modify schedule as p2p type when ppsize is 2

* remove unused code and add docstring

* [Pipeline inference] Refactor code, add docsting, fix bug (#4790)

* add benchmark script

* update argparse

* fix fp16 load

* refactor code style

* add docstring

* polish code

* fix test bug

* [Pipeline inference] Add pipeline inference docs (#4817)

* add readme doc

* add a ico

* Add performance

* update table of contents

* refactor code (#4873)

08a9f76b

10 Oct, 2023 5 commits
- Update README.md · 652adc22
  Camille Zhong authored Oct 10, 2023
  
  652adc22
- Update README.md · afe10a85
  Camille Zhong authored Oct 10, 2023
  
  afe10a85
- Update main README.md · d6c4b9b3
  Camille Zhong authored Oct 10, 2023
```
add modelscope model link
```
  d6c4b9b3
- Update modelscope link in README.md · 3043d5d6
  Camille Zhong authored Oct 10, 2023
```
add modelscope link
```
  3043d5d6
- [doc] update advanced tutorials, training gpt with hybrid parallelism (#4866) · 6a21f96a
  flybird11111 authored Oct 10, 2023
```
* [doc]update advanced tutorials, training gpt with hybrid parallelism

* [doc]update advanced tutorials, training gpt with hybrid parallelism

* update vit tutorials

* update vit tutorials

* update vit tutorials

* update vit tutorials

* update en/train_vit_with_hybrid_parallel.py

* fix

* resolve comments

* fix
```
  6a21f96a
07 Oct, 2023 5 commits
- [nfc] fix minor typo in README (#4846) · 8aed02b9
  Blagoy Simandoff authored Oct 07, 2023
  
  8aed02b9
- [NFC] polish code style (#4799) · cd6a962e
  Camille Zhong authored Sep 27, 2023
  
  cd6a962e
- [NFC] polish colossalai/inference/quant/gptq/cai_gptq/__init__.py code style (#4792) · 07ed155e
  Michelle authored Sep 27, 2023
  
  07ed155e
- polish code for gptq (#4793) · eef96e08
  littsk authored Sep 25, 2023
  
  eef96e08
- [checkpointio] hotfix torch 2.0 compatibility (#4824) · cb3a25a0
  Hongxin Liu authored Oct 07, 2023
  
  cb3a25a0
06 Oct, 2023 2 commits
- Merge pull request #4856 from KKZ20/test/model_support_for_low_level_zero · ad23460c
  ppt0011 authored Oct 06, 2023
```
[test] remove the redundant code of model output transformation in torchrec
```
  ad23460c
- Merge pull request #4858 from Shawlleyw/main · 81ee91f2
  ppt0011 authored Oct 06, 2023
```
[doc]: typo in document of booster low_level_zero plugin
```
  81ee91f2
05 Oct, 2023 1 commit
- fix: typo in comment of low_level_zero plugin · c97a3523
  shaoyuw authored Oct 05, 2023
  
  c97a3523