Commits · b7141c36dd84d025b4aef09da74c7b7ac29010b5 · OpenDAS / ColossalAI

10 May, 2023 3 commits

[CI] fix some spelling errors (#3707) · b7141c36

digger-yu authored May 10, 2023

* fix spelling error with examples/comminity/

* fix spelling error with tests/

* fix some spelling error with tests/ colossalai/ etc.

b7141c36

[chat] fix community example ray (#3719) · f7361ee1
MisterLin1995 authored May 10, 2023
```
Co-authored-by: jiangwen <zxl265370@antgroup.com>
```
f7361ee1

[booster] add tests for ddp and low level zero's checkpointio (#3715) · 20068ba1

jiangmingyan authored May 10, 2023

* [booster] update tests for booster

* [booster] update tests for booster

* [booster] update tests for booster

* [booster] update tests for booster

* [booster] update tests for booster

* [booster] update booster tutorials#3717, fix recursive check

20068ba1

09 May, 2023 1 commit

[booster] fix no_sync method (#3709) · 6552cbf8

Hongxin Liu authored May 09, 2023

* [booster] fix no_sync method

* [booster] add test for ddp no_sync

* [booster] fix merge

* [booster] update unit test

* [booster] update unit test

* [booster] update unit test

6552cbf8

08 May, 2023 2 commits

[booster] update prepare dataloader method for plugin (#3706) · 3bf09efe
Hongxin Liu authored May 08, 2023
```
* [booster] add prepare dataloader method for plug

* [booster] update examples and docstr
```
3bf09efe

[example] add train resnet/vit with booster example (#3694) · f83ea813

Hongxin Liu authored May 08, 2023

* [example] add train vit with booster example

* [example] update readme

* [example] add train resnet with booster example

* [example] enable ci

* [example] enable ci

* [example] add requirements

* [hotfix] fix analyzer init

* [example] update requirements

f83ea813

06 May, 2023 4 commits
- [tensor] Refactor handle_trans_spec in DistSpecManager · 2629f971
  YH authored May 06, 2023
  
  2629f971
- [chat] fix train_prompts.py gemini strategy bug (#3666) · 2da5d81d
  zhang-yi-chi authored May 06, 2023
```
* fix gemini strategy bug

* add comment

* add comment

* better solution
```
  2da5d81d
- [example] add finetune bert with booster example (#3693) · d5566488
  Hongxin Liu authored May 06, 2023
  
  d5566488
- fix some spelling error with applications/Chat/examples/ (#3692) · 65bdc315
  digger-yu authored May 06, 2023
```
* fix spelling error with examples/comminity/

* fix spelling error with example/
```
  65bdc315
05 May, 2023 6 commits

[booster] refactor all dp fashion plugins (#3684) · d0915f54

Hongxin Liu authored May 05, 2023

* [booster] add dp plugin base

* [booster] inherit dp plugin base

* [booster] refactor unit tests

d0915f54

[CI] Update test_sharded_optim_with_sync_bn.py (#3688) · b49020c1
digger-yu authored May 05, 2023
```
fix spelling error in line23
change "cudnn_determinstic"=True to "cudnn_deterministic=True"
```
b49020c1
Merge pull request #3680 from digger-yu/digger-yu-patch-2 · b36e67cb
Tong Li authored May 05, 2023
```
fix spelling error with applications/Chat/evaluate/
```
b36e67cb

[booster] gemini plugin support shard checkpoint (#3610) · 307894f7

jiangmingyan authored May 05, 2023



* gemini plugin add shard checkpoint save/load

* gemini plugin add shard checkpoint save/load

* gemini plugin add shard checkpoint save/load

* gemini plugin add shard checkpoint save/load

* gemini plugin add shard checkpoint save/load

* gemini plugin add shard checkpoint save/load

* gemini plugin add shard checkpoint save/load

* gemini plugin add shard checkpoint save/load

* gemini plugin add shard checkpoint save/load

* gemini plugin add shard checkpoint save/load

* gemini plugin add shard checkpoint save/load

* gemini plugin add shard checkpoint save/load

* gemini plugin add shard checkpoint save/load

* gemini plugin add shard checkpoint save/load

* gemini plugin support shard checkpoint

* [API Refactoring]gemini plugin support shard checkpoint

* [API Refactoring]gemini plugin support shard checkpoint

* [API Refactoring]gemini plugin support shard checkpoint

* [API Refactoring]gemini plugin support shard checkpoint

* [API Refactoring]gemini plugin support shard checkpoint

* [API Refactoring]gemini plugin support shard checkpoint

* [API Refactoring]gemini plugin support shard checkpoint

* [API Refactoring]gemini plugin support shard checkpoint

* [API Refactoring]gemini plugin support shard checkpoint

* [API Refactoring]gemini plugin support shard checkpoint

* [API Refactoring]gemini plugin support shard checkpoint

* [API Refactoring]gemini plugin support shard checkpoint

* [API Refactoring]gemini plugin support shard checkpoint

---------
Co-authored-by: luchen <luchen@luchendeMBP.lan>
Co-authored-by: luchen <luchen@luchendeMacBook-Pro.local>

307894f7

[chat] PPO stage3 doc enhancement (#3679) · 0f785cb1

Camille Zhong authored May 05, 2023

* Add RoBERTa for RLHF Stage 2 & 3 (test)

RoBERTa for RLHF Stage 2 & 3 (still in testing)

Revert "Add RoBERTa for RLHF Stage 2 & 3 (test)"

This reverts commit 06741d894dcbe958acd4e10d771f22275e20e368.

Add RoBERTa for RLHF stage 2 & 3

1. add roberta folder under model folder
2. add  roberta option in train_reward_model.py
3. add some test in testci

Update test_ci.sh

Revert "Update test_ci.sh"

This reverts commit 9c7352b81766f3177d31eeec0ec178a301df966a.

Add RoBERTa for RLHF Stage 2 & 3 (test)

RoBERTa for RLHF Stage 2 & 3 (still in testing)

Revert "Add RoBERTa for RLHF Stage 2 & 3 (test)"

This reverts commit 06741d894dcbe958acd4e10d771f22275e20e368.

Add RoBERTa for RLHF stage 2 & 3

1. add roberta folder under model folder
2. add  roberta option in train_reward_model.py
3. add some test in testci

Update test_ci.sh

Revert "Update test_ci.sh"

This reverts commit 9c7352b81766f3177d31eeec0ec178a301df966a.

update roberta with coati

chat ci update

Revert "chat ci update"

This reverts commit 17ae7ae01fa752bd3289fc39069868fde99cf846.

* Update README.md

Update README.md

* update readme

* Update test_ci.sh

* update readme and add a script

update readme and add a script

modify readme

Update README.md

0f785cb1

[doc] fix chat spelling error (#3671) · 6650daeb

digger-yu authored May 05, 2023

* Update README.md

change "huggingaface" to "huggingface"

* Update README.md

change "Colossa-AI" to "Colossal-AI"

6650daeb

04 May, 2023 3 commits
- [chat] add opt attn kernel (#3655) · 7bd0bee8
  Hongxin Liu authored May 04, 2023
```
* [chat] add opt attn kernel

* [chat] disable xformer during fwd
```
  7bd0bee8
- Update generate_gpt35_answers.py · 8ba78587
  digger-yu authored May 04, 2023
```
fix spelling error with generate_gpt35_answers.py
```
  8ba78587
- fix spelling error · bfbf6505
  digger-yu authored May 04, 2023
```
fix spelling error with evaluate.py
```
  bfbf6505
28 Apr, 2023 5 commits
- [chat] typo accimulation_steps -> accumulation_steps (#3662) · 1a60dc07
  tanitna authored Apr 28, 2023
  
  1a60dc07
- Merge pull request #3656 from TongLi3701/chat/update_eval · 816add7e
  Tong Li authored Apr 28, 2023
```
[Chat]: Remove unnecessary step and update documentation
```
  816add7e
- [chat] set default zero2 strategy (#3667) · 268b3cd8
  binmakeswell authored Apr 28, 2023
```
* [chat] set default gemini strategy

* [chat] set default zero2 strategy

* [chat] set default zero2 strategy
```
  268b3cd8
- update readme · c1a35594
  Tong Li authored Apr 28, 2023
  
  c1a35594
- update documentation · ed3eaa69
  Tong Li authored Apr 28, 2023
  
  ed3eaa69
27 Apr, 2023 6 commits

update questions and readme · c4191173
Tong Li authored Apr 27, 2023

c4191173
remove unnecessary step and update readme · aa77ddae
Tong Li authored Apr 27, 2023

aa77ddae
[zero] Suggests a minor change to confusing variable names in the ZeRO optimizer. (#3173) · a22407cc
YH authored Apr 27, 2023
```
* Fix confusing variable name in zero opt

* Apply lint

* Fix util func

* Fix minor util func

* Fix zero param optimizer name
```
a22407cc

[chat] refactor model save/load logic (#3654) · 842768a1

Hongxin Liu authored Apr 27, 2023

* [chat] strategy refactor unwrap model

* [chat] strategy refactor save model

* [chat] add docstr

* [chat] refactor trainer save model

* [chat] fix strategy typing

* [chat] refactor trainer save model

* [chat] update readme

* [chat] fix unit test

842768a1

[chat] remove lm model class (#3653) · 6ef70114

Hongxin Liu authored Apr 27, 2023

* [chat] refactor lora

* [chat] remove lm class

* [chat] refactor save model

* [chat] refactor train sft

* [chat] fix ci

* [chat] fix ci

6ef70114

[Doc] enhancement on README.md for chat examples (#3646) · 8bccb72c

Camille Zhong authored Apr 27, 2023

* Add RoBERTa for RLHF Stage 2 & 3 (test)

RoBERTa for RLHF Stage 2 & 3 (still in testing)

Revert "Add RoBERTa for RLHF Stage 2 & 3 (test)"

This reverts commit 06741d894dcbe958acd4e10d771f22275e20e368.

Add RoBERTa for RLHF stage 2 & 3

1. add roberta folder under model folder
2. add  roberta option in train_reward_model.py
3. add some test in testci

Update test_ci.sh

Revert "Update test_ci.sh"

This reverts commit 9c7352b81766f3177d31eeec0ec178a301df966a.

Add RoBERTa for RLHF Stage 2 & 3 (test)

RoBERTa for RLHF Stage 2 & 3 (still in testing)

Revert "Add RoBERTa for RLHF Stage 2 & 3 (test)"

This reverts commit 06741d894dcbe958acd4e10d771f22275e20e368.

Add RoBERTa for RLHF stage 2 & 3

1. add roberta folder under model folder
2. add  roberta option in train_reward_model.py
3. add some test in testci

Update test_ci.sh

Revert "Update test_ci.sh"

This reverts commit 9c7352b81766f3177d31eeec0ec178a301df966a.

update roberta with coati

chat ci update

Revert "chat ci update"

This reverts commit 17ae7ae01fa752bd3289fc39069868fde99cf846.

* Update README.md

Update README.md

* update readme

* Update test_ci.sh

8bccb72c

26 Apr, 2023 5 commits

[chat] refactor trainer (#3648) · 2a951955

Hongxin Liu authored Apr 26, 2023

* [chat] ppo trainer remove useless args

* [chat] update examples

* [chat] update benchmark

* [chat] update examples

* [chat] fix sft training with wandb

* [chat] polish docstr

2a951955

[chat] polish performance evaluator (#3647) · f8288315
Hongxin Liu authored Apr 26, 2023

f8288315

[gemini] accelerate inference (#3641) · 50793b35

Hongxin Liu authored Apr 26, 2023

* [gemini] support don't scatter after inference

* [chat] update colossalai strategy

* [chat] fix opt benchmark

* [chat] update opt benchmark

* [gemini] optimize inference

* [test] add gemini inference test

* [chat] fix unit test ci

* [chat] fix ci

* [chat] fix ci

* [chat] skip checkpoint test

50793b35

[booster] add low level zero plugin (#3594) · 4b3240cb

Hongxin Liu authored Apr 26, 2023

* [booster] add low level zero plugin

* [booster] fix gemini plugin test

* [booster] fix precision

* [booster] add low level zero plugin test

* [test] fix booster plugin test oom

* [test] fix booster plugin test oom

* [test] fix googlenet and inception output trans

* [test] fix diffuser clip vision model

* [test] fix torchaudio_wav2vec2_base

* [test] fix low level zero plugin test

4b3240cb

[doc] Fix typo under colossalai and doc(#3618) · b9a8dff7

digger-yu authored Apr 26, 2023

* Fixed several spelling errors under colossalai

* Fix the spelling error in colossalai and docs directory

* Cautious Changed the spelling error under the example folder

* Update runtime_preparation_pass.py

revert autograft to autograd

* Update search_chunk.py

utile to until

* Update check_installation.py

change misteach to mismatch in line 91

* Update 1D_tensor_parallel.md

revert to perceptron

* Update 2D_tensor_parallel.md

revert to perceptron in line 73

* Update 2p5D_tensor_parallel.md

revert to perceptron in line 71

* Update 3D_tensor_parallel.md

revert to perceptron in line 80

* Update README.md

revert to resnet in line 42

* Update reorder_graph.py

revert to indice in line 7

* Update p2p.py

revert to megatron in line 94

* Update initialize.py

revert to torchrun in line 198

* Update routers.py

change to detailed in line 63

* Update routers.py

change to detailed in line 146

* Update README.md

revert  random number in line 402

b9a8dff7

24 Apr, 2023 3 commits
- Merge pull request #3621 from zhang-yi-chi/fix/chat-train-prompts-single-gpu · e1b0a78a
  Tong Li authored Apr 24, 2023
```
[chat] fix single gpu training bug in examples/train_prompts.py
```
  e1b0a78a
- [Chat] Remove duplicate functions (#3625) · df309fc6
  ddobokki authored Apr 24, 2023
  
  df309fc6
- [devops] fix chat ci (#3628) · 179558a8
  Hongxin Liu authored Apr 24, 2023
  
  179558a8
22 Apr, 2023 1 commit
- [chat] fix enable single gpu training bug · 739cfe33
  zhang-yi-chi authored Apr 22, 2023
  
  739cfe33
20 Apr, 2023 1 commit
- [chat] polish code note typo (#3612) · d7bf2847
  digger-yu authored Apr 20, 2023
  
  d7bf2847