Commits · 6e7e43c6fe2d38aa91e57150310708558855da89 · OpenDAS / ColossalAI

17 Apr, 2023 8 commits

[doc] Update .github/workflows/README.md (#3577) · 6e7e43c6
digger-yu authored Apr 17, 2023
```
Optimization Code
I think there were two extra $ entered here, which have been deleted
```
6e7e43c6
[coati] add costom model suppor tguide (#3579) · 6b1a39b1
Fazzie-Maqianli authored Apr 17, 2023

6b1a39b1
[chat] update reward model sh (#3578) · cc1eec2f
binmakeswell authored Apr 17, 2023

cc1eec2f

[chatgpt] Detached PPO Training (#3195) · e3551443

csric authored Apr 17, 2023



* run the base

* working on dist ppo

* sync

* detached trainer

* update detached trainer. no maker update function

* facing init problem

* 1 maker 1 trainer detached run. but no model update

* facing cuda problem

* fix save functions

* verified maker update

* nothing

* add ignore

* analyize loss issue

* remove some debug codes

* facing 2m1t stuck issue

* 2m1t verified

* do not use torchrun

* working on 2m2t

* working on 2m2t

* initialize strategy in ray actor env

* facing actor's init order issue

* facing ddp model update issue (need unwarp ddp)

* unwrap ddp actor

* checking 1m2t stuck problem

* nothing

* set timeout for trainer choosing. It solves the stuck problem!

* delete some debug output

* rename to sync with upstream

* rename to sync with upstream

* coati rename

* nothing

* I am going to detach the replaybuffer from trainer and make it a Ray Actor. Two benefits: 1. support TP trainer. 2. asynchronized buffer operations

* experience_maker_holder performs target-revolving _send_experience() instead of length comparison.

* move code to ray subfolder

* working on pipeline inference

* apply comments

---------
Co-authored-by: csric <richcsr256@gmail.com>

e3551443

Add docstr for zero3 chunk search utils (#3572) · d329c294
YH authored Apr 17, 2023

d329c294
[doc] Update 1D_tensor_parallel.md (#3573) · 9edeadfb
digger-yu authored Apr 17, 2023
```
Display format optimization , same as fix#3562
Simultaneous modification of en version
```
9edeadfb

[misc] add verbose arg for zero and op builder (#3552) · 173dad05

Hongxin Liu authored Apr 17, 2023

* [misc] add print verbose

* [gemini] add print verbose

* [zero] add print verbose for low level

* [misc] add print verbose for op builder

173dad05

[lazyinit] fix clone and deepcopy (#3553) · 4341f5e8
Hongxin Liu authored Apr 17, 2023

4341f5e8

14 Apr, 2023 2 commits

[doc] Update 1D_tensor_parallel.md (#3563) · 1c7734bc

digger-yu authored Apr 14, 2023

Display format optimization, fix bug#3562
Specific changes
1. "This is called a column-parallel fashion" Translate to Chinese
2. use the ```math code block syntax to display a math expression as a block, No modification of formula content

Please check that the math formula is displayed correctly
If OK, I will change the format of the English version of the formula in parallel

1c7734bc

[example] reorganize for community examples (#3557) · f1b3d60c
binmakeswell authored Apr 14, 2023

f1b3d60c

13 Apr, 2023 4 commits

[chat] ChatGPT train prompts on ray example (#3309) · 1a809edd

MisterLin1995 authored Apr 13, 2023



* [feat][chatgpt]train prompts on ray example

* [fix]simplify code

* [fix]remove depreciated parameter

* [fix]add dependencies

* [fix]method calling

* [fix]experience maker

* [fix]missing loss function

* [fix]init optimizer

* [feat]add usage comment

* [fix]rename files

* [fix]add readme

* [fix]file path

* [fix]move directory

---------
Co-authored-by: jiangwen <zxl265370@antgroup.com>

1a809edd

[chat] polish tutorial doc (#3551) · 535b8964

binmakeswell authored Apr 13, 2023

* [chat] clean up duplicate tutorial

* [chat] clean up duplicate tutorial

* [chat] clean up duplicate tutorial

* [chat] clean up duplicate tutorial

535b8964

[doc] Update README.md (#3549) · 77efdfe1
digger-yu authored Apr 13, 2023
```
Format Optimization ,Add [] outside of DeepSpeed
```
77efdfe1
Update README.md (#3548) · 3f760da9
digger-yu authored Apr 13, 2023
```
Delete more ")"
```
3f760da9

12 Apr, 2023 5 commits

[doc] Update README-zh-Hans.md (#3541) · a3ac48ef
digger-yu authored Apr 12, 2023
```
Fixing document link errors using absolute paths
```
a3ac48ef
Polish Code · de84c031
natalie_cao authored Apr 11, 2023

de84c031

[gemini] gemini supports lazy init (#3379) · 152239bb

Hongxin Liu authored Apr 12, 2023

* [gemini] fix nvme optimizer init

* [gemini] gemini supports lazy init

* [gemini] add init example

* [gemini] add fool model

* [zero] update gemini ddp

* [zero] update init example

* add chunk method

* add chunk method

* [lazyinit] fix lazy tensor tolist

* [gemini] fix buffer materialization

* [misc] remove useless file

* [booster] update gemini plugin

* [test] update gemini plugin test

* [test] fix gemini plugin test

* [gemini] fix import

* [gemini] fix import

* [lazyinit] use new metatensor

* [lazyinit] use new metatensor

* [lazyinit] fix __set__ method

152239bb

[checkpoint] Shard saved checkpoint need to be compatible with the naming... · 366a0355

jiangmingyan authored Apr 12, 2023


[checkpoint]  Shard saved checkpoint need to be compatible with the naming format of hf checkpoint files  (#3479)

* [checkpoint] support huggingface style sharded checkpoint, to be compatible with hf file naming format

* [checkpoint] support huggingface style sharded checkpoint, to be compatible with hf file naming format

* [checkpoint] Shard saved checkpoint add 'variant' field to customize filename

* [checkpoint] Shard saved checkpoint add 'variant' field to customize filename

* [checkpoint] Shard saved checkpoint add 'variant' field to customize filename

* [checkpoint] Shard saved checkpoint add 'variant' field to customize filename

---------
Co-authored-by: luchen <luchen@luchendeMacBook-Pro.local>
Co-authored-by: luchen <luchen@luchendeMBP.lan>

366a0355

[chat]add examples of training with limited resources in chat readme (#3536) · 7182ac2a
Yuanchen authored Apr 12, 2023
```
Co-authored-by: Yuanchen Xu <yuanchen.xu00@gmail.com>
```
7182ac2a

11 Apr, 2023 1 commit
- [chat]: add vf_coef argument for PPOTrainer (#3318) · e6a132a4
  zhang-yi-chi authored Apr 11, 2023
  
  e6a132a4
10 Apr, 2023 5 commits
- [chat] add zero2 cpu strategy for sft training (#3520) · 89fd10a1
  ver217 authored Apr 10, 2023
  
  89fd10a1
- [doc] hide diffusion in application path (#3519) · 990d4c3e
  binmakeswell authored Apr 10, 2023
```
- [ ] Stable Diffusion
- [ ] Dreambooth
It's easy for users to think that we don't support them yet. Add them after migrating them from example to application
https://github.com/hpcaitech/ColossalAI/tree/main/examples/images
```
  990d4c3e
- [doc] add requirement and highlight application (#3516) · 0c045570
  binmakeswell authored Apr 10, 2023
```
* [doc] add requirement and highlight application

* [doc] link example and application
```
  0c045570
- [Chat Community] Update README.md (fixed#3487) (#3506) · 635d0a1b
  NatalieC323 authored Apr 10, 2023
```
* Update README.md

* Update README.md

* Update README.md

* Update README.md

---------
Co-authored-by: Fazzie-Maqianli <55798671+Fazziekey@users.noreply.github.com>
```
  635d0a1b
- [doc] Add docs for clip args in zero optim (#3504) · bcf0cbcb
  YH authored Apr 10, 2023
  
  bcf0cbcb
07 Apr, 2023 3 commits
- [coati] Fix LlamaCritic (#3475) · a7ca2972
  gongenlei authored Apr 07, 2023
```
* mv LlamaForCausalLM to LlamaModel

* rm unused imports

---------
Co-authored-by: gongenlei <gongenlei@baidu.com>
```
  a7ca2972
- [example] remove redundant texts & update roberta (#3493) · 8f2c55f9
  mandoxzhang authored Apr 07, 2023
```
* update roberta example

* update roberta example

* modify conflict & update roberta
```
  8f2c55f9
- [example] update roberta with newer ColossalAI (#3472) · ab5fd127
  mandoxzhang authored Apr 07, 2023
```
* update roberta example

* update roberta example
```
  ab5fd127
06 Apr, 2023 12 commits

Revert "[dreambooth] fixing the incompatibity in requirements.txt (#3190) (#3378)" (#3481) · fb8fae6f
NatalieC323 authored Apr 06, 2023

fb8fae6f
[chat] fix stage3 PPO sample sh command (#3477) · 891b8e7f
binmakeswell authored Apr 06, 2023

891b8e7f

[dreambooth] fixing the incompatibity in requirements.txt (#3190) (#3378) · c701b77b

NatalieC323 authored Apr 06, 2023

* Update requirements.txt

* Update environment.yaml

* Update README.md

* Update environment.yaml

* Update README.md

* Update README.md

* Delete requirements_colossalai.txt

* Update requirements.txt

* Update README.md

c701b77b

[doc] updated contributor list (#3474) · 4e998934
Frank Lee authored Apr 06, 2023

4e998934

[checkpoint] support huggingface style sharded checkpoint (#3461) · 52a933e1

jiangmingyan authored Apr 06, 2023



* [checkpoint] support huggingface style sharded checkpoint

* [checkpoint] support huggingface style sharded checkpoint

* [checkpoint] support huggingface style sharded checkpoint

* [checkpoint] support huggingface style sharded checkpoint

* [checkpoint] support huggingface style sharded checkpoint

---------
Co-authored-by: luchen <luchen@luchendeMBP.lan>

52a933e1

add community example dictionary (#3465) · 6afeb120
Fazzie-Maqianli authored Apr 06, 2023

6afeb120

[test] refactor tests with spawn (#3452) · 80eba05b

Frank Lee authored Apr 06, 2023

* [test] added spawn decorator

* polish code

* polish code

* polish code

* polish code

* polish code

* polish code

80eba05b

[Chat]Add Peft support & fix the ptx bug (#3433) · 62f4e2eb

YY Lin authored Apr 06, 2023

* Update ppo.py

Fix the bug of fetching wrong batch data

* Add peft model support in SFT and Prompts training

In stage-1 and stage-3, the peft model supports are added. So the trained artifacts will be only a small lora additions instead of the whole bunch of files.

* Delete test_prompts.txt

* Delete test_pretrained.txt

* Move the peft stuffs to a community folder.

* Move the demo sft to community

* delete dirty files

* Add instructions to install peft using source

* Remove Chinese comments

* remove the Chinese comments

62f4e2eb

[chat]fix save_model(#3377) · 73afb635
Dr-Corgi authored Apr 06, 2023
```
The function save_model should be a part of PPOTrainer.
```
73afb635
[chat]fix readme (#3429) · 57a3c4db
kingkingofall authored Apr 06, 2023
```
* fix stage 2

fix stage 2

* add torch
```
57a3c4db
[booster] fixed the torch ddp plugin with the new checkpoint api (#3442) · 7d8d8256
Frank Lee authored Apr 06, 2023

7d8d8256
Fix typo (#3448) · 8f740deb
YH authored Apr 06, 2023

8f740deb