Commits · c0d1c330220e055d5ff966e3e9b517b5e9dd4802 · chenpangpang / transformers

24 Jul, 2023 16 commits

[i18n-KO] Translated `perf_train_cpu.md` to Korean (#24911) · c0d1c330

seank021 authored Jul 25, 2023



* dos: ko: perf_train_cpu.md

* feat: chatgpt draft

* fix: manual edits

* fix: resolve suggestions

* fix: manual edits
Co-authored-by: Haewon Kim <ehdvkf02@naver.com>

---------
Co-authored-by: Haewon Kim <ehdvkf02@naver.com>

c0d1c330

[`8bit`] Fix 8bit corner case with Blip2 8bit (#25047) · b08f41e6
Younes Belkada authored Jul 24, 2023
```
fix 8bit corner case with Blip2 8bit
```
b08f41e6

compute_loss in trainer failing to label shift for PEFT model when label... · 3611fc90

Nate Brake authored Jul 24, 2023


compute_loss in trainer failing to label shift for PEFT model when label smoothing enabled. (#25044)

* added PeftModelForCausalLM to MODEL_FOR_CAUSAL_LM_MAPPING_NAMES dict

* check for PEFT model in compute_loss section

---------
Co-authored-by: Nathan Brake <nbrake3@mmm.com>

3611fc90

Pvt model (#24720) · a03d13c8

Rinat authored Jul 24, 2023

* pull and push updates

* add docs

* fix modeling

* Add and run test

* make copies

* add task

* fix tests and fix small issues

* Checks on a Pull Request

* fix docs

* add desc pvt.md

a03d13c8

Comment again print statement · afe8bfc0
Sylvain Gugger authored Jul 24, 2023

afe8bfc0
Make more test models smaller (#25005) · 42571f6e
Sylvain Gugger authored Jul 24, 2023
```
* Make more test models tiny

* Make more test models tiny

* More models

* More models
```
42571f6e
Fix typo in LlamaTokenizerFast docstring example (#25018) · 8f1f0bf5
Sören Brunk authored Jul 24, 2023

8f1f0bf5
Add dispatch_batches to training arguments (#25038) · 3b734f50
Zach Mueller authored Jul 24, 2023
```
* Dispatch batches

* Copy items
```
3b734f50

🌐

[i18n-KO] Translated `testing.md` to Korean (#24900) · 9d2b983e

Sunmin Cho authored Jul 24, 2023

* docs: ko: testing.md

* feat: draft

* fix: manual edits

* fix: edit ko/_toctree.yml

* fix: manual edits

* fix: manual edits

* fix: manual edits

* fix: manual edits

* fix: resolve suggestions

9d2b983e

🌐

[i18n-KO] Translated performance.md to Korean (#24883) · 383be1b7

Sangam Lee authored Jul 24, 2023



* dos: ko: performance.md

* feat: chatgpt draft

* fix: manual edits

* fix: manual edits

* Update docs/source/ko/performance.md
Co-authored-by: Kihoon Son <75935546+kihoon71@users.noreply.github.com>

* Update docs/source/ko/performance.md

---------
Co-authored-by: Kihoon Son <75935546+kihoon71@users.noreply.github.com>

383be1b7

Better handling missing SYS in llama conversation tokenizer (#24997) · efb2ba66

Iskren Ivov Chernev authored Jul 24, 2023

* Better handling missing SYS in llama conversation tokenizer

The existing code failed to add SYS if the conversation has history
without SYS, but did modify the passed conversation as it did.

Rearrange the code so modification to the conversation object are taken
into account for token id generation.

* Fix formatting with black

* Avoid one-liners

* Also fix fast tokenizer

* Drop List decl

efb2ba66

Support GatedRepoError + use raise from (#25034) · 67049231

Lucain authored Jul 24, 2023



* Support GatedRepoError + use raise from

* Apply suggestions from code review
Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com>

* Use token instead of use_auth_token in error messages

---------
Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com>

67049231

[docs] Performance docs tidy up, part 1 (#23963) · 75317aef

Maria Khalusova authored Jul 24, 2023



* first pass at the single gpu doc

* overview: improved clarity and navigation

* WIP

* updated intro and deepspeed sections

* improved torch.compile section

* more improvements

* minor improvements

* make style

* Apply suggestions from code review
Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com>

* feedback addressed

* mdx -> md

* link fix

* feedback addressed

---------
Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com>

75317aef

fix(integrations): store serialized `TrainingArgs` to `wandb.config` without sanitization. (#25035) · 54ba8608

Bharat Ramanathan authored Jul 24, 2023

fix: store training args to wandb config without sanitization.

Allows resuming runs by reusing the wandb config.
Co-authored-by: Bharat Ramanathan <ramanathan.parameshwaran@gohuddl.com>

54ba8608

[`logging.py`] set default `stderr` path if `None` (#25033) · 0906d212
Arthur authored Jul 24, 2023
```
set default logger
```
0906d212
[check_config_docstrings.py] improve diagnostics (#25012) · c9a82be5
Stas Bekman authored Jul 23, 2023
```
* [check_config_docstrings.py] improve diagnostics

* style

* rephrase

* fix
```
c9a82be5

21 Jul, 2023 16 commits

🌐 [i18n-KO] Updated Korean `serialization.md` (#24686) · b257c46a
Wonhyeong Seo authored Jul 22, 2023
```
fix: update ko/serialization.md

* chatgpt draft
```
b257c46a
Move template doc file to md (#25004) · 87fba947
Sylvain Gugger authored Jul 21, 2023

87fba947

improve from_pretrained for zero3 multi gpus mode (#24964) · ea41e18c

Ivan Sorokin authored Jul 21, 2023



* improve from_pretrained for zero3 multi gpus mode

* Add check if torch.distributed.is_initialized

* Revert torch.distributed

---------
Co-authored-by: Stas Bekman <stas@stason.org>

ea41e18c

[`Llama`] remove persistent `inv_freq` tensor (#24998) · 95f96b45
Arthur authored Jul 21, 2023
```
remove persistent tensor
```
95f96b45
[`bnb`] Add simple check for bnb import (#24995) · d3ce048c
Younes Belkada authored Jul 21, 2023
```
add simple check for bnb
```
d3ce048c
Fix `llama` tokenization doctest (#24990) · f1a1eb4a
Yih-Dar authored Jul 21, 2023
```
fix
Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>
```
f1a1eb4a
Use main_input_name for include_inputs_for_metrics (#24993) · a7d21318
Sylvain Gugger authored Jul 21, 2023

a7d21318
Fix type annotation for deepspeed training arg (#24988) · a6484c89
Sylvain Gugger authored Jul 21, 2023

a6484c89
Avoid importing all models when instantiating a pipeline (#24960) · 5b7ffd54
Sylvain Gugger authored Jul 21, 2023
```
* Avoid importing all models when instantiating a pipeline

* Remove sums that don't work
```
5b7ffd54
Remove tokenizers from the doc table (#24963) · 640e1b6c
Sylvain Gugger authored Jul 21, 2023

640e1b6c
[`LlamaConfig`] Nit: pad token should be None by default (#24958) · 0511369a
Arthur authored Jul 21, 2023
```
* pad token should be None by default

* fix tests

* nits
```
0511369a

Fix missing spaces in system prompt of Llama2 tokenizer (#24930) · f74560d0

Joya Chen authored Jul 21, 2023



* Update tokenization_llama.py

* Update tokenization_llama_fast.py

* Update src/transformers/models/llama/tokenization_llama_fast.py
Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com>

* Update src/transformers/models/llama/tokenization_llama.py
Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com>

* Update src/transformers/models/llama/tokenization_llama.py
Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com>

* Update src/transformers/models/llama/tokenization_llama_fast.py
Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com>

---------
Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com>

f74560d0

fsdp fixes and enhancements (#24980) · f4eb459e

Sourab Mangrulkar authored Jul 21, 2023

* fix fsdp prepare to remove the warnings and fix excess memory usage

* Update training_args.py

* parity for FSDP+XLA

* Update trainer.py

f4eb459e

🌐

[i18n-KO] Fixed Korean and English `quicktour.md` (#24664) · ec3dfe5e

Wonhyeong Seo authored Jul 21, 2023



* fix: english/korean quicktour.md

* fix: resolve suggestions
Co-authored-by: Hyeonseo Yun <0525yhs@gmail.com>
Co-authored-by: Sohyun Sim <96299403+sim-so@users.noreply.github.com>
Co-authored-by: Kihoon Son <75935546+kihoon71@users.noreply.github.com>

* fix: follow glossary

* 파인튜닝 -> 미세조정

---------
Co-authored-by: Hyeonseo Yun <0525yhs@gmail.com>
Co-authored-by: Sohyun Sim <96299403+sim-so@users.noreply.github.com>
Co-authored-by: Kihoon Son <75935546+kihoon71@users.noreply.github.com>

ec3dfe5e

fix: cast input pixels to appropriate dtype for image_to_text pipelines (#24947) · 83f9314d

Jim Allanson authored Jul 21, 2023

* fix: cast input pixels to appropriate dtype for image_to_text tasks

* fix: add casting to pixel inputs of additional models after running copy checks

83f9314d

fix fsdp checkpointing issues (#24926) · 1c7e5e23
Sourab Mangrulkar authored Jul 21, 2023
```
* fix fsdp load

* Update trainer.py

* remove saving duplicate state_dict
```
1c7e5e23

20 Jul, 2023 8 commits

Fallback for missing attribute `Parameter.ds_numel` (#24942) · 9ef5256d
Apoorv Khandelwal authored Jul 20, 2023
```
* [trainer] fallback for deepspeed param count

* [trainer] more readable numel count
```
9ef5256d
Contrastive Search peak memory reduction (#24120) · caf5e369
Benjamin Badger authored Jul 20, 2023
```
Co-authored-by: Joao Gante <joaofranciscocardosogante@gmail.com>
```
caf5e369
Change logic for logging in the examples (#24956) · aa1b09c5
Zach Mueller authored Jul 20, 2023
```
Change logic
```
aa1b09c5
[`RWKV`] Add Gradient Checkpointing support for RWKV (#24955) · 89a1f342
Younes Belkada authored Jul 20, 2023
```
add GC support for RWKV
```
89a1f342

Bump aiohttp from 3.8.1 to 3.8.5 in /examples/research_projects/decision_transformer (#24954) · 9f912ef6

dependabot[bot] authored Jul 20, 2023

Bump aiohttp in /examples/research_projects/decision_transformer

Bumps [aiohttp](https://github.com/aio-libs/aiohttp) from 3.8.1 to 3.8.5.
- [Release notes](https://github.com/aio-libs/aiohttp/releases)
- [Changelog](https://github.com/aio-libs/aiohttp/blob/v3.8.5/CHANGES.rst)
- [Commits](https://github.com/aio-libs/aiohttp/compare/v3.8.1...v3.8.5

)

---
updated-dependencies:
- dependency-name: aiohttp
  dependency-type: direct:production
...
Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>

9f912ef6

fix type annotations for arguments in training_args (#24550) · e75cb0cb

Shauray Singh authored Jul 20, 2023

* testing

* example script

* fix typehinting

* some tests

* make test

* optional update

* Union of arguments

* does this fix the issue

* remove reports

* set default to False

* documentation change

* None support

* does not need None

* Fix typing annotations for FSDP and DeepSpeed in TrainingArguments (#24549)

* Fix typing annotations for FSDP and DeepSpeed in TrainingArguments

* Change dict to Dict

* Revert "Fix typing annotations for FSDP and DeepSpeed in TrainingArguments" (#24574)

Revert "Fix typing annotations for FSDP and DeepSpeed in TrainingArguments (#24549)"

This reverts commit c5e29d43

.

* Fix typing annotations for FSDP and DeepSpeed in TrainingArguments (#24549)

* Fix typing annotations for FSDP and DeepSpeed in TrainingArguments

* Change dict to Dict

* merge

* hacky fix

* fixup

---------
Co-authored-by: Max Ryabinin <mryabinin0@gmail.com>
Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com>

e75cb0cb

[DOCS] Example for `LogitsProcessor` class (#24848) · 0c41765d

Shauray Singh authored Jul 20, 2023

* make docs

* fixup

* resolved

* remove debugs

* Revert "fixup"

This reverts commit 5e0f636aae0bf8707bc8bdaa6a9427fbf66834ed.

* prev (ignore)

* fixup broke some files

* remove files

* reverting modeling_reformer

* lang fix

0c41765d

Fix `main_input_name` in `src/transformers/keras_callbacks.py` (#24916) · 35c04596
Yih-Dar authored Jul 20, 2023
```
fix
Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>
```
35c04596