Commits · 9aeacb58bab321bc21c24bbdf7a24efdccb1d426 · chenpangpang / transformers

08 Oct, 2020 1 commit

Adding Fast tokenizers for SentencePiece based tokenizers - Breaking: remove... · 9aeacb58

Thomas Wolf authored Oct 08, 2020


Adding Fast tokenizers for SentencePiece based tokenizers - Breaking: remove Transfo-XL fast tokenizer (#7141)

* [WIP] SP tokenizers

* fixing tests for T5

* WIP tokenizers

* serialization

* update T5

* WIP T5 tokenization

* slow to fast conversion script

* Refactoring to move tokenzier implementations inside transformers

* Adding gpt - refactoring - quality

* WIP adding several tokenizers to the fast world

* WIP Roberta - moving implementations

* update to dev4 switch file loading to in-memory loading

* Updating and fixing

* advancing on the tokenizers - updating do_lower_case

* style and quality

* moving forward with tokenizers conversion and tests

* MBart, T5

* dumping the fast version of transformer XL

* Adding to autotokenizers + style/quality

* update init and space_between_special_tokens

* style and quality

* bump up tokenizers version

* add protobuf

* fix pickle Bert JP with Mecab

* fix newly added tokenizers

* style and quality

* fix bert japanese

* fix funnel

* limite tokenizer warning to one occurence

* clean up file

* fix new tokenizers

* fast tokenizers deep tests

* WIP adding all the special fast tests on the new fast tokenizers

* quick fix

* adding more fast tokenizers in the fast tests

* all tokenizers in fast version tested

* Adding BertGenerationFast

* bump up setup.py for CI

* remove BertGenerationFast (too early)

* bump up tokenizers version

* Clean old docstrings

* Typo

* Update following Lysandre comments
Co-authored-by: Sylvain Gugger <sylvain.gugger@gmail.com>

9aeacb58

15 Jun, 2020 1 commit

[HUGE] Refactoring tokenizers backend - padding - truncation - pre-tokenized... · 36434220

Anthony MOI authored Jun 15, 2020


[HUGE] Refactoring tokenizers backend - padding - truncation - pre-tokenized pipeline - fast tokenizers - tests (#4510)

* Use tokenizers pre-tokenized pipeline

* failing pretrokenized test

* Fix is_pretokenized in python

* add pretokenized tests

* style and quality

* better tests for batched pretokenized inputs

* tokenizers clean up - new padding_strategy - split the files

* [HUGE] refactoring tokenizers - padding - truncation - tests

* style and quality

* bump up requied tokenizers version to 0.8.0-rc1

* switched padding/truncation API - simpler better backward compat

* updating tests for custom tokenizers

* style and quality - tests on pad

* fix QA pipeline

* fix backward compatibility for max_length only

* style and quality

* Various cleans up - add verbose

* fix tests

* update docstrings

* Fix tests

* Docs reformatted

* __call__ method documented
Co-authored-by: Thomas Wolf <thomwolf@users.noreply.github.com>
Co-authored-by: Lysandre <lysandre.debut@reseau.eseo.fr>

36434220

15 Jan, 2020 1 commit
- 💄 super · 83a41d39
  Julien Chaumond authored Jan 15, 2020
  
  83a41d39
06 Jan, 2020 2 commits
- GPU text generation: mMoved the encoded_prompt to correct device · 81d6841b
  alberduris authored Dec 31, 2019
  
  81d6841b
- Moved the encoded_prompts to correct device · dd4df80f
  alberduris authored Dec 31, 2019
  
  dd4df80f
22 Dec, 2019 8 commits
- Use built-in open(). · 1c62e87b
  Aymeric Augustin authored Dec 22, 2019
```
On Python 3, `open is io.open`.
```
  1c62e87b
- Remove __future__ imports. · c824d15a
  Aymeric Augustin authored Dec 22, 2019
  
  c824d15a
- Replace CommonTestCases for tokenizers with a mixin. · 00204f2b
  Aymeric Augustin authored Dec 22, 2019
```
This is the same change as for (TF)CommonTestCases for modeling.
```
  00204f2b
- Rename file for consistency. · a3c5883f
  Aymeric Augustin authored Dec 22, 2019
  
  a3c5883f
- Remove unittest.main() in test modules. · 7e98e211
  Aymeric Augustin authored Dec 22, 2019
```
This construct isn't used anymore these days.

Running python tests/test_foo.py puts the tests/ directory on
PYTHONPATH, which isn't representative of how we run tests.

Use python -m unittest tests/test_foo.py instead.
```
  7e98e211
- Switch test files to the standard test_*.py scheme. · ced0a942
  Aymeric Augustin authored Dec 22, 2019
  
  ced0a942
- Move tests outside of library. · 067395d5
  Aymeric Augustin authored Dec 22, 2019
  
  067395d5
- Sort imports with isort. · 158e82e0
  Aymeric Augustin authored Dec 21, 2019
```
This is the result of:

    $ isort --recursive examples templates transformers utils hubconf.py setup.py
```
  158e82e0
21 Dec, 2019 1 commit

Reformat source code with black. · fa84ae26

Aymeric Augustin authored Dec 21, 2019

This is the result of:

    $ black --line-length 119 examples templates transformers utils hubconf.py setup.py

There's a lot of fairly long lines in the project. As a consequence, I'm
picking the longest widely accepted line length, 119 characters.

This is also Thomas' preference, because it allows for explicit variable
names, to make the code easier to understand.

fa84ae26

08 Oct, 2019 1 commit
- fix tokenization · 24831477
  thomwolf authored Oct 08, 2019
  
  24831477
04 Oct, 2019 1 commit

Adding CTRL (squashed commit) · dbed1c5d

keskarnitish authored Sep 30, 2019

adding conversion script

adding first draft of modeling & tokenization

adding placeholder for test files

bunch of changes

registering the tokenizer/model/etc

tests

change link; something is very VERY wrong here

weird end-of-word thingy going on

i think the tokenization works now ; wrote the unit tests

overall structure works;load w next

the monster is alive!

works after some cleanup as well

adding emacs autosave to gitignore

currently only supporting the 48 layer one; seems to infer fine on my macbook

cleanup

fixing some documentation

fixing some documentation

tests passing?

now works on CUDA also

adding greedy?

adding greedy sampling

works well

dbed1c5d

26 Sep, 2019 2 commits
- [BIG] pytorch-transformers => transformers · 31c23bd5
  thomwolf authored Sep 26, 2019
  
  31c23bd5
- fix tokenization tests for gpt2 roberta · f2a337b3
  thomwolf authored Sep 26, 2019
  
  f2a337b3
30 Aug, 2019 5 commits
- added test and debug tokenizer configuration serialization · 69da972a
  thomwolf authored Aug 30, 2019
  
  69da972a
- python2 doesn't spark joy · ce5ef4b3
  thomwolf authored Aug 30, 2019
  
  ce5ef4b3
- clean up all byte-level bpe tests · 5dd7b677
  thomwolf authored Aug 30, 2019
  
  5dd7b677
- fix for python2 · ca1a00a3
  thomwolf authored Aug 30, 2019
  
  ca1a00a3
- fix GPT-2 and RoBERTa tests to be clean now · abe734ca
  thomwolf authored Aug 30, 2019
  
  abe734ca
05 Aug, 2019 1 commit
- cleaning up tokenizer tests structure (at last) - last remaining ppb refs · 328afb70
  thomwolf authored Aug 05, 2019
  
  328afb70
15 Jul, 2019 1 commit
- update tokenizer - update squad example for xlnet · 15d8b126
  thomwolf authored Jul 15, 2019
  
  15d8b126
09 Jul, 2019 2 commits
- fix python 2 tests · c079d7dd
  thomwolf authored Jul 09, 2019
  
  c079d7dd
- unified tokenizer api and serialization + tests · b1978698
  thomwolf authored Jul 09, 2019
  
  b1978698
05 Jul, 2019 3 commits
- tokenization abstract class - tests for examples · 36bca545
  thomwolf authored Jul 05, 2019
  
  36bca545
- [BIG] name change · 0bab55d5
  thomwolf authored Jul 05, 2019
  
  0bab55d5
- standardizing tokenizers API and adding tests · e75c3f70
  thomwolf authored Jul 05, 2019
  
  e75c3f70
02 Jul, 2019 1 commit
- [LARGE] updating all tests and API · 1484d67d
  thomwolf authored Jul 02, 2019
  
  1484d67d
17 Apr, 2019 4 commits
- small clean up in tests · 34ae5bf8
  thomwolf authored Apr 17, 2019
  
  34ae5bf8
- relax network connection requirements · 265550ec
  thomwolf authored Apr 17, 2019
  
  265550ec
- adding s3 model tests with --runslow · 31d38760
  thomwolf authored Apr 17, 2019
  
  31d38760
- fixed GPT-2 tokenization on python 2 · bc70779b
  thomwolf authored Apr 17, 2019
  
  bc70779b
16 Apr, 2019 1 commit
- improving GPT2 tokenization and adding tests · 18a8a15f
  thomwolf authored Apr 16, 2019
  
  18a8a15f
15 Apr, 2019 2 commits
- fixing tests · e8568a3b
  thomwolf authored Apr 15, 2019
  
  e8568a3b
- added tokenizers serialization tests · 870b734b
  thomwolf authored Apr 15, 2019
  
  870b734b
11 Feb, 2019 2 commits
- tests pass on python 2 and 3 · 0a9860da
  thomwolf authored Feb 11, 2019
  
  0a9860da
- fix python 2.7 imports · 2071a9b8
  thomwolf authored Feb 11, 2019
  
  2071a9b8