Commits · e184c501cf8a8dea559ff6a1ce0928a1a54a7f1e · gaoqiong / lm-evaluation-harness

10 Jun, 2024 1 commit
- removed " " in make_table · e184c501
  lintangsutawika authored Jun 10, 2024
  
  e184c501
07 Jun, 2024 1 commit
- fix indentation · 5032ebaf
  lintangsutawika authored Jun 07, 2024
  
  5032ebaf
04 Jun, 2024 1 commit
- format fix · 4eeb8715
  lintangsutawika authored Jun 04, 2024
  
  4eeb8715
03 Jun, 2024 2 commits

KonradSzafer authored Jun 03, 2024



* initial chat template

* tokenizer attribute check

* variable rename

* interface update

* system instruction

* system inst default update

* fewshot as multiturn

* typing update

* indent update

* added comments

* Adding a fewshot in a more readable way

* linting

* Moved apply chat template to LM

* multiturn alternation fix

* cache key update

* apply chat template method fix

* add system prompt hash to cache_key

* tokenizer name property for cache_key

* property name fix

* linting backward compatibility fix

* docs and errors update

* add documentation on adding chat template compatibility to model_guide

* fewshot as multiturn check fix

* saving system inst and chat template in results

* eval tracker update

* docs update

* Apply suggestions from code review
Co-authored-by: Hailey Schoelkopf <65563625+haileyschoelkopf@users.noreply.github.com>

---------
Co-authored-by: haileyschoelkopf <hailey@eleuther.ai>
Co-authored-by: Clémentine Fourrier <22726840+clefourrier@users.noreply.github.com>
Co-authored-by: Hailey Schoelkopf <65563625+haileyschoelkopf@users.noreply.github.com>

070d31df

Fix fewshot seed only set when overriding num_fewshot (#1914) · 42d5c4bf
LSinev authored Jun 03, 2024
```
Fix #1906
```
42d5c4bf

30 May, 2024 1 commit

`higher_is_better` tickers in output table (#1893) · 14221c84

Zafir Stojanovski authored May 30, 2024



* Higher is better tickers in output table

* add extra check for `higher_is_better` not being None already

* Update lm_eval/evaluator.py

* fixup format I messed up

* add comment (and retrigger tests)

---------
Co-authored-by: Hailey Schoelkopf <65563625+haileyschoelkopf@users.noreply.github.com>
Co-authored-by: haileyschoelkopf <hailey@eleuther.ai>

14221c84

26 May, 2024 1 commit
- Rename `lm_eval.logging -> lm_eval.loggers` (#1858) · 0ff6ab99
  Hailey Schoelkopf authored May 26, 2024
```
* rename lm_eval.logging module

* fix evaluation tracker args
```
  0ff6ab99
24 May, 2024 1 commit
- Fix for bootstrap_iters = 0 case (#1715) (#1789) · b043b050
  Hailey Schoelkopf authored May 24, 2024
```
* add handling for bootstrap_iters=0 case

* add more detail to docstring

* run precommit
```
  b043b050
16 May, 2024 1 commit
- add get_subtask_list function to get proper subtask list · 88fea8ad
  lintangsutawika authored May 16, 2024
  
  88fea8ad
10 May, 2024 2 commits
- use task_id · 39c40277
  lintangsutawika authored May 10, 2024
  
  39c40277
- reformat with pre-commit · 13203943
  lintangsutawika authored May 10, 2024
  
  13203943
08 May, 2024 2 commits
- fixed conditional statement · a88886a3
  lintangsutawika authored May 08, 2024
  
  a88886a3
- update additional condition when loading a group in a group yaml · 44d70398
  lintangsutawika authored May 08, 2024
  
  44d70398
07 May, 2024 8 commits
- add version for groups · 2a817535
  lintangsutawika authored May 07, 2024
  
  2a817535
- sort set to False by default, fix predict_only arg · 9f698c20
  lintangsutawika authored May 07, 2024
  
  9f698c20
- move prepare_print_tasks to evaluator_utils · 110e5a28
  lintangsutawika authored May 07, 2024
  
  110e5a28
- readd group_agg · 03982e03
  lintangsutawika authored May 07, 2024
  
  03982e03
- update to work with new group and task configuration · ad70d206
  lintangsutawika authored May 07, 2024
  
  ad70d206
- Fix Caching Tests ; Remove `pretrained=gpt2` default (#1775) · 7fe2b93c
  Hailey Schoelkopf authored May 07, 2024
  
  7fe2b93c
- update mmlu · c23c9305
  lintangsutawika authored May 07, 2024
  
  c23c9305
- adjust group scoring with using ConfigurableGroup · 62572f05
  lintangsutawika authored May 07, 2024
  
  62572f05
06 May, 2024 1 commit

Provide ability for custom sampler for ConfigurableTask (#1616) · ae72cebc

LSinev authored May 06, 2024

* Added fewshot sampling seeds to evaluator.simple_evaluate signature

Way to control seed of fewshot sampling
may help with #1591

* Added ability for custom sampler for ConfigurableTask

May be set in config like
```
fewshot_config:
  sampler: !function utils.MyFewshotSampler
```

* explicitly set fewshot random generator seed for HFLM generate_until_task test

* add backward compatibility for three args seed setup

* save seeds info to logs/reports

ae72cebc

05 May, 2024 1 commit
- limit fix (#1785) · cee785e0
  KonradSzafer authored May 05, 2024
  
  cee785e0
03 May, 2024 1 commit

evaluation tracker implementation (#1766) · 59cf408a

KonradSzafer authored May 03, 2024

* evaluation tracker implementation

* OVModelForCausalLM test fix

* typo fix

* moved methods args

* multiple args in one flag

* loggers moved to dedicated dir

* improved filename sanitization

59cf408a

25 Apr, 2024 1 commit
- fixed args input in aggregate_subtask_metrics · 9551bbf2
  lintangsutawika authored Apr 25, 2024
  
  9551bbf2
24 Apr, 2024 1 commit
- fixed size configuration · 5a98162d
  lintangsutawika authored Apr 24, 2024
  
  5a98162d
23 Apr, 2024 1 commit
- add a group config that allows disabling table for group score and group aggregate in general · 6da6d187
  lintangsutawika authored Apr 23, 2024
  
  6da6d187
22 Mar, 2024 1 commit
- add logging of model args (#1619) · cffc1bd3
  Baber Abbasi authored Mar 22, 2024
```
* add logging of model args

* nit

* Add warnings.

* nit

* add warning

* nit
```
  cffc1bd3
18 Mar, 2024 1 commit

Cleanup for v0.4.2 release (#1573) · 5627e819

Hailey Schoelkopf authored Mar 18, 2024

* Update interface.md

* fix: make caching reqs always work with accelerate launch

* remove stale task migration checklist

* remove deprecation warnings

* make informative TypeErrors for get_task_dict

* bump version metadata

* fix num_fewshot printing bug

* add fewshot value to cache key

5627e819

17 Mar, 2024 1 commit
- Add start date in results.json (#1592) · 6fae67a6
  kwrobel.eth authored Mar 17, 2024
  
  6fae67a6
06 Mar, 2024 1 commit

Cleanup and fixes (Task, Instance, and a little bit of *evaluate) (#1533) · 4ee1b386

LSinev authored Mar 06, 2024



* Remove unused `decontamination_ngrams_path` and all mentions (still no alternative path provided)

* Fix improper import of LM and usage of evaluator in one of scripts

* update type hints in instance and task api

* raising errors in task.py instead of asserts

* Fix warnings from ruff

* raising errors in __main__.py instead of asserts

* raising errors in tasks/__init__.py instead of asserts

* raising errors in evaluator.py instead of asserts

* evaluator: update type hints and remove unused variables in code

* Update lm_eval/__main__.py
Co-authored-by: Hailey Schoelkopf <65563625+haileyschoelkopf@users.noreply.github.com>

* Update lm_eval/__main__.py
Co-authored-by: Hailey Schoelkopf <65563625+haileyschoelkopf@users.noreply.github.com>

* Update lm_eval/api/task.py
Co-authored-by: Hailey Schoelkopf <65563625+haileyschoelkopf@users.noreply.github.com>

* Update lm_eval/api/task.py
Co-authored-by: Hailey Schoelkopf <65563625+haileyschoelkopf@users.noreply.github.com>

* Update lm_eval/api/task.py
Co-authored-by: Hailey Schoelkopf <65563625+haileyschoelkopf@users.noreply.github.com>

* Update lm_eval/evaluator.py
Co-authored-by: Hailey Schoelkopf <65563625+haileyschoelkopf@users.noreply.github.com>

* pre-commit induced fixes

---------
Co-authored-by: Hailey Schoelkopf <65563625+haileyschoelkopf@users.noreply.github.com>

4ee1b386

04 Mar, 2024 1 commit
- Fix minor edge cases (#951 #1503) (#1520) · 292e5814
  Hailey Schoelkopf authored Mar 04, 2024
```
* Fix padding

* Fix elif in model loading

* format
```
  292e5814
27 Feb, 2024 1 commit

Refactor `evaluater.evaluate` (#1441) · 5ccd65d4

Baber Abbasi authored Feb 27, 2024



* change `all_gather` to `gather`

* add TaskOutput utility class

* Add FilterResults class and refactor task handling.

* Rename `key` to `filter_key` for clarity

* Add `print_writeout` function in utils.py

* Add function to calculate limit size.

* Add doc_iterator method to Task class

* Refactor `doc_iterator` and cleanup in Task class

* remove superfluous bits

* change `all_gather` to `gather`

* bugfix

* bugfix

* fix `gather`

* Refactor `gather` loop

* Refactor aggregate metrics calculation

* Refactor and simplify aggregate metrics calculation
Removed unused code

* Simplify metrics calculation and remove unused code.

* simplify the metrics calculation in `utils.py` and `evaluator.py`.

* Fix group metric

* change evaluate to hf_evaluate

* change evaluate to hf_evaluate

* add docs

* add docs

* nits

* make isslice keyword only

* nit

* add todo

* nit

* nit

* nit: swap order samples_metrics tuple

* move instance sorting outside loop

* nit

* nit

* Add __repr__ for ConfigurableTask

* nit

* nit

* Revert "nit"

This reverts commit dab8d9977a643752a17f840fd8cf7e4b107df28f.

* fix some logging

* nit

* fix `predict_only` bug. thanks to `@LSinev`!

* change `print_tasks` to `prepare_print_tasks`

* nits

* move eval utils

* move eval utils

* nit

* add comment

* added tqdm descriptions

* Update lm_eval/evaluator_utils.py
Co-authored-by: Hailey Schoelkopf <65563625+haileyschoelkopf@users.noreply.github.com>

* fix mgsm bug

* nit

* fix `build_all_requests`

* pre-commit

* add ceil to limit

---------
Co-authored-by: Hailey Schoelkopf <65563625+haileyschoelkopf@users.noreply.github.com>

5ccd65d4

26 Feb, 2024 1 commit

Create a means for caching task registration and request building. Ad… (#1372) · 1e6c9272

Aaron V authored Feb 26, 2024



* Create a means for caching task registration and request building. Add the ability to specify an args dict for simple_evaluate().

* Remove extra S in cache path in caching module
Co-authored-by: Hailey Schoelkopf <65563625+haileyschoelkopf@users.noreply.github.com>

* Rename requests cache args, make model_args polymorphic so that a dict can also be accepted.

* Update docs to reflect new caching behavior, add CLI args for requests caching. Create a function for deleting items in the cache.

* Update documentation, fix minor bug with arg parsing for requests caching where an undefined variable was used.

* Remove line from gitignore, add to cli for caching datasets.

* Add hashing suffix to .pickles. Update test script typo.

* Favor isinstance() over type() in evaluator.py

* Add tests for caching, gets tests working, remove unneeded arg from build_all_requests().

* Update arg description to simple_evaluate.

* Update pyproject.toml

* Fix typehint

* Remove the use of random() for creating default cache pickle hash.

* Check that cache dir exists before clearing it in request cache tests.

* Fix linting problems.

* Fix additional formatting errors.

* Remove trailing whitespace.

* Add new line to the end of .gitignore.

---------
Co-authored-by: Hailey Schoelkopf <65563625+haileyschoelkopf@users.noreply.github.com>

1e6c9272

24 Feb, 2024 1 commit

Add environment and transformers version logging in results dump (#1464) · f78e2da4

LSinev authored Feb 24, 2024

* Save git_hash to results even if git is not available to call as subprocess

* Store more info about environment and transformers version in results to help researchers track inconsistencies

* moved added logging to logging_utils

* moved get_git_commit_hash to logging_utils.py

* moved add_env_info inside evaluator

f78e2da4

22 Feb, 2024 1 commit
- Log which subtasks were called with which groups (#1456) · 00dc9960
  Hailey Schoelkopf authored Feb 22, 2024
```
* log group membership

* no stray prints

* Update evaluator.py
```
  00dc9960
12 Feb, 2024 1 commit
- Added seeds to `evaluator.simple_evaluate` signature (#1412) · bfbd0325
  Amine Elhattami authored Feb 12, 2024
```
* Added seeds to `evaluator.simple_evaluate` signature

* Added  CLI argument

* Updated  to add  arg.
```
  bfbd0325
11 Feb, 2024 1 commit

Evaluate (#1385) · 1ff84897

Baber Abbasi authored Feb 11, 2024

* un-exclude `evaluate.py` from linting

* readability

* readability

* add task name to build info message

* fix link

* nit

* add functions for var and mean pooling

* add functions for var and mean pooling

* metadata compatibility with task

* rename `override_config` to `set_config` and move to `Task`

* add unit test

* nit

* nit

* bugfix

* nit

* nit

* nit

* add docstrings

* fix metadata-fewshot

* revert metric refactor

* nit

* type checking

* type hints

* type hints

* move `override_metric` to `Task`

* change metadata

* change name

* pre-commit

* rename

* remove

* remove

* `override_metric` backwards compatible with `Task`

* type hints

* use generic

* type hint

1ff84897

09 Feb, 2024 1 commit
- use reversed task hierarchy for print (#1414) · ab4dba8f
  Hailey Schoelkopf authored Feb 09, 2024
  
  ab4dba8f
06 Feb, 2024 1 commit

Use Pooled rather than Combined Variance for calculating stderr of task groupings (#1390) · 94cc1850

Hailey Schoelkopf authored Feb 06, 2024

* update formula for stderr aggregation

* hack: see what happens when using stderr_for_metric bootstrapping on a group

* undo bootstrap_for_stderr test

* factor out variance-aggregation formulas into api.metrics

* fix failing tests

* remove stray print

* update comment

* further detail in comment

* add back initialize_tasks() call

* fix format

94cc1850