Commits · 9bd948df0d4d2f3503bbcbf2e78091cf2881e1fa · gaoqiong / lm-evaluation-harness

22 May, 2024 1 commit
- Update polemo2_out.yaml (#1871) · 70e1de09
  zhabuye authored May 22, 2024
  
  70e1de09
21 May, 2024 1 commit
- fixed incorrect check for task type (replace `~` with `not`) (#1865) · 00b7a61c
  Zafir Stojanovski authored May 21, 2024
  
  00b7a61c
13 May, 2024 1 commit

Adding tinyBenchmarks datasets (#1545) · fe9fef4e

Lucas Weber authored May 13, 2024



* Add tinyBenchmarks

* Add acknowledgements

* Add ordering of outputs for data-parallel

* Run pre-commit

* Add few_shot specifications

* Add tinyBenchmarks post-processing

* add conditional import ; fix task names

---------
Co-authored-by: haileyschoelkopf <hailey@eleuther.ai>

fe9fef4e

09 May, 2024 1 commit

Copal task (#1803) · 1980a13c

Edd authored May 10, 2024

* add copal

* change name to copal id for clarity and the task name

* remove `copal_id...` to yaml to make it work

* checkmark on README

* change group name to `copal_id`

1980a13c

08 May, 2024 1 commit

add task for mmlu evaluation in arc multiple choice format (#1745) · 9097ad3e

jonabur authored May 08, 2024



* add mmlu arc style evaluation

* rename arc_style to continuation

---------
Co-authored-by: Jonathan Burdge <jburdge@mahti-login11.mahti.csc.fi>
Co-authored-by: Jonathan Burdge <jburdge@mahti-login12.mahti.csc.fi>

9097ad3e

07 May, 2024 3 commits

Initial integration of the Unitxt to LM eval harness (#1615) · 885f48d6

Yoav Katz authored May 08, 2024

* Initial support for Unitxt datasets in LM Eval Harness

See  https://github.com/IBM/unitxt



The script 'generate_yamls.py' creates LM Eval Harness yaml files corresponding to Unitxt datasets specified in the 'unitxt_datasets' file.

The glue code required to register Unitxt metrics is in 'unitxt_wrapper.py'.

* Added dataset loading check to generate_yaml

Improved error messages.

* Speed up generate_yaml

Added printouts and improved error message

* Added output printout

* Simplified integration of unitxt datasets

Store all the common yaml configuration in a yaml include shared by all datasets of the same task.

* Post code review comments - part 1

1. Made sure include files don't end wth 'yaml' so they won't be marked as tasks
2. Added more datasets and tasks (NER, GEC)
3. Added README

* Post code review comments - part 2

1. Added install unitxt install option in pyproject.toml:
pip install 'lm_eval[unitxt]'
2. Added a check that unitxt is installed and print a clear error message if not

* Commited missing pyproject change

* Added documentation on adding datasets

* More doc changes

* add unitxt extra to readme

* run precommit

---------
Co-authored-by: haileyschoelkopf <hailey@eleuther.ai>

885f48d6

Re-add Hendrycks MATH (no sympy checking, no Minerva hardcoded prompt) variant (#1793) · d42a3e44
Hailey Schoelkopf authored May 07, 2024
```
* add Hendrycks MATH (no sympy checking) variant

* add readmes for MATH tasks
```
d42a3e44
Fix Caching Tests ; Remove `pretrained=gpt2` default (#1775) · 7fe2b93c
Hailey Schoelkopf authored May 07, 2024

7fe2b93c

01 May, 2024 4 commits

upload new tasks (#1728) · caaf9ab6

Simran Arora authored May 01, 2024



* upload new tasks

* add readmes

* run linters

---------
Co-authored-by: haileyschoelkopf <hailey@eleuther.ai>

caaf9ab6

Fix m_arc choices (#1760) · f27c4050

Zehan Li authored May 02, 2024



* Update utils.py

This is a 4-choice task, option_e is null for all but 3 samples

* Fix options

Adaptive choices

* add option e

* bump multilingual arc version

---------
Co-authored-by: Hailey Schoelkopf <65563625+haileyschoelkopf@users.noreply.github.com>

f27c4050

Pile 10k new task (#1758) · b898bdaa
Gabriel Mukobi authored May 01, 2024
```
* Add Pile-10k readme

* Add Pile-10k task configuration file
```
b898bdaa
remove duplicated `num_fewshot: 0` (#1769) · 552eeae7
Chujie Zheng authored May 01, 2024

552eeae7

26 Apr, 2024 1 commit
- Support individual scrolls datasets (#1740) · 9b49556a
  giorgossideris authored Apr 26, 2024
```
* Support individual scrolls datasets

* Add qmsum context

* Fix formatting
```
  9b49556a
25 Apr, 2024 2 commits
- Fix Parameter Propagation for Tasks that have `include` (#1749) · 0bafcef0
  Lintang Sutawika authored Apr 26, 2024
```
* Update task.py

* Update __init__.py
```
  0bafcef0
- Add XNLIeu: a dataset for cross-lingual NLI in Basque (#1694) · 594015b6
  Julen Etxaniz authored Apr 25, 2024
```
* add xnli_eu tasks

* update tasks readme

* update readme
```
  594015b6
18 Apr, 2024 1 commit
- Adding retries and rate limit to toxicity tasks (#1620) · 3196e907
  sator-labs authored Apr 18, 2024
  
  3196e907
05 Apr, 2024 1 commit

TMMLU+ implementation (#1394) · 9ae96cdf

ZoneTwelve authored Apr 06, 2024



* implementation of TMMLU+

* implemented: TMMLU+

****TMMLU+ : large-scale Traditional chinese Massive Multitask language Understanding****

- 4 categories
    - STEM
    - Social Science
    - Humanities
    - Other

The TMMLU+ dataset, encompassing over 67 subjects and 20160 tasks, is six times larger and more balanced than its predecessor, TMMLU, and includes benchmark results from both closed-source and 20 open-weight Chinese large language models with 1.8B to 72B parameters. However, Traditional Chinese variants continue to underperform compared to major Simplified Chinese models.

```markdown
Total number of tasks in the 'test' sets: 20160
Total number of tasks in the 'validation' sets: 2247
Total number of tasks in the 'train' sets: 335
```

* Remove print from __init__.py

There was my mistake in forgetting to remove the debug print from the code.

* update: move TMMLU+ config generation program into default

* fix: we should use training set as few shots example

* update: README for TMMLU+

* update: a small changes of TMMLU+ README file

* pre-commit run thought

* Add README for TMMLU+ dataset

* run precommit

* trigger precommit again

* trigger precommit again

* isort is fussy

* isort is fussy

* format, again

* oops

* oops

---------
Co-authored-by: lintang <lintang@eleuther.ai>
Co-authored-by: haileyschoelkopf <hailey@eleuther.ai>

9ae96cdf

04 Apr, 2024 1 commit
- Patch QQP prompt (#1661) · ff24e992
  Hailey Schoelkopf authored Apr 04, 2024
  
  ff24e992
01 Apr, 2024 1 commit

Add Latxa paper evaluation tasks for Basque (#1654) · c2c8e238

Julen Etxaniz authored Apr 01, 2024

* add basqueglue

* add eus_exams

* add eus_proficiency

* add eus_reading

* add eus_trivia

* run pre-commit

c2c8e238

28 Mar, 2024 1 commit
- Fix SuperGlue's ReCoRD task following regression in v0.4 refactoring (#1647) · ab7cc6b1
  Or Sharir authored Mar 28, 2024
  
  ab7cc6b1
21 Mar, 2024 1 commit
- Add ACLUE task (#1614) · 65546905
  Haonan Li authored Mar 21, 2024
```
* Add task ACLUE

* fix minor bug

* fix code style

* fix code style
```
  65546905
18 Mar, 2024 2 commits

Fix eval_logger import for mmlu/_generate_configs.py (#1593) · 4600d6bf

Nouf M. Alotaibi authored Mar 18, 2024



* Fix eval_logger import for mmlu/_generate_configs.py

* linter

---------
Co-authored-by: Hailey Schoelkopf <65563625+haileyschoelkopf@users.noreply.github.com>

4600d6bf

Cleanup for v0.4.2 release (#1573) · 5627e819

Hailey Schoelkopf authored Mar 18, 2024

* Update interface.md

* fix: make caching reqs always work with accelerate launch

* remove stale task migration checklist

* remove deprecation warnings

* make informative TypeErrors for get_task_dict

* bump version metadata

* fix num_fewshot printing bug

* add fewshot value to cache key

5627e819

15 Mar, 2024 1 commit
- Fix Jinja template for Advanced AI Risk (#1587) · dc90fecc
  Rylan Schaeffer authored Mar 15, 2024
  
  dc90fecc
13 Mar, 2024 1 commit

add manual tqdm disabling management (#1569) · e74ec966

achervyakov authored Mar 13, 2024



* add manual tqdm disabling management

* add typing to all new args

* apply precommit changes

---------
Co-authored-by: haileyschoelkopf <hailey@eleuther.ai>

e74ec966

11 Mar, 2024 4 commits

AGIEval (#1359) · a3e56afe

Hailey Schoelkopf authored Mar 11, 2024



* add agieval

* fix typo

* add cloze / math exactmatch agieval tasks, rename

* update exact-match agieval tasks, allow for multiple-correct answers

* add more detail to readme

* don't parse_math_answer twice

---------
Co-authored-by: Alex Bäuerle <alex@a13x.io>

a3e56afe

add Arabic EXAMS benchmark (#1498) · 4ab07597

khalil authored Mar 11, 2024



* add Arabic EXAMS benchmark

* fixed the linter issue, and add more information on the readme

* Update README.md

---------
Co-authored-by: Lintang Sutawika <lintang@sutawika.com>

4ab07597

Update ifeval.yaml (#1506) · 282b9e76
Hailey Schoelkopf authored Mar 11, 2024

282b9e76
Update generate_until_template_yaml (#1546) · a79a7c33
Hailey Schoelkopf authored Mar 11, 2024

a79a7c33

09 Mar, 2024 1 commit
- Fix incorrect `max_gen_toks` generation kwarg default in code2_text. (#1551) · f518228f
  Piyush Thakur authored Mar 09, 2024
```
* update gen_kwargs in code2-text-go.yaml

* update gen_kwargs in rest code2-text
```
  f518228f
06 Mar, 2024 5 commits

Cleanup and fixes (Task, Instance, and a little bit of *evaluate) (#1533) · 4ee1b386

LSinev authored Mar 06, 2024



* Remove unused `decontamination_ngrams_path` and all mentions (still no alternative path provided)

* Fix improper import of LM and usage of evaluator in one of scripts

* update type hints in instance and task api

* raising errors in task.py instead of asserts

* Fix warnings from ruff

* raising errors in __main__.py instead of asserts

* raising errors in tasks/__init__.py instead of asserts

* raising errors in evaluator.py instead of asserts

* evaluator: update type hints and remove unused variables in code

* Update lm_eval/__main__.py
Co-authored-by: Hailey Schoelkopf <65563625+haileyschoelkopf@users.noreply.github.com>

* Update lm_eval/__main__.py
Co-authored-by: Hailey Schoelkopf <65563625+haileyschoelkopf@users.noreply.github.com>

* Update lm_eval/api/task.py
Co-authored-by: Hailey Schoelkopf <65563625+haileyschoelkopf@users.noreply.github.com>

* Update lm_eval/api/task.py
Co-authored-by: Hailey Schoelkopf <65563625+haileyschoelkopf@users.noreply.github.com>

* Update lm_eval/api/task.py
Co-authored-by: Hailey Schoelkopf <65563625+haileyschoelkopf@users.noreply.github.com>

* Update lm_eval/evaluator.py
Co-authored-by: Hailey Schoelkopf <65563625+haileyschoelkopf@users.noreply.github.com>

* pre-commit induced fixes

---------
Co-authored-by: Hailey Schoelkopf <65563625+haileyschoelkopf@users.noreply.github.com>

4ee1b386

update printed num-fewshot ; prevent fewshots from erroneously being used by... · 02705057
Hailey Schoelkopf authored Mar 06, 2024
```
update printed num-fewshot ; prevent fewshots from erroneously being used by cot which hardcodes fewshot prompt (#1502)
```
02705057
Adding new task : KorMedMCQA (#1530) · faee1adf
sean0042 authored Mar 06, 2024

faee1adf

Add WMDP Multiple-choice (#1534) · 29b2b013

Long Phan authored Mar 05, 2024



* init wmdp yaml file

* Add WMDP Multiple-choice

* fix linter issues

* Delete lm_eval/tasks/wmdp/_wmdp.yaml

---------
Co-authored-by: Lintang Sutawika <lintang@sutawika.com>

29b2b013

Add EQ-Bench as per #1459 (#1511) · c5acce0c

Peter Bevan authored Mar 06, 2024

* Start adding eq-bench

* Start adding to yaml and utils

* Get metric working

* Add README

* Handle cases where answer is not parseable

* Deal with unparseable answers and add percent_parseable metric

* Update README

c5acce0c

05 Mar, 2024 2 commits

Add a new task GPQA (the part CoT and generative) (#1482) · 01108aca

Uanu authored Mar 06, 2024



* Add new tasks of GPQA

* Add README

* Remove unused functions

* Remove unused functions

* Linters

* Add flexible match

* update

* Remove deplicate function

* Linter

* update

* Update lm_eval/filters/extraction.py
Co-authored-by: Hailey Schoelkopf <65563625+haileyschoelkopf@users.noreply.github.com>

* register multi_choice_regex

* Update

* run precommit

---------
Co-authored-by: Hailey Schoelkopf <65563625+haileyschoelkopf@users.noreply.github.com>
Co-authored-by: haileyschoelkopf <hailey@eleuther.ai>

01108aca

Openllm benchmark (#1526) · 8a875e9a
Baber Abbasi authored Mar 05, 2024

8a875e9a

04 Mar, 2024 1 commit

French Bench (#1500) · 48476c4c

Manuel Faysse authored Mar 04, 2024



* add french-bench

* rename arc easy

* linting

* update datasets for no remote code exec

* fix string delimiter

* add info to readmr

* trim trailing whitespace

* add detailed groups

* add info to readme

* remove orangesum title from fbench main

* Force PPL tasks to be 0-shot

---------
Co-authored-by: Hailey Schoelkopf <65563625+haileyschoelkopf@users.noreply.github.com>

48476c4c

03 Mar, 2024 1 commit

Setting trust_remote_code to True for HuggingFace datasets compatibility (#1487) · 95167926

Vicki Boykis authored Mar 03, 2024

* setting trust_remote_code

* dataset list no notebooks

* respect trust remote code

* Address changes, move cli options and change datasets

* fix task for tests

* headqa

* remove kobest

* pin datasets and address comments

* clean up space

95167926

01 Mar, 2024 1 commit
- Add multilingual truthfulqa targets (#1499) · d272c19f
  Zehan Li authored Mar 01, 2024
  
  d272c19f