Commits · 5a85f9bbd5602673d2b24ee64637ff87e03dfd5a · gaoqiong / lm-evaluation-harness

15 Mar, 2024 1 commit
- fixed encoding for seq2seq models · 5a85f9bb
  lintangsutawika authored Mar 15, 2024
  
  5a85f9bb
14 Mar, 2024 3 commits
- fixed stop token · a85e0150
  lintangsutawika authored Mar 14, 2024
  
  a85e0150
- fixed eos decoding · cf4cd770
  lintangsutawika authored Mar 14, 2024
  
  cf4cd770
- Merge branch 'main' of https://github.com/EleutherAI/lm-evaluation-harness into t5v2-alt-plus · 02e841ce
  lintangsutawika authored Mar 14, 2024
  
  02e841ce
13 Mar, 2024 1 commit

add manual tqdm disabling management (#1569) · e74ec966

achervyakov authored Mar 13, 2024



* add manual tqdm disabling management

* add typing to all new args

* apply precommit changes

---------
Co-authored-by: haileyschoelkopf <hailey@eleuther.ai>

e74ec966

12 Mar, 2024 1 commit
- cli_evaluate calls simple_evaluate with the same verbosity. (#1563) · 49695e8d
  Wongboo authored Mar 12, 2024
  
  49695e8d
11 Mar, 2024 4 commits

AGIEval (#1359) · a3e56afe

Hailey Schoelkopf authored Mar 11, 2024



* add agieval

* fix typo

* add cloze / math exactmatch agieval tasks, rename

* update exact-match agieval tasks, allow for multiple-correct answers

* add more detail to readme

* don't parse_math_answer twice

---------
Co-authored-by: Alex Bäuerle <alex@a13x.io>

a3e56afe

add Arabic EXAMS benchmark (#1498) · 4ab07597

khalil authored Mar 11, 2024



* add Arabic EXAMS benchmark

* fixed the linter issue, and add more information on the readme

* Update README.md

---------
Co-authored-by: Lintang Sutawika <lintang@sutawika.com>

4ab07597

Update ifeval.yaml (#1506) · 282b9e76
Hailey Schoelkopf authored Mar 11, 2024

282b9e76
Update generate_until_template_yaml (#1546) · a79a7c33
Hailey Schoelkopf authored Mar 11, 2024

a79a7c33

10 Mar, 2024 1 commit

Support jinja templating for task descriptions (#1553) · 3bdf25ec

Hisham Alyahya authored Mar 10, 2024



* Support jinja templating for "description"

* Update task_guide.md

* Update lm_eval/api/task.py

* fix format?

* whitespace errors

* fix whitespace

* fix bad variable reference

---------
Co-authored-by: Hailey Schoelkopf <65563625+haileyschoelkopf@users.noreply.github.com>
Co-authored-by: haileyschoelkopf <hailey@eleuther.ai>

3bdf25ec

09 Mar, 2024 2 commits

Fix incorrect `max_gen_toks` generation kwarg default in code2_text. (#1551) · f518228f
Piyush Thakur authored Mar 09, 2024
```
* update gen_kwargs in code2-text-go.yaml

* update gen_kwargs in rest code2-text
```
f518228f

Add compatibility for vLLM's new Logprob object (#1549) · 8051d954

Antoni Baum authored Mar 09, 2024



* Add compatibility for vLLM's new Logprob object

* Fix

* Update lm_eval/models/vllm_causallms.py

* fix format?

* trailing whitespace

---------
Co-authored-by: Hailey Schoelkopf <65563625+haileyschoelkopf@users.noreply.github.com>

8051d954

06 Mar, 2024 7 commits

Update installation commands in openai_completions.py and contributing... · 9e6e2402

Sungho Park authored Mar 07, 2024


Update installation commands in openai_completions.py and contributing document and, update wandb_args description (#1536)

* Update openai completions and docs/CONTRIBUTING.md

* Update wandb args description

* Update docs/interface.md

---------
Co-authored-by: Hailey Schoelkopf <65563625+haileyschoelkopf@users.noreply.github.com>

9e6e2402

Cleanup and fixes (Task, Instance, and a little bit of *evaluate) (#1533) · 4ee1b386

LSinev authored Mar 06, 2024



* Remove unused `decontamination_ngrams_path` and all mentions (still no alternative path provided)

* Fix improper import of LM and usage of evaluator in one of scripts

* update type hints in instance and task api

* raising errors in task.py instead of asserts

* Fix warnings from ruff

* raising errors in __main__.py instead of asserts

* raising errors in tasks/__init__.py instead of asserts

* raising errors in evaluator.py instead of asserts

* evaluator: update type hints and remove unused variables in code

* Update lm_eval/__main__.py
Co-authored-by: Hailey Schoelkopf <65563625+haileyschoelkopf@users.noreply.github.com>

* Update lm_eval/__main__.py
Co-authored-by: Hailey Schoelkopf <65563625+haileyschoelkopf@users.noreply.github.com>

* Update lm_eval/api/task.py
Co-authored-by: Hailey Schoelkopf <65563625+haileyschoelkopf@users.noreply.github.com>

* Update lm_eval/api/task.py
Co-authored-by: Hailey Schoelkopf <65563625+haileyschoelkopf@users.noreply.github.com>

* Update lm_eval/api/task.py
Co-authored-by: Hailey Schoelkopf <65563625+haileyschoelkopf@users.noreply.github.com>

* Update lm_eval/evaluator.py
Co-authored-by: Hailey Schoelkopf <65563625+haileyschoelkopf@users.noreply.github.com>

* pre-commit induced fixes

---------
Co-authored-by: Hailey Schoelkopf <65563625+haileyschoelkopf@users.noreply.github.com>

4ee1b386

update printed num-fewshot ; prevent fewshots from erroneously being used by... · 02705057
Hailey Schoelkopf authored Mar 06, 2024
```
update printed num-fewshot ; prevent fewshots from erroneously being used by cot which hardcodes fewshot prompt (#1502)
```
02705057
Update docs on LM.loglikelihood_rolling abstract method (#1532) · 525b8f5d
Hailey Schoelkopf authored Mar 06, 2024

525b8f5d
Adding new task : KorMedMCQA (#1530) · faee1adf
sean0042 authored Mar 06, 2024

faee1adf

Add WMDP Multiple-choice (#1534) · 29b2b013

Long Phan authored Mar 05, 2024



* init wmdp yaml file

* Add WMDP Multiple-choice

* fix linter issues

* Delete lm_eval/tasks/wmdp/_wmdp.yaml

---------
Co-authored-by: Lintang Sutawika <lintang@sutawika.com>

29b2b013

Add EQ-Bench as per #1459 (#1511) · c5acce0c

Peter Bevan authored Mar 06, 2024

* Start adding eq-bench

* Start adding to yaml and utils

* Get metric working

* Add README

* Handle cases where answer is not parseable

* Deal with unparseable answers and add percent_parseable metric

* Update README

c5acce0c

05 Mar, 2024 2 commits

Add a new task GPQA (the part CoT and generative) (#1482) · 01108aca

Uanu authored Mar 06, 2024



* Add new tasks of GPQA

* Add README

* Remove unused functions

* Remove unused functions

* Linters

* Add flexible match

* update

* Remove deplicate function

* Linter

* update

* Update lm_eval/filters/extraction.py
Co-authored-by: Hailey Schoelkopf <65563625+haileyschoelkopf@users.noreply.github.com>

* register multi_choice_regex

* Update

* run precommit

---------
Co-authored-by: Hailey Schoelkopf <65563625+haileyschoelkopf@users.noreply.github.com>
Co-authored-by: haileyschoelkopf <hailey@eleuther.ai>

01108aca

Openllm benchmark (#1526) · 8a875e9a
Baber Abbasi authored Mar 05, 2024

8a875e9a

04 Mar, 2024 4 commits

Fix minor edge cases (#951 #1503) (#1520) · 292e5814
Hailey Schoelkopf authored Mar 04, 2024
```
* Fix padding

* Fix elif in model loading

* format
```
292e5814
Hotfix: fix TypeError in `--trust_remote_code` (#1517) · 45823914
Hailey Schoelkopf authored Mar 04, 2024

45823914

French Bench (#1500) · 48476c4c

Manuel Faysse authored Mar 04, 2024



* add french-bench

* rename arc easy

* linting

* update datasets for no remote code exec

* fix string delimiter

* add info to readmr

* trim trailing whitespace

* add detailed groups

* add info to readme

* remove orangesum title from fbench main

* Force PPL tasks to be 0-shot

---------
Co-authored-by: Hailey Schoelkopf <65563625+haileyschoelkopf@users.noreply.github.com>

48476c4c

Cleaning up unused unit tests (#1516) · 4eba9cf3
Vicki Boykis authored Mar 03, 2024

4eba9cf3

03 Mar, 2024 2 commits

Setting trust_remote_code to True for HuggingFace datasets compatibility (#1487) · 95167926

Vicki Boykis authored Mar 03, 2024

* setting trust_remote_code

* dataset list no notebooks

* respect trust remote code

* Address changes, move cli options and change datasets

* fix task for tests

* headqa

* remove kobest

* pin datasets and address comments

* clean up space

95167926

Vllm update DP+TP (#1508) · e5e35fca

Baber Abbasi authored Mar 03, 2024

* use `@ray.remote` with distributed vLLM

* update versions

* bugfix

* unpin vllm

* fix pre-commit

* added version assertion error

* Revert "added version assertion error"

This reverts commit 8041e9b78e95eea9f4f4d0dc260115ba8698e9cc.

* added version assertion for DP

* expand DP note

* add warning

* nit

* pin vllm

* fix typos

e5e35fca

01 Mar, 2024 5 commits
- modify `WandbLogger` to accept arbitrary kwargs (#1491) · ae79b121
  Baber Abbasi authored Mar 01, 2024
```
* make `WandbLogger` init args optional

* nit

* nit

* nit

* move import warning to `WandbLogger`

* nit

* update docs

* nit
```
  ae79b121
- Improve data-parallel request partitioning for VLLM (#1477) · 27a3da96
  Hailey Schoelkopf authored Mar 01, 2024
```
* add undistribute + use more_itertools

* remove divide() util fn

* add more_itertools as dependency
```
  27a3da96
- always include EOS token in stopsequences if possible (#1480) · 284dd80d
  Hailey Schoelkopf authored Mar 01, 2024
  
  284dd80d
- Add multilingual truthfulqa targets (#1499) · d272c19f
  Zehan Li authored Mar 01, 2024
  
  d272c19f
- merged main · 90ad5db7
  lintangsutawika authored Mar 01, 2024
  
  90ad5db7
28 Feb, 2024 1 commit
- fix duplicated kwargs in some model init (#1495) · b177c82c
  Linsong Chu authored Feb 28, 2024
  
  b177c82c
27 Feb, 2024 4 commits

Fix AttributeError in huggingface.py When 'model_type' is Missing (#1489) · cc771eca

Rich authored Feb 27, 2024



* model_type attribute error

Getting attribute error when using a model without a 'model_type'

* fix w/ and w/out the 'model_type' specification

* use getattr(), also fix other config.model_type reference

* Update huggingface.py

---------
Co-authored-by: Hailey Schoelkopf <65563625+haileyschoelkopf@users.noreply.github.com>

cc771eca

update name of val split in truthfulqa multilingual (#1488) · a08eb870
Hailey Schoelkopf authored Feb 27, 2024

a08eb870
add multilingual mmlu eval (#1484) · 7cd004c4
Zehan Li authored Feb 27, 2024

7cd004c4

Refactor `evaluater.evaluate` (#1441) · 5ccd65d4

Baber Abbasi authored Feb 27, 2024



* change `all_gather` to `gather`

* add TaskOutput utility class

* Add FilterResults class and refactor task handling.

* Rename `key` to `filter_key` for clarity

* Add `print_writeout` function in utils.py

* Add function to calculate limit size.

* Add doc_iterator method to Task class

* Refactor `doc_iterator` and cleanup in Task class

* remove superfluous bits

* change `all_gather` to `gather`

* bugfix

* bugfix

* fix `gather`

* Refactor `gather` loop

* Refactor aggregate metrics calculation

* Refactor and simplify aggregate metrics calculation
Removed unused code

* Simplify metrics calculation and remove unused code.

* simplify the metrics calculation in `utils.py` and `evaluator.py`.

* Fix group metric

* change evaluate to hf_evaluate

* change evaluate to hf_evaluate

* add docs

* add docs

* nits

* make isslice keyword only

* nit

* add todo

* nit

* nit

* nit: swap order samples_metrics tuple

* move instance sorting outside loop

* nit

* nit

* Add __repr__ for ConfigurableTask

* nit

* nit

* Revert "nit"

This reverts commit dab8d9977a643752a17f840fd8cf7e4b107df28f.

* fix some logging

* nit

* fix `predict_only` bug. thanks to `@LSinev`!

* change `print_tasks` to `prepare_print_tasks`

* nits

* move eval utils

* move eval utils

* nit

* add comment

* added tqdm descriptions

* Update lm_eval/evaluator_utils.py
Co-authored-by: Hailey Schoelkopf <65563625+haileyschoelkopf@users.noreply.github.com>

* fix mgsm bug

* nit

* fix `build_all_requests`

* pre-commit

* add ceil to limit

---------
Co-authored-by: Hailey Schoelkopf <65563625+haileyschoelkopf@users.noreply.github.com>

5ccd65d4

26 Feb, 2024 2 commits

Cont metrics (#1475) · 96d185fa

Lintang Sutawika authored Feb 26, 2024



* add brier_score

* process brier_score

* brier score is working for N-sized class

* fxied brier score

* add TED to BigBench and Brier score to MMLU

* format

* Update metrics.py

* Update task.py

* Update generate_until_template_yaml

* Delete lm_eval/tasks/bigbench/aux_metric.py

* Update generate_until_template_yaml

* Update _default_template_yaml

* Update _generate_configs.py

* Update _generate_configs.py

* Update _generate_configs.py

* fix (format?)

* format?

* format, once more

---------
Co-authored-by: Hailey Schoelkopf <65563625+haileyschoelkopf@users.noreply.github.com>

96d185fa

Create a means for caching task registration and request building. Ad… (#1372) · 1e6c9272

Aaron V authored Feb 26, 2024



* Create a means for caching task registration and request building. Add the ability to specify an args dict for simple_evaluate().

* Remove extra S in cache path in caching module
Co-authored-by: Hailey Schoelkopf <65563625+haileyschoelkopf@users.noreply.github.com>

* Rename requests cache args, make model_args polymorphic so that a dict can also be accepted.

* Update docs to reflect new caching behavior, add CLI args for requests caching. Create a function for deleting items in the cache.

* Update documentation, fix minor bug with arg parsing for requests caching where an undefined variable was used.

* Remove line from gitignore, add to cli for caching datasets.

* Add hashing suffix to .pickles. Update test script typo.

* Favor isinstance() over type() in evaluator.py

* Add tests for caching, gets tests working, remove unneeded arg from build_all_requests().

* Update arg description to simple_evaluate.

* Update pyproject.toml

* Fix typehint

* Remove the use of random() for creating default cache pickle hash.

* Check that cache dir exists before clearing it in request cache tests.

* Fix linting problems.

* Fix additional formatting errors.

* Remove trailing whitespace.

* Add new line to the end of .gitignore.

---------
Co-authored-by: Hailey Schoelkopf <65563625+haileyschoelkopf@users.noreply.github.com>

1e6c9272