Commits · 96e499baf4fb9a382d7fa3f0bc533d3d20ea72fc · gaoqiong / lm-evaluation-harness

28 Jan, 2025 3 commits

fix multiple input chat tempalte (#2576) · 96e499ba

Baber Abbasi authored Jan 28, 2025

* feat: drop Python 3.8 support

* feat: drop Python 3.8 tests

* pre-commit

* handle chat_template for multiple iput

96e499ba

add TransformerLens example (#2651) · 42f79131

Nicky Pochinkov authored Jan 28, 2025

* add TransformerLens example

Many people use TransformerLens to do interpretability and interventions on models, and then need to test the model.

Here is a simple script that allows one to pass in the TransformerLens model and run evaluations on it.

* Ran pre-commit checks

42f79131

Add Moral Stories (#2653) · a0466f01

Irina Proskurina authored Jan 28, 2025

* Add moral stories task

* Add moral stories task

* Create README.md

* Update README.md

* Update line endings in moral_stories files

a0466f01

24 Jan, 2025 1 commit
- separate category for `global_mmlu` (#2652) · 5c006ed4
  Minho Ryu authored Jan 25, 2025
```
* separate category

* set version 0.0

* apply precommit
```
  5c006ed4
21 Jan, 2025 3 commits
- Fix max_tokens handling in vllm_vlms.py (#2637) · 370e2f9e
  Jan Kaniecki authored Jan 21, 2025
```
* Update vllm_vlms.py

* pre-commit

---------
Co-authored-by: Baber <baber@hey.com>
```
  370e2f9e
- aggregate by group (total and categories) (#2643) · b2c090cc
  Minho Ryu authored Jan 22, 2025
  
  b2c090cc
- revise mbpp prompt (#2645) · ed9c6fc8
  Minho Ryu authored Jan 22, 2025
  
  ed9c6fc8
20 Jan, 2025 6 commits

fixed mmlu generative response extraction (#2503) · 12b6eeb5

Ramiro R. C. authored Jan 20, 2025



* fixed mmlu generative response extraction

* updated file version | added args to exact_match

* fix

* fix

* pre-commit

* fix groups

---------
Co-authored-by: Baber <baber@hey.com>

12b6eeb5

fix tmlu tmlu_taiwan_specific_tasks tag (#2420) · 88144079
nike00811 authored Jan 21, 2025

88144079

Update KorMedMCQA: ver 2.0 (#2540) · ff2c49ff

Gyouk Chu authored Jan 21, 2025

* Update KorMedMCQA: ver 2.0

* Fix pre-commit formatting issues

* Update KorMedMCQA v2.0

* pre-commit

ff2c49ff

apply precommit (#2636) · 3a4e4674
Minho Ryu authored Jan 21, 2025

3a4e4674

New arabicmmlu (#2541) · 6dac8c69

Boda Sadallah authored Jan 21, 2025

* point to the original ArabicMMLU dataset

* create the new subtasks files

* fix bug when the context filed is empty

6dac8c69

add hrm8k benchmark for both Korean and English (#2627) · a5c344cf

Minho Ryu authored Jan 21, 2025



* add hrm8k benchmark for both Korean and English

* apply precommit

* revise tasks to make models not to directly answer; use zeroshot_cot if possible

* add README

* Add hrm8k on the task-list

---------
Co-authored-by: Baber <baber@hey.com>

a5c344cf

19 Jan, 2025 1 commit
- update pre-commit (#2632) · f724be69
  Baber Abbasi authored Jan 19, 2025
```
* update pre-commit
```
  f724be69
17 Jan, 2025 1 commit
- fix gen_prefix (#2630) · 9dda03d6
  Baber Abbasi authored Jan 17, 2025
```
* switch arg
```
  9dda03d6
15 Jan, 2025 4 commits

assistant prefill (#2615) · 703fbffd

Baber Abbasi authored Jan 15, 2025

* add assistant prefix

* add arc_challenge from llama

* nit

* nit

* nit

* add assistant prefix

* add mmlu_llama

* nit

* nit

* Revert "nit"

This reverts commit 6a97f8356237305e375212b966b30e8de59dd4bc.

* fix regex bug

* add assistant_prefix to vllm

* add `Question:`

* add mmlu_pro

* add fewshot assistant_prefix

* use `assistant_prefill`

* typehints

* nits

* nits

* add to docs

* add readme

703fbffd

Add MLQA (#2622) · e86cece6

Shivansh Pachnanda authored Jan 16, 2025

* Add MLQA
* add mlqa_common_yaml

* add 49 tests of mlqa family

* update tasks/README.md

---------

* fix: mlqa ast error

* nit: removed .yaml ext from template_yaml

* nit changes: minor modifications generate_tasks.py

* deleted    lm_eval/tasks/mlqa/mlqa_common_yaml.yaml

* tests updated

* nit

e86cece6

Add MBPP (#2247) · 5db23e2c

Hojin Lee authored Jan 16, 2025



* add mbpp

* fix some bugs

* add README for mbpp

* update README

* nits

---------
Co-authored-by: Hojin Lee <19949034+hjlee1371@users.noreply.github.com>
Co-authored-by: Baber <baber@hey.com>

5db23e2c

Add HumanEval (#1992) · 4c11206b

Hojin Lee authored Jan 16, 2025



* add custom filter

* fix type casting of references

* add humaneval

* fix a bug in humaneval

* add greedy version of humaneval

* update tasks README

* test humaneval

* return multiple metrics

* nit

* add confirmation to run code tasks

* nit

* nit

---------
Co-authored-by: Hojin Lee <19949034+hjlee1371@users.noreply.github.com>
Co-authored-by: Baber <baber@hey.com>

4c11206b

07 Jan, 2025 3 commits

Fix the format of mgsm zh and ja. (#2587) · bb098f13

Wenyang LUO authored Jan 07, 2025

* Fix the format of mgsm zh and ja.

* Add change log to mgsm.

* Add newline after changelog.

bb098f13

Fix Zeno visualizer on tasks like GSM8k (#2599) · 6d62a69c

Petr Baudis authored Jan 07, 2025



* fix(zeno): Generate unique ids in case of multiple filters

* fix(zeno): Report even non-aggregable metrics, just not as metrics

* pre-commit

---------
Co-authored-by: Baber <baber@hey.com>

6d62a69c

Fix gguf loading via Transformers (#2596) · 16cfe464

CL-ModelCloud authored Jan 07, 2025



* hf support load gguf file

* code review

* code review

* code clean up

* note about use_fast compat with gguf

---------
Co-authored-by: Qubitium-ModelCloud <qubitium@modelcloud.ai>

16cfe464

04 Jan, 2025 1 commit

some minor logging nits (#2609) · 888ac292

Baber Abbasi authored Jan 04, 2025

* remove yaml extension from phraes_va_common

* remove yaml extension from winogenerated

* remove yaml extension from phrases_es

* no cache debug logging when not used

888ac292

02 Jan, 2025 1 commit

update scrolls (#2602) · 1044db95

Baber Abbasi authored Jan 02, 2025

* update evaluate; update construct requests

* update construct requests to handle `apply_chat_template` kwarg

1044db95

30 Dec, 2024 1 commit
- fix model tests (#2604) · aa72104b
  Baber Abbasi authored Dec 30, 2024
```
upgrade transformers and peft in CI
```
  aa72104b
25 Dec, 2024 1 commit

fix extra_match low if batch_size > 1 (#2595) · 59f9ad4b

Wang, Yi authored Dec 25, 2024



* fix extra_match low if batch_size > 1
Signed-off-by: Wang, Yi A <yi.a.wang@intel.com>

* add sorting to logprobs

* nit

---------
Signed-off-by: Wang, Yi A <yi.a.wang@intel.com>
Co-authored-by: Baber <baber@hey.com>

59f9ad4b

24 Dec, 2024 1 commit

AraDICE task config file (#2507) · 932e8f9e

Firoj Alam, Scientist, QCRI authored Dec 24, 2024



* added aradice

* Added ArabicMMLU Lev Configs

* added ArabicMMLU egy configs

* Added boolq configs

* Added cultural bench configs

* added openbookqa configs

* Added PiQA configs

* added winogrande configs

* Added truthfulQA configs

* Added aradice group config

* Remove deleted files from repository

* modified arabimmlu configs

* modified metadata versions

* fixed formatting using ruff

* added aradice tasks information

* pre-commit

* Uptaded openbookqa utils

* fixed formatting on obqa

---------
Co-authored-by: Basel Mousi <bmousi@hbku.edu.qa>
Co-authored-by: Baber <baber@hey.com>

932e8f9e

20 Dec, 2024 1 commit
- Wandb step handling bugfix and feature (#2580) · b86aa213
  Sabrina J. Mielke authored Dec 20, 2024
  
  b86aa213
19 Dec, 2024 2 commits

add warning for truncation (#2585) · 6ccd520f
Baber Abbasi authored Dec 19, 2024
```
* add warning for truncation
```
6ccd520f

Add Global MMLU Lite (#2567) · 2b75b110

shivalika-singh authored Dec 19, 2024



* add global mmlu lite

* add global mmlu lite

* fix bugs

* add task README.md

* Update README.md

* Update tasks README.md

* Update README.md

* update readme

---------
Co-authored-by: shivi <shivalikasingh95@gmail.com>

2b75b110

17 Dec, 2024 2 commits
- drop python 3.8 support (#2575) · 8558b8d4
  Baber Abbasi authored Dec 17, 2024
```
* feat: drop Python 3.8 support

* feat: drop Python 3.8 tests

* pre-commit
```
  8558b8d4
- increment version (#2574) · 4c26a9c1
  Baber Abbasi authored Dec 17, 2024
```
forgot to increment 0.4.6!
```
  4c26a9c1
16 Dec, 2024 3 commits

fix `DeprecationWarning: invalid escape sequence '\s'` for whitespace filter (#2560) · 8d2f64c1

Baber Abbasi authored Dec 16, 2024

* fix `DeprecationWarning: invalid escape sequence '\s'`

* add type hints

* Revert "add type hints"

This reverts commit 15d8abc626a84e97f8c238ddfbf9e243d6f6eb5c.

8d2f64c1

batch `loglikelihood_rolling` across requests (#2559) · 0bfb0220

Baber Abbasi authored Dec 16, 2024

* batch all rolling token windows

* nit

* copy to vllm

* fix max_length for `get_rolling_token_windows`

* bugfix

* bugfix

* add type hints

0bfb0220

Adding new subtask to SCORE tasks: non greedy robustness (#2558) · 976d8a0b

Rima Shahbazyan authored Dec 16, 2024

* score readme added

* generate until task's "until" parameter's default value fixed.

* score mmlu-pro and agieval added

* changed macro accuracy to micro for agieval

* Always E removed from agi eval

* redundancies removed

* MATH added

* minor cosmetic changes for math

* Licenses added Readme updated

* changes for flake8 + license header on math

* Score added to readme and precommit was run.

* Score added to readme and precommit was run.

* Import error fixed

* math task bugfix
postprocess minor fix

* CR for math added

* math CR

* math task bugfix
postprocess minor fix

CR for math added

* Math cr fixed

* mmlu_pro non_greedy task added

* non greedy summarizer added

* Non greedy for all score tasks

* Bugfixes for non-greedy

* fixing the until argument

* undoing the change to "until" arguments default behaviour

* minor fix in summarizer

* log naming changes for better readability

* math subtasks naming fix

* agieval subtask naming fix

* logging added for debugging

* path issue fixed

* minor fix

* path fix

* path fix

* non_greedy_math minor fix

* final changes

* changed readme for non-greedy
added Nvidia header
added wxample script for non_greedy
changed prompts to match that fo trt runs

* non greedy summarizer bugfix

* non_greedy summarizer fixed

976d8a0b

14 Dec, 2024 1 commit
- add warning to readme (#2568) · 8de772f9
  Baber Abbasi authored Dec 14, 2024
```
* make warning prominent

* make warning prominent
```
  8de772f9
13 Dec, 2024 1 commit

add optimum-intel ipex model (#2566) · 919470a1

Yao Matrix authored Dec 14, 2024



* initial support for optimum-intel ipex model. LM model as first step

* format
Signed-off-by: Yao Matrix <matrix.yao@intel.com>

* pass dtype
Signed-off-by: Yao Matrix <matrix.yao@intel.com>

* update README
Signed-off-by: Yao, Matrix <matrix.yao@intel.com>

---------
Signed-off-by: Yao Matrix <matrix.yao@intel.com>

919470a1

09 Dec, 2024 2 commits

Update Lightning import (#2549) · 0b994433

Maanu Grover authored Dec 09, 2024



* update import
Signed-off-by: Maanu Grover <maanug@nvidia.com>

* run formatting

---------
Signed-off-by: Maanu Grover <maanug@nvidia.com>

0b994433

[API] left truncate for generate_until (#2554) · 2d11f2e5
Baber Abbasi authored Dec 09, 2024
```
* left truncate for generate_until

* pre-commit
```
2d11f2e5

05 Dec, 2024 1 commit
- Update README.md (#2546) · bcb4cbf4
  fzyzcjy authored Dec 05, 2024
  
  bcb4cbf4