README.md 1.85 KB
Newer Older
haileyschoelkopf's avatar
haileyschoelkopf committed
1
2
3
# v1.0 Tasks
This list keeps track of which tasks' implementations have been ported to YAML / v2.0 of the Eval Harness.

4
Boxes should be checked iff tasks are implemented in the refactor and tested for regression. Tasks should be struck through if checked *against original introducing paper* implementation or popularizing implementation. (WIP) Denotes that there exists a PR or person working on this task already.
haileyschoelkopf's avatar
haileyschoelkopf committed
5

lintangsutawika's avatar
lintangsutawika committed
6
- [x] Glue
haileyschoelkopf's avatar
haileyschoelkopf committed
7
- [x] SuperGlue
lintangsutawika's avatar
lintangsutawika committed
8
- [x] CoQA
lintangsutawika's avatar
lintangsutawika committed
9
- [x] DROP
haileyschoelkopf's avatar
haileyschoelkopf committed
10
11
- [x] ~~Lambada~~
- [x] Lambada (Cloze variants)
12
- [x] ~~Lambada (Multilingual)~~
haileyschoelkopf's avatar
haileyschoelkopf committed
13
14
- [x] Wikitext
- [x] PiQA
Hailey Schoelkopf's avatar
Hailey Schoelkopf committed
15
- [x] PROST
lintangsutawika's avatar
lintangsutawika committed
16
- [x] MCTACO
17
- [x] Pubmed QA
haileyschoelkopf's avatar
haileyschoelkopf committed
18
- [x] SciQ
lintangsutawika's avatar
lintangsutawika committed
19
- [x] QASPER
Hailey Schoelkopf's avatar
Hailey Schoelkopf committed
20
- [x] QA4MRE
Lintang Sutawika's avatar
Lintang Sutawika committed
21
- [x] TriviaQA
haileyschoelkopf's avatar
haileyschoelkopf committed
22
- [x] AI2 ARC
Hailey Schoelkopf's avatar
Hailey Schoelkopf committed
23
- [x] LogiQA
24
- [x] HellaSwag
Benjamin Fattori's avatar
Benjamin Fattori committed
25
- [x] SWAG
Hailey Schoelkopf's avatar
Hailey Schoelkopf committed
26
- [x] OpenBookQA
Lintang Sutawika's avatar
Lintang Sutawika committed
27
- [ ] SQuADv2 (Lintang)
Hailey Schoelkopf's avatar
Hailey Schoelkopf committed
28
- [x] RACE
Hailey Schoelkopf's avatar
Hailey Schoelkopf committed
29
- [x] HeadQA
Lintang Sutawika's avatar
Lintang Sutawika committed
30
- [x] MathQA
Hailey Schoelkopf's avatar
Hailey Schoelkopf committed
31
- [x] WebQs
lintangsutawika's avatar
lintangsutawika committed
32
- [x] WSC273
33
- [x] Winogrande
haileyschoelkopf's avatar
haileyschoelkopf committed
34
- [x] ANLI
Hailey Schoelkopf's avatar
Hailey Schoelkopf committed
35
- [x] Hendrycks Ethics (missing some tasks/metrics, see PR 660: <https://github.com/EleutherAI/lm-evaluation-harness/pull/660> for more info)
lintangsutawika's avatar
lintangsutawika committed
36
37
38
- [x] TruthfulQA (mc1)
- [x] TruthfulQA (mc2)
- [x] TruthfulQA (gen)
lintangsutawika's avatar
lintangsutawika committed
39
- [x] MuTual
Lintang Sutawika's avatar
Lintang Sutawika committed
40
- [ ] Hendrycks Math (Hailey)
lintangsutawika's avatar
lintangsutawika committed
41
- [x] Asdiv
haileyschoelkopf's avatar
haileyschoelkopf committed
42
- [ ] GSM8k
43
- [x] Arithmetic
Hailey Schoelkopf's avatar
Hailey Schoelkopf committed
44
- [ ] MMMLU (Hailey)
Lintang Sutawika's avatar
Lintang Sutawika committed
45
- [x] Translation (WMT) suite
haileyschoelkopf's avatar
haileyschoelkopf committed
46
- [x] Unscramble
haileyschoelkopf's avatar
haileyschoelkopf committed
47
- [x] ~~Pile (perplexity)~~
lintangsutawika's avatar
lintangsutawika committed
48
- [x] BLiMP
haileyschoelkopf's avatar
haileyschoelkopf committed
49
- [x] ToxiGen
lintangsutawika's avatar
lintangsutawika committed
50
- [x] StoryCloze
Lintang Sutawika's avatar
Lintang Sutawika committed
51
- [ ] NaturalQs (Hailey)
haileyschoelkopf's avatar
haileyschoelkopf committed
52
- [x] CrowS-Pairs
lintangsutawika's avatar
lintangsutawika committed
53
- [x] XCopa
Hailey Schoelkopf's avatar
Hailey Schoelkopf committed
54
- [ ] BIG-Bench (Hailey)
lintangsutawika's avatar
lintangsutawika committed
55
- [x] XStoryCloze
Lintang Sutawika's avatar
Lintang Sutawika committed
56
- [x] XWinograd
Lintang Sutawika's avatar
Lintang Sutawika committed
57
- [x] PAWS-X
lintangsutawika's avatar
lintangsutawika committed
58
- [x] XNLI
lintangsutawika's avatar
lintangsutawika committed
59
- [x] MGSM
Hailey Schoelkopf's avatar
Hailey Schoelkopf committed
60
- [ ] SCROLLS
61
- [x] Babi
haileyschoelkopf's avatar
haileyschoelkopf committed
62
63
64
65
66
67
68
69
70
71
72

# Novel Tasks
Tasks added in the revamped harness that were not previously available. Again, a strikethrough denotes checking performed *against the original task's implementation or published results introducing the task*.

# Task Wishlist

- [ ] TheoremQA
- [ ] Theorem Proving evaluations
- [ ] Chain of Thought
- [ ] Self-consistency ; Least-to-Most prompting, etc.
- [ ] Summarization Tasks
lintangsutawika's avatar
lintangsutawika committed
73
- [ ] Anthropic Model-Written Evals