README.md 1.84 KB
Newer Older
haileyschoelkopf's avatar
haileyschoelkopf committed
1
2
3
# v1.0 Tasks
This list keeps track of which tasks' implementations have been ported to YAML / v2.0 of the Eval Harness.

4
Boxes should be checked iff tasks are implemented in the refactor and tested for regression. Tasks should be struck through if checked *against original introducing paper* implementation or popularizing implementation. (WIP) Denotes that there exists a PR or person working on this task already.
haileyschoelkopf's avatar
haileyschoelkopf committed
5

Hailey Schoelkopf's avatar
Hailey Schoelkopf committed
6
- [ ] Glue (WIP)
haileyschoelkopf's avatar
haileyschoelkopf committed
7
- [x] SuperGlue
haileyschoelkopf's avatar
haileyschoelkopf committed
8
9
10
11
- [ ] CoQA
- [ ] DROP
- [x] ~~Lambada~~
- [x] Lambada (Cloze variants)
12
- [x] ~~Lambada (Multilingual)~~
haileyschoelkopf's avatar
haileyschoelkopf committed
13
14
- [x] Wikitext
- [x] PiQA
Hailey Schoelkopf's avatar
Hailey Schoelkopf committed
15
- [x] PROST
haileyschoelkopf's avatar
haileyschoelkopf committed
16
- [ ] MCTACO
17
- [x] Pubmed QA
haileyschoelkopf's avatar
haileyschoelkopf committed
18
19
- [x] SciQ
- [ ] QASPER
Hailey Schoelkopf's avatar
Hailey Schoelkopf committed
20
- [x] QA4MRE
haileyschoelkopf's avatar
haileyschoelkopf committed
21
22
- [ ] TriviaQA
- [x] AI2 ARC
Hailey Schoelkopf's avatar
Hailey Schoelkopf committed
23
- [ ] LogiQA (WIP)
24
- [x] HellaSwag
Benjamin Fattori's avatar
Benjamin Fattori committed
25
- [x] SWAG
Hailey Schoelkopf's avatar
Hailey Schoelkopf committed
26
- [x] OpenBookQA
27
- [x] RACE
Hailey Schoelkopf's avatar
Hailey Schoelkopf committed
28
- [ ] LogiQA (WIP)
29
- [x] HellaSwag
Hailey Schoelkopf's avatar
Hailey Schoelkopf committed
30
- [x] SWAG
Hailey Schoelkopf's avatar
Hailey Schoelkopf committed
31
- [x] OpenBookQA
Hailey Schoelkopf's avatar
Hailey Schoelkopf committed
32
- [ ] SQuADv2 (WIP)
Hailey Schoelkopf's avatar
Hailey Schoelkopf committed
33
- [x] RACE
34
35
- [x] HeadQA (WIP)
- [ ] MathQA (WIP)
haileyschoelkopf's avatar
haileyschoelkopf committed
36
37
- [ ] WebQs
- [ ] WSC273
38
- [x] Winogrande
haileyschoelkopf's avatar
haileyschoelkopf committed
39
- [x] ANLI
haileyschoelkopf's avatar
haileyschoelkopf committed
40
41
42
- [ ] Hendrycks Ethics
- [ ] TruthfulQA
- [ ] MuTual
Hailey Schoelkopf's avatar
Hailey Schoelkopf committed
43
- [ ] Hendrycks Math (WIP)
Hailey Schoelkopf's avatar
Hailey Schoelkopf committed
44
- [ ] Asdiv (WIP)
haileyschoelkopf's avatar
haileyschoelkopf committed
45
- [ ] GSM8k
46
- [x] Arithmetic
haileyschoelkopf's avatar
haileyschoelkopf committed
47
48
- [ ] MMMLU
- [ ] Translation (WMT) suite
haileyschoelkopf's avatar
haileyschoelkopf committed
49
- [x] Unscramble
haileyschoelkopf's avatar
haileyschoelkopf committed
50
51
- [x] ~~Pile (perplexity)~~
- [ ] BLiMP
haileyschoelkopf's avatar
haileyschoelkopf committed
52
- [x] ToxiGen
Hailey Schoelkopf's avatar
Hailey Schoelkopf committed
53
54
- [ ] StoryCloze
- [ ] NaturalQs
haileyschoelkopf's avatar
haileyschoelkopf committed
55
56
57
58
59
60
61
62
- [ ] CrowS-Pairs
- [ ] XCopa
- [ ] BIG-Bench
- [ ] XStoryCloze
- [ ] XWinograd
- [ ] PAWS-X
- [ ] XNLI
- [ ] MGSM
Hailey Schoelkopf's avatar
Hailey Schoelkopf committed
63
64
- [ ] SCROLLS
- [ ] JSON Task (reference: https://github.com/EleutherAI/lm-evaluation-harness/pull/481)
Hailey Schoelkopf's avatar
Hailey Schoelkopf committed
65
- [ ] Babi
haileyschoelkopf's avatar
haileyschoelkopf committed
66
67
68
69
70
71
72
73
74
75
76

# Novel Tasks
Tasks added in the revamped harness that were not previously available. Again, a strikethrough denotes checking performed *against the original task's implementation or published results introducing the task*.

# Task Wishlist

- [ ] TheoremQA
- [ ] Theorem Proving evaluations
- [ ] Chain of Thought
- [ ] Self-consistency ; Least-to-Most prompting, etc.
- [ ] Summarization Tasks
lintangsutawika's avatar
lintangsutawika committed
77
- [ ] Anthropic Model-Written Evals