Skip to content
GitLab
Menu
Projects
Groups
Snippets
Loading...
Help
Help
Support
Community forum
Keyboard shortcuts
?
Submit feedback
Contribute to GitLab
Sign in / Register
Toggle navigation
Menu
Open sidebar
gaoqiong
lm-evaluation-harness
Commits
da8af971
Commit
da8af971
authored
Aug 15, 2023
by
lintangsutawika
Browse files
update readme docs
parent
cfbc5a8a
Changes
2
Hide whitespace changes
Inline
Side-by-side
Showing
2 changed files
with
27 additions
and
10 deletions
+27
-10
lm_eval/tasks/README.md
lm_eval/tasks/README.md
+1
-1
lm_eval/tasks/mc_taco/README.md
lm_eval/tasks/mc_taco/README.md
+26
-9
No files found.
lm_eval/tasks/README.md
View file @
da8af971
...
@@ -13,7 +13,7 @@ Boxes should be checked iff tasks are implemented in the refactor and tested for
...
@@ -13,7 +13,7 @@ Boxes should be checked iff tasks are implemented in the refactor and tested for
-
[x] Wikitext
-
[x] Wikitext
-
[x] PiQA
-
[x] PiQA
-
[x] PROST
-
[x] PROST
-
[
] MCTACO
(Lintang)
-
[
x
] MCTACO
-
[x] Pubmed QA
-
[x] Pubmed QA
-
[x] SciQ
-
[x] SciQ
-
[ ] QASPER
-
[ ] QASPER
...
...
lm_eval/tasks/mc_taco/README.md
View file @
da8af971
#
Task-name
#
MC Taco
### Paper
### Paper
Title:
`
paper title goes here
`
Title:
`
"Going on a vacation" takes longer than "Going for a walk": A Study of Temporal Commonsense Understanding
`
Abstract:
`link to paper PDF or arXiv abstract goes here`
Abstract:
https://arxiv.org/abs/1909.03065
`Short description of paper / benchmark goes here:`
MC-TACO is a dataset of 13k question-answer pairs that require temporal commonsense
comprehension. The dataset contains five temporal properties, (1) duration (how long
an event takes), (2) temporal ordering (typical order of events), (3) typical time
(when an event occurs), (4) frequency (how often an event occurs), and (5) stationarity
(whether a state is maintained for a very long time or indefinitely).
Homepage:
`homepage to the benchmark's website goes here, if applicable`
WARNING: Running this task with a
`--limit`
arg will give misleading results! The
corresponding dataset is structured such that each multiple-choice-question gathered
by the authors is split into question-option pairs, where each such pair gets
siloed into an individual document for plausibility testing. Because the harness
shuffles these documents, setting
`--limit`
will likely "cut off" certain candidate
answers. This is a problem because the task's metrics require an exhaustive evaluation
of a question's options. See section 4 of the paper for details.
Homepage: https://leaderboard.allenai.org/mctaco/submissions/public
### Citation
### Citation
...
@@ -16,11 +28,16 @@ Homepage: `homepage to the benchmark's website goes here, if applicable`
...
@@ -16,11 +28,16 @@ Homepage: `homepage to the benchmark's website goes here, if applicable`
BibTeX-formatted citation goes here
BibTeX-formatted citation goes here
```
```
### Subtasks
### Groups and Tasks
#### Groups
*
Not part of a group yet.
#### Tasks
*
`mc_taco`
List or describe tasks defined in this folder, and their names here:
*
`task_name`
:
`1-sentence description of what this particular task does`
*
`task_name2`
: .....
### Checklist
### Checklist
...
...
Write
Preview
Markdown
is supported
0%
Try again
or
attach a new file
.
Attach a file
Cancel
You are about to add
0
people
to the discussion. Proceed with caution.
Finish editing this message first!
Cancel
Please
register
or
sign in
to comment