merged with latest

d627333a · lintangsutawika · 4156a005 · 4cda3a1c · d627333a · d627333a
Commit d627333a authored Aug 29, 2023 by lintangsutawika
20 changed files
--- a/lm_eval/tasks/lambada_multilingual/lambada_mt_it.yaml
+++ b/lm_eval/tasks/lambada_multilingual/lambada_mt_it.yaml
 include: lambada_mt_en.yaml
-group:
-  - lambada_multilingual
-  - loglikelihood
-  - perplexity
 task: lambada_openai_mt_it
 dataset_name: it
--- a/lm_eval/tasks/logiqa/README.md
+++ b/lm_eval/tasks/logiqa/README.md
+# LogiQA
+### Paper
+Title: `LogiQA: A Challenge Dataset for Machine Reading Comprehension with Logical Reasoning`
+Abstract: https://arxiv.org/abs/2007.08124
+LogiQA is a dataset for testing human logical reasoning. It consists of 8,678 QA
+instances, covering multiple types of deductive reasoning. Results show that state-
+of-the-art neural models perform by far worse than human ceiling. The dataset can
+also serve as a benchmark for reinvestigating logical AI under the deep learning
+NLP setting.
+Homepage: https://github.com/lgw863/LogiQA-dataset
+### Citation
+```
+@misc{liu2020logiqa,
+    title={LogiQA: A Challenge Dataset for Machine Reading Comprehension with Logical Reasoning},
+    author={Jian Liu and Leyang Cui and Hanmeng Liu and Dandan Huang and Yile Wang and Yue Zhang},
+    year={2020},
+    eprint={2007.08124},
+    archivePrefix={arXiv},
+    primaryClass={cs.CL}
+}
+```
+### Groups and Tasks
+#### Groups
+* Not part of a group yet
+#### Tasks
+* `logiqa`
+### Checklist
+For adding novel benchmarks/datasets to the library:
+* [ ] Is the task an existing benchmark in the literature?
+  * [ ] Have you referenced the original paper that introduced the task?
+  * [ ] If yes, does the original paper provide a reference implementation? If so, have you checked against the reference implementation and documented how to run such a test?
+If other tasks on this dataset are already supported:
+* [ ] Is the "Main" variant of this task clearly denoted?
+* [ ] Have you provided a short sentence in a README on what each new variant adds / evaluates?
+* [ ] Have you noted which, if any, published evaluation setups are matched by this variant?
--- a/lm_eval/tasks/logiqa/logiqa.yaml
+++ b/lm_eval/tasks/logiqa/logiqa.yaml
-group:
-  - multiple_choice
 task: logiqa
 dataset_path: EleutherAI/logiqa
 dataset_name: logiqa

--- a/lm_eval/tasks/logiqa2/README.md
+++ b/lm_eval/tasks/logiqa2/README.md
@@ -25,15 +25,19 @@ Homepage: https://github.com/csitfun/LogiQA2.0
  doi={10.1109/TASLP.2023.3293046}}
 ```
-### Subtasks
+### Groups and Tasks
-`logiqa2_zh`: The original dataset in Chinese.
+#### Groups
-`logiqa2_NLI`: The NLI version of the dataset converted from the MRC version.
+* Not part of a group yet
-`logieval`: Prompt based; https://github.com/csitfun/LogiEval
+#### Tasks
-The subtasks have not been verified yet.
+* `logiqa2_zh`: The original dataset in Chinese.
+* `logiqa2_NLI`: The NLI version of the dataset converted from the MRC version.
+* `logieval`: Prompt based; https://github.com/csitfun/LogiEval
+NOTE! The subtasks have not been verified yet.
 ### Checklist

--- a/lm_eval/tasks/logiqa2/logieval.yaml
+++ b/lm_eval/tasks/logiqa2/logieval.yaml
-group:
-  - greedy_until
 task: logieval
 dataset_path: baber/logiqa2
 dataset_name: logieval

--- a/lm_eval/tasks/logiqa2/logiqa2.yaml
+++ b/lm_eval/tasks/logiqa2/logiqa2.yaml
-group:
-  - multiple_choice
 task: logiqa2
 dataset_path: baber/logiqa2
 dataset_name: logiqa2

--- a/lm_eval/tasks/mathqa/README.md
+++ b/lm_eval/tasks/mathqa/README.md
@@ -25,7 +25,13 @@ Homepage: https://math-qa.github.io/math-QA/
 }
 ```
-### Subtasks
+### Groups and Tasks
+#### Groups
+* `math_word_problems`
+#### Tasks
 * `mathqa`: The MathQA dataset, as a multiple choice dataset where the answer choices are not in context.

--- a/lm_eval/tasks/mathqa/mathqa.yaml
+++ b/lm_eval/tasks/mathqa/mathqa.yaml
 group:
-  - multiple_choice
  - math_word_problems
 task: mathqa
 dataset_path: math_qa

--- a/lm_eval/tasks/mc_taco/README.md
+++ b/lm_eval/tasks/mc_taco/README.md
+# MC Taco
+### Paper
+Title: `"Going on a vacation" takes longer than "Going for a walk": A Study of Temporal Commonsense Understanding`
+Abstract: https://arxiv.org/abs/1909.03065
+MC-TACO is a dataset of 13k question-answer pairs that require temporal commonsense
+comprehension. The dataset contains five temporal properties, (1) duration (how long
+an event takes), (2) temporal ordering (typical order of events), (3) typical time
+(when an event occurs), (4) frequency (how often an event occurs), and (5) stationarity
+(whether a state is maintained for a very long time or indefinitely).
+WARNING: Running this task with a `--limit` arg will give misleading results! The
+corresponding dataset is structured such that each multiple-choice-question gathered
+by the authors is split into question-option pairs, where each such pair gets
+siloed into an individual document for plausibility testing. Because the harness
+shuffles these documents, setting `--limit` will likely "cut off" certain candidate
+answers. This is a problem because the task's metrics require an exhaustive evaluation
+of a question's options. See section 4 of the paper for details.
+Homepage: https://leaderboard.allenai.org/mctaco/submissions/public
+### Citation
+```
+BibTeX-formatted citation goes here
+```
+### Groups and Tasks
+#### Groups
+* Not part of a group yet.
+#### Tasks
+* `mc_taco`
+### Checklist
+For adding novel benchmarks/datasets to the library:
+* [ ] Is the task an existing benchmark in the literature?
+  * [ ] Have you referenced the original paper that introduced the task?
+  * [ ] If yes, does the original paper provide a reference implementation? If so, have you checked against the reference implementation and documented how to run such a test?
+If other tasks on this dataset are already supported:
+* [ ] Is the "Main" variant of this task clearly denoted?
+* [ ] Have you provided a short sentence in a README on what each new variant adds / evaluates?
+* [ ] Have you noted which, if any, published evaluation setups are matched by this variant?
--- a/lm_eval/tasks/mc_taco/default.yaml
+++ b/lm_eval/tasks/mc_taco/default.yaml
+task: mc_taco
+dataset_path: mc_taco
+output_type: multiple_choice
+validation_split: validation
+test_split: test
+doc_to_text: "{{sentence}}\nQuestion: {{question}}\nAnswer: {{answer}}\nPlausible:"
+doc_to_target: label
+doc_to_choice: ["no", "yes"]
+should_decontaminate: true
+doc_to_decontamination_query: "{{question}} {{sentence}}"
+metric_list:
+  - metric: acc
+  - metric: f1
--- a/lm_eval/tasks/openbookqa/README.md
+++ b/lm_eval/tasks/openbookqa/README.md
+# Task-name
+### Paper
+Title: `Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering`
+Abstract: https://arxiv.org/abs/1809.02789
+OpenBookQA is a question-answering dataset modeled after open book exams for
+assessing human understanding of a subject. It consists of 5,957 multiple-choice
+elementary-level science questions (4,957 train, 500 dev, 500 test), which probe
+the understanding of a small “book” of 1,326 core science facts and the application
+of these facts to novel situations. For training, the dataset includes a mapping
+from each question to the core science fact it was designed to probe. Answering
+OpenBookQA questions requires additional broad common knowledge, not contained
+in the book. The questions, by design, are answered incorrectly by both a retrieval-
+based algorithm and a word co-occurrence algorithm.
+Homepage: https://allenai.org/data/open-book-qa
+### Citation
+```
+@inproceedings{OpenBookQA2018,
+    title={Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering},
+    author={Todor Mihaylov and Peter Clark and Tushar Khot and Ashish Sabharwal},
+    booktitle={EMNLP},
+    year={2018}
+}
+```
+### Groups and Tasks
+#### Groups
+* Not part of a group yet
+#### Tasks
+* `openbookqa`
+### Checklist
+For adding novel benchmarks/datasets to the library:
+* [ ] Is the task an existing benchmark in the literature?
+  * [ ] Have you referenced the original paper that introduced the task?
+  * [ ] If yes, does the original paper provide a reference implementation? If so, have you checked against the reference implementation and documented how to run such a test?
+If other tasks on this dataset are already supported:
+* [ ] Is the "Main" variant of this task clearly denoted?
+* [ ] Have you provided a short sentence in a README on what each new variant adds / evaluates?
+* [ ] Have you noted which, if any, published evaluation setups are matched by this variant?
--- a/lm_eval/tasks/openbookqa/openbookqa.yaml
+++ b/lm_eval/tasks/openbookqa/openbookqa.yaml
-group:
-  - multiple_choice
 task: openbookqa
 dataset_path: openbookqa
 dataset_name: main

--- a/lm_eval/tasks/pile/README.md
+++ b/lm_eval/tasks/pile/README.md
 # The Pile
 ### Paper
-The Pile: An 800GB Dataset of Diverse Text for Language Modeling
+Title: The Pile: An 800GB Dataset of Diverse Text for Language Modeling
-https://arxiv.org/pdf/2101.00027.pdf
+Abstract: https://arxiv.org/abs/2101.00027
 The Pile is a 825 GiB diverse, open source language modelling data set that consists
 of 22 smaller, high-quality datasets combined together. To score well on Pile
@@ -21,3 +22,47 @@ Homepage: https://pile.eleuther.ai/
  year={2020}
 }
 ```
+### Groups and Tasks
+#### Groups
+* `pile`
+#### Tasks
+* `pile_arxiv`
+* `pile_bookcorpus2`
+* `pile_books3`
+* `pile_dm-mathematics`
+* `pile_enron`
+* `pile_europarl`
+* `pile_freelaw`
+* `pile_github`
+* `pile_gutenberg`
+* `pile_hackernews`
+* `pile_nih-exporter`
+* `pile_opensubtitles`
+* `pile_openwebtext2`
+* `pile_philpapers`
+* `pile_pile-cc`
+* `pile_pubmed-abstracts`
+* `pile_pubmed-central`
+* `pile_stackexchange`
+* `pile_ubuntu-irc`
+* `pile_uspto`
+* `pile_wikipedia`
+* `pile_youtubesubtitles`
+### Checklist
+For adding novel benchmarks/datasets to the library:
+* [ ] Is the task an existing benchmark in the literature?
+  * [ ] Have you referenced the original paper that introduced the task?
+  * [ ] If yes, does the original paper provide a reference implementation? If so, have you checked against the reference implementation and documented how to run such a test?
+If other tasks on this dataset are already supported:
+* [ ] Is the "Main" variant of this task clearly denoted?
+* [ ] Have you provided a short sentence in a README on what each new variant adds / evaluates?
+* [ ] Have you noted which, if any, published evaluation setups are matched by this variant?
--- a/lm_eval/tasks/pile/pile_arxiv.yaml
+++ b/lm_eval/tasks/pile/pile_arxiv.yaml
 group:
  - pile
-  - perplexity
-  - loglikelihood_rolling
 task: pile_arxiv
 dataset_path: EleutherAI/pile
 dataset_name: pile_arxiv

--- a/lm_eval/tasks/piqa/README.md
+++ b/lm_eval/tasks/piqa/README.md
+# PIQA
+### Paper
+Title: `PIQA: Reasoning about Physical Commonsense in Natural Language`
+Abstract: https://arxiv.org/abs/1911.11641
+Physical Interaction: Question Answering (PIQA) is a physical commonsense
+reasoning and a corresponding benchmark dataset. PIQA was designed to investigate
+the physical knowledge of existing models. To what extent are current approaches
+actually learning about the world?
+Homepage: https://yonatanbisk.com/piqa/
+### Citation
+```
+@inproceedings{Bisk2020,
+    author = {Yonatan Bisk and Rowan Zellers and
+            Ronan Le Bras and Jianfeng Gao
+            and Yejin Choi},
+    title = {PIQA: Reasoning about Physical Commonsense in
+           Natural Language},
+    booktitle = {Thirty-Fourth AAAI Conference on
+               Artificial Intelligence},
+    year = {2020},
+}
+```
+### Groups and Tasks
+#### Groups
+* Not part of a group yet.
+#### Tasks
+* `piqa`
+### Checklist
+For adding novel benchmarks/datasets to the library:
+* [ ] Is the task an existing benchmark in the literature?
+  * [ ] Have you referenced the original paper that introduced the task?
+  * [ ] If yes, does the original paper provide a reference implementation? If so, have you checked against the reference implementation and documented how to run such a test?
+If other tasks on this dataset are already supported:
+* [ ] Is the "Main" variant of this task clearly denoted?
+* [ ] Have you provided a short sentence in a README on what each new variant adds / evaluates?
+* [ ] Have you noted which, if any, published evaluation setups are matched by this variant?
--- a/lm_eval/tasks/piqa/piqa.yaml
+++ b/lm_eval/tasks/piqa/piqa.yaml
-group:
-  - multiple_choice
 task: piqa
 dataset_path: piqa
 dataset_name: null

--- a/lm_eval/tasks/prost/README.md
+++ b/lm_eval/tasks/prost/README.md
+# PROST
+### Paper
+Title: `PROST: Physical Reasoning about Objects Through Space and Time`
+Abstract: https://arxiv.org/abs/2106.03634
+PROST, Physical Reasoning about Objects Through Space and Time, is a dataset
+consisting of 18,736 multiple-choice questions made from 14 manually curated
+templates, covering 10 physical reasoning concepts. All questions are designed
+to probe both causal and masked language models in a zero-shot setting.
+NOTE: PROST is limited to the zero-shot setting to adhere to authors' intentions
+as discussed in section 7 of the paper: "We hope that the community will use
+this dataset in the intended way: in a zero-shot setting to probe models which
+have been trained on data not specifically collected to succeed on PROST."
+Homepage: https://github.com/nala-cub/prost
+### Citation
+```
+@inproceedings{aroca-ouellette-etal-2021-prost,
+    title = "{PROST}: {P}hysical Reasoning about Objects through Space and Time",
+    author = "Aroca-Ouellette, St{\'e}phane  and
+      Paik, Cory  and
+      Roncone, Alessandro  and
+      Kann, Katharina",
+    booktitle = "Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021",
+    month = aug,
+    year = "2021",
+    address = "Online",
+    publisher = "Association for Computational Linguistics",
+    url = "https://aclanthology.org/2021.findings-acl.404",
+    pages = "4597--4608",
+}
+```
+### Groups and Tasks
+#### Groups
+* Not part of a group yet.
+#### Tasks
+* `prost`
+### Checklist
+For adding novel benchmarks/datasets to the library:
+* [ ] Is the task an existing benchmark in the literature?
+  * [ ] Have you referenced the original paper that introduced the task?
+  * [ ] If yes, does the original paper provide a reference implementation? If so, have you checked against the reference implementation and documented how to run such a test?
+If other tasks on this dataset are already supported:
+* [ ] Is the "Main" variant of this task clearly denoted?
+* [ ] Have you provided a short sentence in a README on what each new variant adds / evaluates?
+* [ ] Have you noted which, if any, published evaluation setups are matched by this variant?
--- a/lm_eval/tasks/prost/corypaik_prost.yaml
+++ b/lm_eval/tasks/prost/corypaik_prost.yaml
-group:
-  - multiple_choice
 task: prost
 dataset_path: corypaik/prost
 dataset_name: null

--- a/lm_eval/tasks/pubmedqa/README.md
+++ b/lm_eval/tasks/pubmedqa/README.md
+# PubMedQA
+### Paper
+Title: `PubMedQA: A Dataset for Biomedical Research Question Answering`
+Abstract: https://arxiv.org/abs/1909.06146
+PubMedQA is a novel biomedical question answering (QA) dataset collected from
+PubMed abstracts. The task of PubMedQA is to answer research questions with
+yes/no/maybe (e.g.: Do preoperative statins reduce atrial fibrillation after
+coronary artery bypass grafting?) using the corresponding abstracts. PubMedQA
+has 1k expert-annotated, 61.2k unlabeled and 211.3k artificially generated QA
+instances. Each PubMedQA instance is composed of (1) a question which is either
+an existing research article title or derived from one, (2) a context which is
+the corresponding abstract without its conclusion, (3) a long answer, which is
+the conclusion of the abstract and, presumably, answers the research question,
+and (4) a yes/no/maybe answer which summarizes the conclusion.
+Homepage: https://pubmedqa.github.io/
+### Citation
+```
+@inproceedings{jin2019pubmedqa,
+    title={PubMedQA: A Dataset for Biomedical Research Question Answering},
+    author={Jin, Qiao and Dhingra, Bhuwan and Liu, Zhengping and Cohen, William and Lu, Xinghua},
+    booktitle={Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)},
+    pages={2567--2577},
+    year={2019}
+}
+```
+### Groups and Tasks
+#### Groups
+* Not part of a group yet
+#### Tasks
+* `pubmed_qa`
+### Checklist
+For adding novel benchmarks/datasets to the library:
+* [ ] Is the task an existing benchmark in the literature?
+  * [ ] Have you referenced the original paper that introduced the task?
+  * [ ] If yes, does the original paper provide a reference implementation? If so, have you checked against the reference implementation and documented how to run such a test?
+If other tasks on this dataset are already supported:
+* [ ] Is the "Main" variant of this task clearly denoted?
+* [ ] Have you provided a short sentence in a README on what each new variant adds / evaluates?
+* [ ] Have you noted which, if any, published evaluation setups are matched by this variant?
--- a/lm_eval/tasks/pubmedqa/pubmedqa.yaml
+++ b/lm_eval/tasks/pubmedqa/pubmedqa.yaml
-group:
-  - multiple_choice
 task: pubmed_qa
 dataset_path: pubmed_qa
 dataset_name: pqa_labeled