Group agg rework (#1741)

* add greoup_config arg * add a group config that allows disabling table for group score and group aggregate in general * fixed size configuration * adjust config * add group config * adjust mmlu to use group_config * fixed args input in aggregate_subtask_metrics * fixed issues related to printing alias of group and updated yaml * update all mmlu variants to include group_config * edit format * modify mmlu tasks * adjust group to also be a configurable group * add configurable group * simplify get_task_list * adjust group scoring with using ConfigurableGroup * adjust args * update mmlu * update mmlu * update to work with new group and task configuration * readd group_agg * readd files * move prepare_print_tasks to evaluator_utils * sort set to False by default, fix predict_only arg * add version for groups * reversed task list * update additional condition when loading a group in a group yaml * update truthfulqa * add description regarding tags replacing group * replace group to tag * fixed conditional statement * remove warning * update loading of task group and newly added tags * reformat with pre-commit * fixed info log * update * fix bug * fix bug * use task id to differentiate tasks * convert all groups to configurable groups * use task_id * reformat * add task_id for python tasks as well * add task_id for python tasks as well * add task_id for python tasks as well * revert truthfulqa * revert mmlu tasks * new mmlu config * new group config parameter `tag_to_task` * Update truthfulqa_mc2.yaml * reformate * add _process_group_config * adjust task_id * add get_subtask_list function to get proper subtask list * group config to_dict update * remove tag check * update mmlu * fix config passing issues * add test yaml * format fix * add documentation * corner case for single tag being called * fix indentation * formatting * update all mmlu variants * Update docs/task_guide.md Co-authored-by: Hailey Schoelkopf <65563625+haileyschoelkopf@users.noreply.github.com> * remove group_alias * Update docs/task_guide.md Co-authored-by: Hailey Schoelkopf <65563625+haileyschoelkopf@users.noreply.github.com> * remove version for metadata * Update docs/task_guide.md Co-authored-by: Hailey Schoelkopf <65563625+haileyschoelkopf@users.noreply.github.com> * update mmlu/ * removed " " in make_table * change how aggregate_metric is loaded * change how aggregate_metric is loaded * update aggregate_metric arg * update format * update format * some docs fixes * add groups for agieval, aexams, aclue * add more explicit aggregation groups * add more groupings / tags distinctions * add more groupings * more groupings * add many explicit group configs * add many explicit group configs * add more explicit group configs * add more explicit group configs * add more error msgs, agg_metric -> agg_metric_list * some docs updates * update task_id to be updateable and uses group:task format * make KMMLU a tag for now * update docs * don't duplicate task names * fix merge conflicts? * giving this a try * clean up diff * switch mmlu variants over to using * don't use to-be-deprecated group: config field in overview notebook * Python tasks which subclass ConfigurableTask now run * update mmlu * pre-commit format * fixed sorting for multi-level printing * move group api to separate file * fix bbh aggregation filter usage * track api/group.py * adjust group and tags loading * make explicit group configs for leaderboard and other newer tasks * fix arabicmmlu * update * change arabicmmlu template name??? * update group alias * fix printing bugs * check table printing is correct ; update tests * use mmlu_stem to have a group included in print tests --------- Co-authored-by: Hailey Schoelkopf <65563625+haileyschoelkopf@users.noreply.github.com> Co-authored-by: haileyschoelkopf <hailey@eleuther.ai>

Group agg rework (#1741)
* add greoup_config arg * add a group config that allows disabling table for group score and group aggregate in general * fixed size configuration * adjust config * add group config * adjust mmlu to use group_config * fixed args input in aggregate_subtask_metrics * fixed issues related to printing alias of group and updated yaml * update all mmlu variants to include group_config * edit format * modify mmlu tasks * adjust group to also be a configurable group * add configurable group * simplify get_task_list * adjust group scoring with using ConfigurableGroup * adjust args * update mmlu * update mmlu * update to work with new group and task configuration * readd group_agg * readd files * move prepare_print_tasks to evaluator_utils * sort set to False by default, fix predict_only arg * add version for groups * reversed task list * update additional condition when loading a group in a group yaml * update truthfulqa * add description regarding tags replacing group * replace group to tag * fixed conditional statement * remove warning * update loading of task group and newly added tags * reformat with pre-commit * fixed info log * update * fix bug * fix bug * use task id to differentiate tasks * convert all groups to configurable groups * use task_id * reformat * add task_id for python tasks as well * add task_id for python tasks as well * add task_id for python tasks as well * revert truthfulqa * revert mmlu tasks * new mmlu config * new group config parameter `tag_to_task` * Update truthfulqa_mc2.yaml * reformate * add _process_group_config * adjust task_id * add get_subtask_list function to get proper subtask list * group config to_dict update * remove tag check * update mmlu * fix config passing issues * add test yaml * format fix * add documentation * corner case for single tag being called * fix indentation * formatting * update all mmlu variants * Update docs/task_guide.md Co-authored-by: Hailey Schoelkopf <65563625+haileyschoelkopf@users.noreply.github.com> * remove group_alias * Update docs/task_guide.md Co-authored-by: Hailey Schoelkopf <65563625+haileyschoelkopf@users.noreply.github.com> * remove version for metadata * Update docs/task_guide.md Co-authored-by: Hailey Schoelkopf <65563625+haileyschoelkopf@users.noreply.github.com> * update mmlu/ * removed " " in make_table * change how aggregate_metric is loaded * change how aggregate_metric is loaded * update aggregate_metric arg * update format * update format * some docs fixes * add groups for agieval, aexams, aclue * add more explicit aggregation groups * add more groupings / tags distinctions * add more groupings * more groupings * add many explicit group configs * add many explicit group configs * add more explicit group configs * add more explicit group configs * add more error msgs, agg_metric -> agg_metric_list * some docs updates * update task_id to be updateable and uses group:task format * make KMMLU a tag for now * update docs * don't duplicate task names * fix merge conflicts? * giving this a try * clean up diff * switch mmlu variants over to using * don't use to-be-deprecated group: config field in overview notebook * Python tasks which subclass ConfigurableTask now run * update mmlu * pre-commit format * fixed sorting for multi-level printing * move group api to separate file * fix bbh aggregation filter usage * track api/group.py * adjust group and tags loading * make explicit group configs for leaderboard and other newer tasks * fix arabicmmlu * update * change arabicmmlu template name??? * update group alias * fix printing bugs * check table printing is correct ; update tests * use mmlu_stem to have a group included in print tests --------- Co-authored-by: Hailey Schoelkopf <65563625+haileyschoelkopf@users.noreply.github.com> Co-authored-by: haileyschoelkopf <hailey@eleuther.ai>
517aadc4 · Lintang Sutawika · GitHub · 5a7ed3ee · 517aadc4 · 517aadc4
Unverified Commit 517aadc4 authored Jul 08, 2024 by Lintang Sutawika Committed by GitHub Jul 08, 2024
20 changed files
--- a/lm_eval/tasks/mmlu/default/mmlu_logical_fallacies.yaml
+++ b/lm_eval/tasks/mmlu/default/mmlu_logical_fallacies.yaml
 "dataset_name": "logical_fallacies"
 "description": "The following are multiple choice questions (with answers) about logical\
  \ fallacies.\n\n"
-"group": "mmlu_humanities"
-"group_alias": "humanities"
+"tag": "mmlu_humanities_tasks"
 "include": "_default_template_yaml"
 "task": "mmlu_logical_fallacies"
 "task_alias": "logical_fallacies"
--- a/lm_eval/tasks/mmlu/default/mmlu_machine_learning.yaml
+++ b/lm_eval/tasks/mmlu/default/mmlu_machine_learning.yaml
 "dataset_name": "machine_learning"
 "description": "The following are multiple choice questions (with answers) about machine\
  \ learning.\n\n"
-"group": "mmlu_stem"
-"group_alias": "stem"
+"tag": "mmlu_stem_tasks"
 "include": "_default_template_yaml"
 "task": "mmlu_machine_learning"
 "task_alias": "machine_learning"
--- a/lm_eval/tasks/mmlu/default/mmlu_management.yaml
+++ b/lm_eval/tasks/mmlu/default/mmlu_management.yaml
 "dataset_name": "management"
 "description": "The following are multiple choice questions (with answers) about management.\n\
  \n"
-"group": "mmlu_other"
-"group_alias": "other"
+"tag": "mmlu_other_tasks"
 "include": "_default_template_yaml"
 "task": "mmlu_management"
 "task_alias": "management"
--- a/lm_eval/tasks/mmlu/default/mmlu_marketing.yaml
+++ b/lm_eval/tasks/mmlu/default/mmlu_marketing.yaml
 "dataset_name": "marketing"
 "description": "The following are multiple choice questions (with answers) about marketing.\n\
  \n"
-"group": "mmlu_other"
-"group_alias": "other"
+"tag": "mmlu_other_tasks"
 "include": "_default_template_yaml"
 "task": "mmlu_marketing"
 "task_alias": "marketing"
--- a/lm_eval/tasks/mmlu/default/mmlu_medical_genetics.yaml
+++ b/lm_eval/tasks/mmlu/default/mmlu_medical_genetics.yaml
 "dataset_name": "medical_genetics"
 "description": "The following are multiple choice questions (with answers) about medical\
  \ genetics.\n\n"
-"group": "mmlu_other"
-"group_alias": "other"
+"tag": "mmlu_other_tasks"
 "include": "_default_template_yaml"
 "task": "mmlu_medical_genetics"
 "task_alias": "medical_genetics"
--- a/lm_eval/tasks/mmlu/default/mmlu_miscellaneous.yaml
+++ b/lm_eval/tasks/mmlu/default/mmlu_miscellaneous.yaml
 "dataset_name": "miscellaneous"
 "description": "The following are multiple choice questions (with answers) about miscellaneous.\n\
  \n"
-"group": "mmlu_other"
-"group_alias": "other"
+"tag": "mmlu_other_tasks"
 "include": "_default_template_yaml"
 "task": "mmlu_miscellaneous"
 "task_alias": "miscellaneous"
--- a/lm_eval/tasks/mmlu/default/mmlu_moral_disputes.yaml
+++ b/lm_eval/tasks/mmlu/default/mmlu_moral_disputes.yaml
 "dataset_name": "moral_disputes"
 "description": "The following are multiple choice questions (with answers) about moral\
  \ disputes.\n\n"
-"group": "mmlu_humanities"
-"group_alias": "humanities"
+"tag": "mmlu_humanities_tasks"
 "include": "_default_template_yaml"
 "task": "mmlu_moral_disputes"
 "task_alias": "moral_disputes"
--- a/lm_eval/tasks/mmlu/default/mmlu_moral_scenarios.yaml
+++ b/lm_eval/tasks/mmlu/default/mmlu_moral_scenarios.yaml
 "dataset_name": "moral_scenarios"
 "description": "The following are multiple choice questions (with answers) about moral\
  \ scenarios.\n\n"
-"group": "mmlu_humanities"
-"group_alias": "humanities"
+"tag": "mmlu_humanities_tasks"
 "include": "_default_template_yaml"
 "task": "mmlu_moral_scenarios"
 "task_alias": "moral_scenarios"
--- a/lm_eval/tasks/mmlu/default/mmlu_nutrition.yaml
+++ b/lm_eval/tasks/mmlu/default/mmlu_nutrition.yaml
 "dataset_name": "nutrition"
 "description": "The following are multiple choice questions (with answers) about nutrition.\n\
  \n"
-"group": "mmlu_other"
-"group_alias": "other"
+"tag": "mmlu_other_tasks"
 "include": "_default_template_yaml"
 "task": "mmlu_nutrition"
 "task_alias": "nutrition"
--- a/lm_eval/tasks/mmlu/default/mmlu_philosophy.yaml
+++ b/lm_eval/tasks/mmlu/default/mmlu_philosophy.yaml
 "dataset_name": "philosophy"
 "description": "The following are multiple choice questions (with answers) about philosophy.\n\
  \n"
-"group": "mmlu_humanities"
-"group_alias": "humanities"
+"tag": "mmlu_humanities_tasks"
 "include": "_default_template_yaml"
 "task": "mmlu_philosophy"
 "task_alias": "philosophy"
--- a/lm_eval/tasks/mmlu/default/mmlu_prehistory.yaml
+++ b/lm_eval/tasks/mmlu/default/mmlu_prehistory.yaml
 "dataset_name": "prehistory"
 "description": "The following are multiple choice questions (with answers) about prehistory.\n\
  \n"
-"group": "mmlu_humanities"
-"group_alias": "humanities"
+"tag": "mmlu_humanities_tasks"
 "include": "_default_template_yaml"
 "task": "mmlu_prehistory"
 "task_alias": "prehistory"
--- a/lm_eval/tasks/mmlu/default/mmlu_professional_accounting.yaml
+++ b/lm_eval/tasks/mmlu/default/mmlu_professional_accounting.yaml
 "dataset_name": "professional_accounting"
 "description": "The following are multiple choice questions (with answers) about professional\
  \ accounting.\n\n"
-"group": "mmlu_other"
-"group_alias": "other"
+"tag": "mmlu_other_tasks"
 "include": "_default_template_yaml"
 "task": "mmlu_professional_accounting"
 "task_alias": "professional_accounting"
--- a/lm_eval/tasks/mmlu/default/mmlu_professional_law.yaml
+++ b/lm_eval/tasks/mmlu/default/mmlu_professional_law.yaml
 "dataset_name": "professional_law"
 "description": "The following are multiple choice questions (with answers) about professional\
  \ law.\n\n"
-"group": "mmlu_humanities"
-"group_alias": "humanities"
+"tag": "mmlu_humanities_tasks"
 "include": "_default_template_yaml"
 "task": "mmlu_professional_law"
 "task_alias": "professional_law"
--- a/lm_eval/tasks/mmlu/default/mmlu_professional_medicine.yaml
+++ b/lm_eval/tasks/mmlu/default/mmlu_professional_medicine.yaml
 "dataset_name": "professional_medicine"
 "description": "The following are multiple choice questions (with answers) about professional\
  \ medicine.\n\n"
-"group": "mmlu_other"
-"group_alias": "other"
+"tag": "mmlu_other_tasks"
 "include": "_default_template_yaml"
 "task": "mmlu_professional_medicine"
 "task_alias": "professional_medicine"
--- a/lm_eval/tasks/mmlu/default/mmlu_professional_psychology.yaml
+++ b/lm_eval/tasks/mmlu/default/mmlu_professional_psychology.yaml
 "dataset_name": "professional_psychology"
 "description": "The following are multiple choice questions (with answers) about professional\
  \ psychology.\n\n"
-"group": "mmlu_social_sciences"
-"group_alias": "social_sciences"
+"tag": "mmlu_social_sciences_tasks"
 "include": "_default_template_yaml"
 "task": "mmlu_professional_psychology"
 "task_alias": "professional_psychology"
--- a/lm_eval/tasks/mmlu/default/mmlu_public_relations.yaml
+++ b/lm_eval/tasks/mmlu/default/mmlu_public_relations.yaml
 "dataset_name": "public_relations"
 "description": "The following are multiple choice questions (with answers) about public\
  \ relations.\n\n"
-"group": "mmlu_social_sciences"
-"group_alias": "social_sciences"
+"tag": "mmlu_social_sciences_tasks"
 "include": "_default_template_yaml"
 "task": "mmlu_public_relations"
 "task_alias": "public_relations"
--- a/lm_eval/tasks/mmlu/default/mmlu_security_studies.yaml
+++ b/lm_eval/tasks/mmlu/default/mmlu_security_studies.yaml
 "dataset_name": "security_studies"
 "description": "The following are multiple choice questions (with answers) about security\
  \ studies.\n\n"
-"group": "mmlu_social_sciences"
-"group_alias": "social_sciences"
+"tag": "mmlu_social_sciences_tasks"
 "include": "_default_template_yaml"
 "task": "mmlu_security_studies"
 "task_alias": "security_studies"
--- a/lm_eval/tasks/mmlu/default/mmlu_sociology.yaml
+++ b/lm_eval/tasks/mmlu/default/mmlu_sociology.yaml
 "dataset_name": "sociology"
 "description": "The following are multiple choice questions (with answers) about sociology.\n\
  \n"
-"group": "mmlu_social_sciences"
-"group_alias": "social_sciences"
+"tag": "mmlu_social_sciences_tasks"
 "include": "_default_template_yaml"
 "task": "mmlu_sociology"
 "task_alias": "sociology"
--- a/lm_eval/tasks/mmlu/default/mmlu_us_foreign_policy.yaml
+++ b/lm_eval/tasks/mmlu/default/mmlu_us_foreign_policy.yaml
 "dataset_name": "us_foreign_policy"
 "description": "The following are multiple choice questions (with answers) about us\
  \ foreign policy.\n\n"
-"group": "mmlu_social_sciences"
-"group_alias": "social_sciences"
+"tag": "mmlu_social_sciences_tasks"
 "include": "_default_template_yaml"
 "task": "mmlu_us_foreign_policy"
 "task_alias": "us_foreign_policy"
--- a/lm_eval/tasks/mmlu/default/mmlu_virology.yaml
+++ b/lm_eval/tasks/mmlu/default/mmlu_virology.yaml
 "dataset_name": "virology"
 "description": "The following are multiple choice questions (with answers) about virology.\n\
  \n"
-"group": "mmlu_other"
-"group_alias": "other"
+"tag": "mmlu_other_tasks"
 "include": "_default_template_yaml"
 "task": "mmlu_virology"
 "task_alias": "virology"