INSAIT-Institute/rit-hellaswag
收藏资源简介:
HellaSwag Multilingual 是一个多语言常识句子补全基准测试数据集,专注于日常情境中的合理延续。该数据集通过自动翻译框架(RiTranslation)从原始HellaSwag基准测试翻译而来,支持多种东欧和南欧语言,包括乌克兰语、希腊语、爱沙尼亚语、立陶宛语、罗马尼亚语、斯洛伐克语和土耳其语。每个语言作为一个独立的配置(config),数据集结构包含ind、activity_label、ctx_a、ctx_b、ctx、endings、source_id、split、split_type和label等字段。该数据集旨在便于跨语言比较模型性能,适用于多语言基准评估和分析。
HellaSwag Multilingual is a commonsense sentence-completion benchmark centered on plausible continuations for everyday situations. It is a multilingual version of the HellaSwag benchmark, translated through an automated framework (RiTranslation) and includes languages such as Ukrainian, Greek, Estonian, Lithuanian, Romanian, Slovak, and Turkish. Each language is provided as a separate config, with dataset fields including ind, activity_label, ctx_a, ctx_b, ctx, endings, source_id, split, split_type, and label. The dataset is intended for multilingual benchmark evaluation and analysis, facilitating cross-language model comparisons.




