遇见数据集

Evaluation Results of English / Japanese LLMs Using Swallow-Evaluation ver.202407

收藏
Zenodo2024-10-21 更新2026-05-26 收录
官方服务:

资源简介:

Evaluation Results of English / Japanese LLMs Using Swallow-Evaluation ver.202407 This dataset is the source material of our observational analysis paper, "Significance and Effectiveness of Training LLM with Japanese Texts [Saito+, 2024]." It includes the evaluation results of 35 LLMs on 19 Japanese and English tasks, using Swallow-evaluation ver.202407. As part of the Swallow project at Tokyo Institute of Technology, this dataset was developed to enable rigorous and comprehensive comparison of Japanese and English LLMs developed in Japan and worldwide. Details Tasks Evaluation experiments are conducted on LLMs using 10 datasets for Japanese language understanding and generation tasks, and 9 datasets for English language understanding and generation tasks. All evaluation scores are normalized within a range from 0 (lowest) to 1 (highest). Refer to the reference [Saito+, 2024] for the complete list of evaluation tasks and datasets, evaluation metrics, and task configurations. Environment The evaluaitons were primarily conducted on A100 nodes (AIST), using Python as the programming language. Limitation While efforts were made to evaluate under fair conditions, considering the unique specifications of each LLM (such as tokenization and system prompts), minor differences in evaluation specifics (like prompt formatting and dependence on eval. environment) may cause task evaluation scores to change independently of the LLM’s performance. Reference ```ここにNL研究会予稿書誌情報を載せる``` License This dataset is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License. Creators Swallow LLM (GitHub, Official Web Page)

提供机构:
Zenodo
创建时间:
2024-08-02
二维码
社区交流群
二维码
科研交流群
商业服务