CT-Eval
收藏资源简介:
CT-Eval是一个专为评估大型语言模型在中文文本到表格任务上性能而设计的数据集。该数据集由早稻田大学创建,涵盖了28个不同领域,确保了数据的多样性。数据集的构建过程中,首先从百度百科这一流行的中文多学科在线百科中收集文档-表格对,然后使用大型语言模型作为幻觉判断器,过滤掉含有幻觉的任务样本,最后通过人工标注者进一步清理验证和测试集中的幻觉信息。CT-Eval包含88.6K任务样本,旨在帮助研究人员评估和快速理解现有大型语言模型的中文文本到表格能力,并作为提升文本到表格性能的重要资源。
CT-Eval is a dedicated dataset for evaluating the performance of large language models (LLMs) on Chinese text-to-table tasks. Developed by Waseda University, this dataset covers 28 distinct domains to ensure comprehensive data diversity. In its construction workflow, document-table pairs were first collected from Baidu Baike, a widely used Chinese multi-disciplinary online encyclopedia. Subsequently, large language models were employed as hallucination detectors to filter out task samples containing hallucinatory content. Finally, human annotators further cleaned and validated the hallucination-related information in the validation and test sets. CT-Eval comprises 88.6K task samples, and aims to assist researchers in evaluating and rapidly comprehending the Chinese text-to-table capabilities of existing large language models, serving as a critical resource for enhancing text-to-table performance.




