DocIE@XLLM25 Synthetic Dataset
收藏资源简介:
DocIE@XLLM25 Synthetic Dataset 是一个用于文档级别实体和关系抽取的合成数据集,由 ScaDS.AI 和 TU Dresden 的研究团队创建。数据集包含超过 5,000 篇维基百科摘要,其中包含大约 59,000 个实体和 30,000 个关系三元组。数据集的创建过程采用了一个自动化的 LLM 驱动的合成数据生成流程,包括基于 LLM 的标注和基于规则的验证两个阶段。该数据集旨在解决在零样本或少样本设置下文档级别实体和关系抽取任务中高质量标注语料稀缺的问题,并应用于评估和促进少样本和零样本文档级别信息抽取的研究。
DocIE@XLLM25 Synthetic Dataset is a synthetic dataset for document-level entity and relation extraction, created by the research teams from ScaDS.AI and TU Dresden. It contains over 5,000 Wikipedia abstracts, with approximately 59,000 entities and 30,000 relation triples. The dataset was developed using an automated LLM-driven synthetic data generation pipeline, which consists of two stages: LLM-based annotation and rule-based validation. This dataset aims to address the scarcity of high-quality annotated corpora for document-level entity and relation extraction tasks under zero-shot or few-shot settings, and is applied to evaluate and promote research on few-shot and zero-shot document-level information extraction.




