PreCo
收藏资源简介:
PreCo是由依图科技的幼儿实验室创建的大规模英语数据集,专注于解决共指消解问题。该数据集包含约38,000份文档和1240万词,主要来自英语为母语的学龄前儿童的词汇。PreCo通过提高训练与测试集之间的重叠,解决了训练与测试集之间低重叠的挑战,并首次量化了提及检测器对共指消解性能的影响。数据集的应用领域包括阅读理解、翻译和文本摘要等,旨在提高共指消解算法的效率和准确性。
PreCo is a large-scale English dataset created by the Early Childhood Lab of Yitu Technology, focusing on addressing the coreference resolution task. This dataset contains approximately 38,000 documents and 12.4 million words, mainly sourced from the vocabulary of native English-speaking preschool children. PreCo solves the challenge of low overlap between training and test sets by increasing the overlap between them, and for the first time quantifies the impact of mention detectors on the performance of coreference resolution models. The application fields of this dataset include reading comprehension, translation, text summarization and other areas, aiming to improve the efficiency and accuracy of coreference resolution algorithms.




