QuoteR
收藏资源简介:
QuoteR是由清华大学人工智能研究院构建的一个大型开放引用推荐数据集,旨在帮助写作者高效找到合适的引用。该数据集包含三个部分:英语、标准中文和古典中文,总计13,550条引用,分别来自Wikiquote、Juzimi和Gushiwenwang等资源。数据集的创建过程涉及从多个高质量语料库中提取引用及其上下文,确保数据的多样性和实用性。QuoteR主要用于解决引用推荐任务中的挑战,如提高推荐准确性和效率,适用于自然语言处理和机器学习研究。
QuoteR is a large-scale open citation recommendation dataset developed by the Institute of Artificial Intelligence at Tsinghua University, aiming to assist writers in efficiently locating suitable citations. It consists of three parts: English, standard Chinese, and classical Chinese, with a total of 13,550 citations sourced from resources including Wikiquote, Juzimi, and Gushiwenwang. The development of the dataset involves extracting citations and their contextual information from multiple high-quality corpora to ensure data diversity and practicality. QuoteR is primarily intended to address challenges in citation recommendation tasks, such as enhancing recommendation accuracy and efficiency, and is applicable to research in natural language processing and machine learning.

- 1QuoteR: A Benchmark of Quote Recommendation for Writing清华大学人工智能研究院 · 2022年



