Grammarly’s Yahoo Answers Formality Corpus (GYAFC)
收藏资源简介:
Grammarly’s Yahoo Answers Formality Corpus (GYAFC)是由马里兰大学帕克分校的研究人员创建的一个大规模数据集,专注于非正式到正式的文本风格转换。该数据集包含110,000对非正式和正式的句子对,这些句子来源于Yahoo Answers平台,并通过人工重写确保了正式性。数据集的创建过程包括筛选、预处理和人工重写,旨在为机器翻译和文本简化等领域的研究提供高质量的训练和评估资源。GYAFC数据集的应用领域包括自然语言处理、机器翻译和文本生成,旨在解决文本风格转换中的自动评估和模型训练问题。
Grammarly’s Yahoo Answers Formality Corpus (GYAFC) is a large-scale dataset created by researchers at the University of Maryland, College Park, focusing on informal-to-formal text style transfer. It contains 110,000 informal-formal sentence pairs sourced from the Yahoo Answers platform, with the formal versions manually rewritten to guarantee their formality. The dataset's creation process includes filtering, preprocessing and manual rewriting, aiming to provide high-quality training and evaluation resources for research in fields such as machine translation and text simplification. The GYAFC dataset is applied across natural language processing, machine translation and text generation, targeting the resolution of challenges related to automatic evaluation and model training in text style transfer tasks.




