遇见数据集

GPT-NL/DuidelijkeTaal-v1.0-split

收藏
Hugging Face2025-12-23 更新2025-12-20 收录
官方服务:

资源简介:

该数据集是Duidelijk Taal数据集的副本,添加了训练和测试分割,由GPT-NL的基准测试团队创建,以便于重用和记录数据分割。原始描述:语言材料“自动文本简化的人类评估:众包结果”是在“Duidelijke taal”项目期间创建的,包含电子表格CrowdsourcingResults.cvs。该电子表格包含来自SoNaR语料库的句子/文本,这些句子/文本经过GPT-4简化,并包含人类对这些简化在复杂性、准确性和流畅性方面的评分。

This dataset is a copy of the Duidelijk Taal dataset with training and test splits, created by the benchmarking team of GPT-NL to allow for re-use with documented data splits. Original description: The language material "Human evaluation of automated text simplification: crowdsourcing results" was created during the project "Duidelijke taal" and consists of the spreadsheet CrowdsourcingResults.cvs. That spreadsheet contains sentences/text from the SoNaR corpus, a version of those sentences / that text simplified by GPT-4 and the human ratings of those simplifications with respect to complexity, accuracy and fluency.

提供机构:
GPT-NL
二维码
社区交流群
二维码
科研交流群
商业服务