遇见数据集

ClassiCC-Corpus/curio-rewrite-edu-dataset

收藏
Hugging Face2026-05-11 更新2026-05-31 收录
官方服务:

资源简介:

该数据集名为Curio Rewrite — Educational,包含葡萄牙语网页文本(从ClassiCC数据集中筛选出的教育子集),这些文本由Qwen2.5-7B-Instruct模型根据四种提示风格进行重写,用于训练Curio重写模型。数据集分为四个配置:easy(使用简单词汇、适合儿童的改写)、medium(中等程度的改写)、hard(复杂改写)和qa(重新格式化为问答形式),每个配置包含7,777,128行数据,源文档相同但重写风格不同。字段包括原始文本、文档ID、元数据、提示、提示类型、聚类ID、类别标签和重写文本。

Portuguese web texts (educational subset, filtered from ClassiCC) rewritten by Qwen2.5-7B-Instruct under four prompt styles. Used to train the Curio rewrite models. Each config holds the same source documents with a different rewrite style: easy (simple vocabulary, child-friendly paraphrase), medium (moderate paraphrase), hard (sophisticated paraphrase), and qa (reformatted as question/answer), with 7,777,128 rows each.

提供机构:
ClassiCC-Corpus
二维码
社区交流群
二维码
科研交流群
商业服务