BiSECT
收藏官方服务:
资源简介:
BiSECT 是一个用于句子简化的数据集,它能够将一个长而复杂的句子分成更短的句子,并根据需要重新措辞。 BiSECT 训练数据由 100 万个长英语句子和较短的、意义等效的英语句子组成。这些是通过在双语平行语料库中提取 1-2 个句子对齐,然后使用机器翻译将语料库的两侧转换为相同的语言而获得的。
BiSECT is a sentence simplification dataset that splits long and complex sentences into shorter ones and rephrases them as needed. The training data of BiSECT consists of 1 million long English sentences paired with shorter, semantically equivalent English sentences. These are obtained by extracting 1-2 sentence alignments from bilingual parallel corpora, then translating both sides of the corpus into the same language using machine translation.
提供机构:
OpenDataLab创建时间:
2022-05-23
搜集汇总
数据集介绍

背景与挑战
背景概述
BiSECT是一个句子简化数据集,包含100万个长英语句子及其简化版本,通过双语语料库对齐和机器翻译构建。该数据集由宾夕法尼亚大学等机构于2021年发布。
以上内容由遇见数据集搜集并总结生成



