遇见数据集

BiSECT

收藏
OpenDataLab2026-07-12 更新2024-05-09 收录
官方服务:

资源简介:

BiSECT 是一个用于句子简化的数据集,它能够将一个长而复杂的句子分成更短的句子,并根据需要重新措辞。 BiSECT 训练数据由 100 万个长英语句子和较短的、意义等效的英语句子组成。这些是通过在双语平行语料库中提取 1-2 个句子对齐,然后使用机器翻译将语料库的两侧转换为相同的语言而获得的。

BiSECT is a sentence simplification dataset that splits long and complex sentences into shorter ones and rephrases them as needed. The training data of BiSECT consists of 1 million long English sentences paired with shorter, semantically equivalent English sentences. These are obtained by extracting 1-2 sentence alignments from bilingual parallel corpora, then translating both sides of the corpus into the same language using machine translation.

提供机构:
OpenDataLab
创建时间:
2022-05-23
搜集汇总
数据集介绍
BiSECT 数据集图片
背景与挑战
背景概述
BiSECT是一个句子简化数据集,包含100万个长英语句子及其简化版本,通过双语语料库对齐和机器翻译构建。该数据集由宾夕法尼亚大学等机构于2021年发布。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务