遇见数据集

St Andrews Corpus

收藏
arXiv2025-09-30 收录
官方服务:

资源简介:

该数据集包含了从圣安德鲁斯语料库中提取的891个平行数据样本,这些样本是利用基于规则的分割器进行分割的。其规模为891个示例,任务是对这些数据进行翻译。

This dataset includes 891 parallel data samples extracted from the St Andrews Corpus, which were segmented with a rule-based segmenter. It has a total of 891 instances, and the task is to translate these data.

提供机构:
St Andrews University
二维码
社区交流群
二维码
科研交流群
商业服务