遇见数据集
官方服务:

资源简介:

c4-200m is a collection of 185 million sentence pairs generated from the cleaned English dataset from C4. This dataset can be used in grammatical error correction (GEC) tasks. The corruption edits and scripts used to synthesize this dataset is referenced from: C4-200M Synthetic Dataset

提供机构:
OpenDataLab
创建时间:
2023-12-07
二维码
社区交流群
二维码
科研交流群
商业服务