遇见数据集

mimir-lcm/fineweb-2-sentence-split

收藏
Hugging Face2026-05-26 更新2026-05-31 收录
官方服务:

资源简介:

该数据集是Fineweb 2的分句版本,将文本分割成句子。我们使用wtpsplit库的sat3-l模型进行句子分割,设定句子阈值为0.02,最大句子长度为256。实例按语言采样,以平衡与Fineweb-edu相关的数据。

This dataset is a sentence-split version of FineWeb 2, where text is divided into sentences. We used the sat3-l model from the wtpsplit library for sentence splitting, with a sentence threshold of 0.02 and a maximum sentence length of 256. Instances per language were sampled to balance the data with respect to Fineweb-edu.

提供机构:
mimir-lcm
二维码
社区交流群
二维码
科研交流群
商业服务