遇见数据集

mimir-lcm/fineweb-edu-350BT-sentence-split

收藏
Hugging Face2026-05-26 更新2026-05-31 收录
官方服务:

资源简介:

该数据集是Fineweb-edu 350BT子集的句子分割版本,专门用于自然语言处理任务。它基于教育内容,使用wtpsplit库的sat3-l模型将原始文本分割成句子,设定句子阈值为0.02和最大句子长度为256字符,以优化句子边界检测。数据集旨在提供高质量、结构化的教育文本句子,适用于语言建模、文本分析等应用。

This dataset is a sentence-split version of the Fineweb-edu 350BT subset, designed for natural language processing tasks. It is based on educational content, using the sat3-l model from the wtpsplit library to split raw text into sentences, with a sentence threshold of 0.02 and a maximum sentence length of 256 characters to optimize sentence boundary detection. The dataset aims to provide high-quality, structured educational text sentences suitable for applications such as language modeling and text analysis.

提供机构:
mimir-lcm
二维码
社区交流群
二维码
科研交流群
商业服务