mimir-lcm/fineweb-edu-350BT-sentence-split
收藏官方服务:
资源简介:
该数据集是Fineweb-edu 350BT子集的句子分割版本,专门用于自然语言处理任务。它基于教育内容,使用wtpsplit库的sat3-l模型将原始文本分割成句子,设定句子阈值为0.02和最大句子长度为256字符,以优化句子边界检测。数据集旨在提供高质量、结构化的教育文本句子,适用于语言建模、文本分析等应用。
This dataset is a sentence-split version of the Fineweb-edu 350BT subset, designed for natural language processing tasks. It is based on educational content, using the sat3-l model from the wtpsplit library to split raw text into sentences, with a sentence threshold of 0.02 and a maximum sentence length of 256 characters to optimize sentence boundary detection. The dataset aims to provide high-quality, structured educational text sentences suitable for applications such as language modeling and text analysis.
提供机构:
mimir-lcm


