遇见数据集

MLHermit/9ja-bookcorpus

收藏
Hugging Face2026-05-15 更新2026-05-31 收录
官方服务:

资源简介:

该数据集包含从160本尼日利亚作者书籍的PDF文档中提取的文本,经过清理和精炼处理。处理步骤包括:基本完整性检查(如主要为字母字符、合理的单词长度)、语言检测以筛选英文文本,以及使用NLTK英文单词库进行拼写纠正。数据集适用于语言建模、文本生成和其他自然语言处理任务。数据集结构为单一“训练”分割,每个示例包含一个“文本”特征,代表从PDF提取的单行原始文本。

A corpus of text data scraped from 160 PDFs from Nigerian Authors, cleaned and refined. It has undergone several refinement steps including: basic sanity checks (e.g., mostly alphabetic characters, reasonable word lengths), language detection to filter for English text, and spell correction using NLTKs English word corpus. This dataset is intended for use in language modeling, text generation, and other NLP tasks. The dataset consists of a single train split with a text feature per example representing a single line of raw text from the PDFs.

提供机构:
MLHermit
二维码
社区交流群
二维码
科研交流群
商业服务