遇见数据集

soundstarrain/ro-lightnovels-clean

收藏
Hugging Face2026-03-22 更新2026-03-29 收录
官方服务:

资源简介:

--- language: - ro pretty_name: ro-lightnovels-clean license: other license_name: custom-restricted-fan-translation tags: - romanian - light-novel - fan-translation - curated - restricted size_categories: - 10K<n<100K configs: - config_name: default data_files: - split: train path: lines-*.parquet --- ro-lightnovels-clean is a cleaned Romanian light novel dataset built from Baka-Tsuki Romanian project pages and recoverable linked Romanian fan-translation sources. The dataset contains 3 series, 111 chapters, and 57,035 line-level records organized from a series / volume / chapter corpus, with 2,729,563 characters, 466,523 words, and 924,109 tokens measured with the Qwen/Qwen3-8B tokenizer. The default Hugging Face dataset view uses one row per text line. The original hierarchical text layout under `novels/` is preserved alongside the standard `lines-*.parquet` export. The source texts were manually reviewed and additionally strict-cleaned to remove obvious wrong-language pages, placeholder pages, glossary pages, Wikipedia/archive/project pages, translator-note pages, duplicated entries, and other non-story material, while keeping in-book forewords, afterwords, author notes, and commentary when they belonged to the original volume. Copyright and translation rights remain with the original rightsholders, publishers, and/or translators. This dataset should be treated as a restricted research dataset and should not be assumed to be freely redistributable or commercially reusable.

--- 语言: - 罗马尼亚语(ro) 数据集名称:ro-lightnovels-clean 授权协议类型:其他 授权协议名称:自定义受限粉丝翻译协议(custom-restricted-fan-translation) 标签: - 罗马尼亚语 - 轻小说 - 粉丝翻译 - 精选整理 - 受限使用 数据规模分类: - 10K<n<100K 配置项: - 配置名称:default 数据文件: - 拆分方式:训练集(train) 路径:lines-*.parquet --- ro-lightnovels-clean是一款经过严格清洗的罗马尼亚语轻小说数据集,其数据源自Baka-Tsuki罗马尼亚语项目页面及可恢复的关联罗马尼亚语粉丝翻译源。 该数据集包含3部作品、111个章节,共计57035条行级文本记录,以作品/卷/章节的层级语料结构组织,总字符数达2729563,单词数466523,使用Qwen/Qwen3-8B分词器(tokenizer)统计得到的Token数为924109。 默认的Hugging Face数据集视图采用每行对应一条文本行的展示形式。标准`lines-*.parquet`导出文件同步保留了`novels/`目录下的原始层级文本布局。 所有源文本均经过人工审核与严格清洗:移除了明显的语言错误页面、占位页面、术语表页面、维基百科/存档/项目页面、译者注页面、重复条目及其他非故事类内容,但保留了原卷册附带的卷首语、卷尾语、作者注与评论内容(若其属于原卷册的组成部分)。 本数据集的版权与翻译权归属于原版权方、出版方及/或译者。该数据集仅可用于受限的研究用途,不得被视为可自由分发或用于商业复用。

提供机构:
soundstarrain
二维码
社区交流群
二维码
科研交流群
商业服务