soundstarrain/cs-lightnovels-clean
收藏资源简介:
--- language: - cs pretty_name: cs-lightnovels-clean license: other license_name: custom-restricted-fan-translation tags: - czech - light-novel - fan-translation - curated - restricted size_categories: - 1K<n<10K configs: - config_name: default data_files: - split: train path: lines-*.parquet --- cs-lightnovels-clean is a cleaned Czech light novel dataset built from Baka-Tsuki Czech project pages and recoverable linked Czech fan-translation sources. The dataset contains 2 series, 16 chapters, and 8,512 line-level records organized from a series / volume / chapter corpus, with 296,672 characters, 48,842 words, and 134,703 tokens measured with the Qwen/Qwen3-8B tokenizer. The default Hugging Face dataset view uses one row per text line. The original hierarchical text layout under `novels/` is preserved alongside the standard `lines-*.parquet` export. The source texts were manually reviewed and additionally strict-cleaned to remove obvious wrong-language pages, placeholder pages, glossary pages, Wikipedia/archive/project pages, translator-note pages, duplicated entries, and other non-story material, while keeping in-book forewords, afterwords, author notes, and commentary when they belonged to the original volume. Copyright and translation rights remain with the original rightsholders, publishers, and/or translators. This dataset should be treated as a restricted research dataset and should not be assumed to be freely redistributable or commercially reusable.
--- 语言: - cs 展示名称:cs-lightnovels-clean 许可协议:其他 许可名称:custom-restricted-fan-translation 标签: - 捷克语 - 轻小说 - 粉丝翻译 - 精选整理 - 受限 样本规模类别: - 1K<n<10K 配置项: - 配置名称:default 数据文件: - 拆分方式:训练集 路径:lines-*.parquet --- cs-lightnovels-clean是一款经过严格清洗的捷克语轻小说数据集,其数据来源于Baka-Tsuki捷克语项目页面以及可恢复的关联捷克语粉丝翻译资源。 该数据集包含2部系列作品、16个章节,共计8512条行级记录,基于系列/卷/章节的语料库进行组织;经Qwen/Qwen3-8B分词器(tokenizer)统计,数据集共包含296672个字符、48842个单词以及134703个Token。 默认的Hugging Face数据集视图采用每行对应一条文本行的格式。同时,`novels/`目录下的原始层级文本结构将与标准的`lines-*.parquet`导出文件一并保留。 源文本均经过人工审核与严格清洗,移除了明显的语言错误页面、占位页面、词汇表页面、维基百科/归档/项目页面、译者注释页面、重复条目以及其他非故事类内容;同时保留了属于原版卷册的书中前言、后记、作者注释与评论文字。 版权与翻译权归原权利持有人、出版方及/或译者所有。本数据集属于受限研究用数据集,不得被视为可自由分发或用于商业用途。



