soundstarrain/fr-lightnovels-clean
收藏资源简介:
--- language: - fr pretty_name: fr-lightnovels-clean license: other license_name: custom-restricted-fan-translation tags: - french - light-novel - fan-translation - curated - restricted size_categories: - 100K<n<1M configs: - config_name: default data_files: - split: train path: train-*.parquet --- fr-lightnovels-clean is a cleaned French light novel dataset built from Baka-Tsuki French project pages and recoverable linked French fan-translation sources. The dataset contains 31 series, 850 chapters, and 460,266 line-level records organized from a series / volume / chapter corpus, with 25,148,311 characters, 4,255,333 words, and 7,317,880 tokens measured with the Qwen/Qwen3-8B tokenizer. The default Hugging Face dataset view uses one row per text line. The original hierarchical text layout under `novels/` is preserved alongside the standard `train-*.parquet` export. The source texts were manually reviewed and additionally strict-cleaned to remove obvious wrong-language pages, placeholder pages, glossary pages, Wikipedia/archive/project pages, translator-note pages, duplicated entries, and other non-story material, while keeping in-book forewords, afterwords, author notes, and commentary when they belonged to the original volume. Copyright and translation rights remain with the original rightsholders, publishers, and/or translators. This dataset should be treated as a restricted research dataset and should not be assumed to be freely redistributable or commercially reusable.



