soundstarrain/el-lightnovels-clean
收藏资源简介:
--- language: - el pretty_name: el-lightnovels-clean license: other license_name: custom-restricted-fan-translation tags: - greek - light-novel - fan-translation - curated - restricted size_categories: - 10K<n<100K configs: - config_name: default data_files: - split: train path: lines-*.parquet --- el-lightnovels-clean is a cleaned Greek light novel dataset built from Baka-Tsuki Greek project pages and recoverable linked Greek fan-translation sources. The dataset contains 3 series, 22 chapters, and 16,232 line-level records organized from a series / volume / chapter corpus, with 979,106 characters, 160,373 words, and 818,186 tokens measured with the Qwen/Qwen3-8B tokenizer. The default Hugging Face dataset view uses one row per text line. The original hierarchical text layout under `novels/` is preserved alongside the standard `lines-*.parquet` export. The source texts were manually reviewed and additionally strict-cleaned to remove obvious wrong-language pages, placeholder pages, glossary pages, Wikipedia/archive/project pages, translator-note pages, duplicated entries, and other non-story material, while keeping in-book forewords, afterwords, author notes, and commentary when they belonged to the original volume. Copyright and translation rights remain with the original rightsholders, publishers, and/or translators. This dataset should be treated as a restricted research dataset and should not be assumed to be freely redistributable or commercially reusable.
--- language: - 希腊语(el) pretty_name: el-lightnovels-clean license: 其他 license_name: 自定义受限粉丝翻译许可(custom-restricted-fan-translation) tags: - 希腊语 - 轻小说 - 粉丝翻译 - 精选 - 受限 size_categories: - 10000 < 样本量 < 100000 configs: - config_name: 默认配置 data_files: - split: 训练集 path: lines-*.parquet --- el-lightnovels-clean 是一款经过标准化清洗的希腊语轻小说数据集,其数据源自巴卡月译站(Baka-Tsuki)希腊语项目页面,以及可追溯的关联希腊语粉丝翻译资源。 该数据集涵盖3部系列作品、22个章节,总计16,232条行级记录,数据按照「系列-卷-章节」的层级语料结构进行组织;经Qwen/Qwen3-8B分词器统计,数据集共包含979,106个字符、160,373个单词,以及818,186个Token。 默认的Hugging Face数据集视图采用「每行对应一条文本行」的展示形式。数据集同时保留了`novels/`目录下的原始层级文本结构,以及标准格式的`lines-*.parquet`导出文件。 源文本均经过人工审核与严格清洗:移除了明显的语言错误页面、占位页面、术语表页面、维基百科/存档/项目页面、译者注释页面、重复条目及其他非故事类内容;同时保留了属于原卷的卷首语、卷尾语、作者注释与评论内容。 本数据集的版权与翻译权归原权利人、出版方及/或译者所有。该数据集仅可用于受限的研究用途,不得被视为可自由分发或商业复用的资源。



