ngocdang83-tran-vi-teacher
收藏资源简介:
# tran-vi-teacher: Chinese→Vietnamese Web-novel Teacher Dataset A **350,751-row strict-clean Chinese-to-Vietnamese parallel corpus** for web-novel translation training, distilled from Gemini 2.5/3.0/3.1 teacher models on raw Chinese web-novel paragraphs. ## Overview - **Source language**: Simplified Chinese (zh) - **Target language**: Vietnamese (vi) - **Domain**: Web-novel (xianxia, urban, school, fantasy, history, sci-fi) - **Format**: Multi-line paragraph chunks (median 7 lines/row) - **Total paired rows**: 364,025 - **Strict-clean rows**: 350,751 (96.3% pass rate) - **Dedup-source rows**: 347,259 (recommended for training) ## Quality Filter (Strict-Clean) All rows in `tran_vi_teacher_strict_clean_*.jsonl` pass these gates: | Gate | Description | |---|---| | `finish_reason == stop` | Teacher completed naturally, not truncated | | Closed target tag | `<TRANSLATION>` or `<OUTPUT_VI>` fully closed | | Line-count match | Source and target line counts equal | | No Han leak | Target Vietnamese contains zero Chinese chars | | Target ≤3000 chars | Reasonable length cap | | Char ratio ≤5.0 | Target/source ratio sanity check | Reject reasons documented in the original audit (8,563 Han-leak, 4,065 line mismatch, etc.) ## Teacher Model Mix | Tier | Rows | % | Models | |---|---:|---:|---| | Flash Lite | 299,113 | 86.1% | gemini-2.5-flash-lite, gemini-3.1-flash-lite | | Flash | 54,490 | 15.7% | gemini-2.5-flash, gemini-3-flash, gemini-3.1-flash | | Pro | 10,422 | 3.0% | gemini-2.5-pro, gemini-3-pro, gemini-3.1-pro | ## File Variants | File | Rows | Use | |---|---:|---| | `tran_vi_teacher_strict_clean_all.jsonl` | 350,751 | Includes duplicates | | `tran_vi_teacher_strict_clean_dedup_source.jsonl` | **347,259** | **Recommended.** Deduped by source hash, preference: pro > flash > flash_lite. | ## Schema Each row is a JSON object: ```json { "id": "tran-vi-11116AcGtTNs", "source": "他必须得抓紧时间了。\n凌伊山掏出手机...", "target": "Hắn phải nhanh chóng lên đường.\nLăng Y Sơn lấy điện thoại ra...", "source_zh": "他必须得抓紧时间了。...", "target_vi": "Hắn phải nhanh chóng lên đường....", "meta": { "source_dataset": "novel-data/tran-vi", "stem": "11116AcGtTNs", "model": "gemini-2.5-flash-lite-preview-09-2025", "teacher_tier": "flash_lite", "finish_reason": "stop", "source_tag": "SOURCE", "target_tag": "TRANSLATION", "source_hash": "b9d49985de...", "pair_hash": "2118585f56...", "source_line_count": 8, "target_line_count": 8, "source_chars": 201, "target_chars": 660, "target_source_char_ratio": 3.2836, "quality_bucket": "strict_clean" } } ``` Both `source`/`target` and `source_zh`/`target_vi` carry the same text — keep whichever pair matches your trainer's schema. ## Statistics (Dedup Source Variant) ### Target Length Distribution Per Tier | Tier | N | mean chars | p50 | p99 | % >500 chars | |---|---:|---:|---:|---:|---:| | pro | 9,602 | 1,067 | 959 | 1,802 | 93.02% | | flash | 51,073 | 730 | 713 | 1,583 | 88.18% | | flash_lite | 286,584 | 869 | 769 | 1,768 | 89.05% | This dataset has **dense long-target signal** (86-93% of rows have >500 char targets), making it well-suited for paragraph-level MT training and avoiding the "trained EOS bias" pathology seen with chat-extracted single-sentence SFT datasets. ### Line Count Per Row - Median: 7 lines - p90: 13 lines - p99: 19 lines - Max: 85 lines ## Intended Uses ### Recommended 1. **Broad teacher/distill pool** for Chinese→Vietnamese MT training 2. **Cross-domain coverage** for urban, school, history, fantasy, sci-fi 3. **Paragraph-level training** (avoid sentence-segmentation collapse) 4. **OOD/lexical mining** (modern terms, western names, classical patterns) ### With Caution 1. **Style polish/literary final-output** — most rows are Flash-Lite, not literary gold 2. **Mixed with human-curated gold** in pool — do not weight equally 3. **Pro subset** is highest-quality candidate for review/gold seeding ### Not Recommended 1. As primary style reference (use literary corpora for that) 2. Without sentence-resegmentation if your trainer needs short sentences 3. As-is for max_position ≤256 model (target p99 ≈ 1500-1800 chars) ## How This Was Built 1. **Source**: Raw Chinese web-novel paragraphs from various sources 2. **Teacher**: OpenAI-compatible Gemini API calls with `<SOURCE>` / `<INPUT_ZH>` wrapped prompts 3. **Extraction**: Tags parsed (`<TRANSLATION>`, `<OUTPUT_VI>`, `<VI>`) 4. **Filtering**: Strict gates applied (see Quality Filter section) 5. **Dedup**: Source-hash dedup with teacher-tier preference Full audit available in original extraction manifest. ## License CC-BY-4.0 — free for any use with attribution. Note: Gemini API output is derivative work; consult Google's Gemini API terms for downstream model training (Google permits Gemini outputs for training in most cases as of late 2025/early 2026, but verify current TOS). ## Attribution Shared by **[chi-vi](https://huggingface.co/chi-vi)** — Chinese↔Vietnamese translation research community.



