遇见数据集

AhmedZaky1/authorship-style-transfer-multilangual

收藏
Hugging Face2026-04-06 更新2026-04-12 收录
官方服务:

资源简介:

--- language: - en - zh license: other license_name: mixed-sources-unspecified license_details: "Several source styles/languages; confirm rights for your use case (see Limitations)." task_categories: - text-generation tags: - style-transfer - parallel-corpus - instruction-tuning - llm pretty_name: Parallel neutral vs author-style text size_categories: - "1K<n<10K" --- # Parallel neutral / author-style fine-tuning dataset Tabular parallel text built from matched **neutral** (“standard”) and **author-style** sources. Each row is one **chunk** of several consecutive non-empty lines, paired so that the same semantic content appears in both columns. ## Dataset statistics | | | |--|--| | **Samples (CSV rows)** | 4,868 | | **Hub size bucket** | `1K<n<10K` (matches sample count) | | **Primary file** | `fine_tune_dataset.csv` (UTF-8) | The metadata field `size_categories` refers to **number of rows / examples**, not file size in bytes or token count. ## Data files | File | Description | |------|-------------| | `fine_tune_dataset.csv` | Main split (UTF-8 CSV) | ## Data fields | Column | Type | Description | |--------|------|-------------| | `id` | int | Row index (1-based). | | `author` | string | Source label from the original filename (stem of `.txt`, underscores → spaces). | | `text_in_msa` | string | Text in a **neutral / standardized** register (chunked; lines joined with newlines inside the cell). | | `text_in_author_style` | string | Same content in the **author- or source-specific** style (aligned chunk). | Despite the column name, `text_in_msa` is used here in the sense of **neutral / MSA-like standard wording** where applicable; some authors/languages are not Arabic. Treat the column as **neutral reference text** paired with `text_in_author_style`. ## Construction - **Alignment:** For each author file, lines are read in lockstep from the neutral and original folders (same filename, case-insensitive match). - **Cleaning:** Empty lines are skipped; each kept line is stripped of leading/trailing whitespace. - **Chunking:** Non-empty lines are grouped into fixed-size chunks (default **5** lines per chunk in the build script). The last chunk may be shorter. - **Format:** Within a cell, chunk lines are stored separated by newline characters. ## Intended use - Supervised fine-tuning or preference modeling for **style transfer**, **rewriting**, or **parallel sequence** tasks where a model learns to map neutral text ↔ stylistic text. ## Limitations - **License:** Source licensing varies; confirm rights before redistribution or commercial use. The YAML `license: other` marks a non-SPDX / mixed situation—replace with a concrete SPDX id on the Hub when you know it. - **Quality:** Neutral text may be model- or rule-generated; verify for your application. - **Language mix:** Authors include multiple languages (e.g. English, Chinese); not single-locale. ## Citation If you use this dataset, cite or link this dataset on the Hugging Face Hub and credit underlying works as appropriate. ## Contact Dataset owner: [AhmedZaky1](https://huggingface.co/AhmedZaky1) on Hugging Face.

提供机构:
AhmedZaky1
二维码
社区交流群
二维码
科研交流群
商业服务