C-RANS: A Benchmark Dataset of Learner Chinese with Register-Specific Naturalness Ratings and Revisions
收藏资源简介:
This dataset contains sentence-level annotations of learner Chinese for evaluating grammaticality, register-specific naturalness, and naturalness-oriented revision. The dataset is built from the Guangwai-Lancaster Chinese Learner Corpus (GLCLC) and includes 6,358 learner sentences across four communicative registers: written formal, written informal, spoken formal, and spoken informal. Each sentence is annotated by certified teachers of Teaching Chinese as a Foreign Language (TCFL). The annotations include five-point ratings for grammaticality and register-specific naturalness, expert revisions that improve naturalness while preserving meaning, and comments describing register-related issues in learner production. Dataset contents The released dataset provides: - sentence-level learner Chinese samples- original corpus metadata, including sentence ID, text ID, file ID, file path, topic, and data type- full-text context for each target sentence- register labels: written formal, written informal, spoken formal, and spoken informal- five-point grammaticality ratings- five-point register-specific naturalness ratings- expert naturalness revisions- annotator comments on register-specific issues Register categories The dataset operationalizes register along two dimensions: mode and formality. This results in four register categories: - written formal- written informal- spoken formal- spoken informal These categories were derived from the text types in GLCLC, including exam writing, free writing, oral exams, oral interviews, oral counselling, and spoken classroom exercises. Purpose C-RANS is designed to support research and evaluation in three main areas: - second language acquisition research on Chinese learner language and register competence- development of pedagogical materials for Teaching Chinese as a Foreign Language- evaluation of large language models on grammaticality assessment, register-specific naturalness assessment, and naturalness-oriented rewriting Benchmark evaluation The repository also includes scripts and baseline results for three benchmark tasks: - Task A: grammaticality scoring- Task B: register-specific naturalness scoring- Task C: naturalness-oriented rewriting The benchmark evaluation uses Quadratic Weighted Kappa (QWK) for rating tasks and multiple revision metrics for rewriting, including GLEU, BGE-based cosine similarity, Expert-Aligned Perplexity (EAP), and Normalized Compression Distance (NCD). Data structure The main dataset is provided in JSON format. Each record contains standardized fields for the target sentence, context, metadata, register category, human ratings, expert revision, and annotator comments. The repository also contains evaluation scripts and baseline model outputs used in the technical validation of the dataset. License The dataset is released under the Creative Commons Attribution 4.0 International (CC BY 4.0) license.



