Beijing-AISI/C-VARC
收藏资源简介:
--- dataset: C-VARC language: - zh - en license: cc-by-4.0 task_categories: - text-generation - multiple-choice multilinguality: monolingual size_categories: - 100K<n<1M annotations_creators: - expert-annotated - machine-generated source_datasets: - Social Chemistry 101 - Moral Integrity Corpus - Flames pretty_name: Chinese Value Rule Corpus (C-VARC) tags: - chinese-values - ethics - moral-dilemmas - llm-alignment - cultural-alignment configs: - config_name: default data_files: - split: c_varc_zh path: C-VARC.jsonl - split: c_varc_en path: C-VARC(EN).jsonl --- This repository contains all the data associated with the paper "**C-VARC: A Large-Scale Chinese Value Rule Corpus for Value Alignment of Large Language Models**".  We propose a three-tier value classification framework based on core Chinese values, which includes three dimensions, twelve core values, and fifty derived values. With the assistance of large language models and manual verification, we constructed a large-scale, refined, and high-quality value corpus containing over 250,000 rules. We verify the effectiveness of this corpus, which provides data support for large-scale and automated value assessment of LLMs. Main contributions: - **Construction of the first large-scale, refined Chinese Value Rule Corpus (C-VARC):** Based on the core socialist values, we developed a localized value classification framework covering national, societal, and personal levels, with 12 core values and 50 derived values. Using this framework, we built the first large-scale Chinese Value Rule Corpus (C-VARC), comprising over 250,000 high-quality, manually annotated normative rules, filling an important gap in the field. - **Systematic validation of C-VARC's generation guidance advantages and cross-model applicability:** We validated C-VARC's effectiveness in guiding scenario generation for the 12 core values. Quantitative analysis shows that C-VARC guided scenes exhibit more compact clustering and clearer boundaries in *t*-SNE space. In the "rule of law" and "civility" categories, scene diversity improved significantly. In tests on six ethical themes, seven major LLMs chose C-VARC generated options over 70% of the time, and the consistency with five Chinese annotators exceeded 0.87, confirming C-VARC's strong guidance capability and its clear representation of Chinese values. - **Proposal of a rule-driven method for large-scale moral dilemma generation:** Leveraging C-VARC, we propose a method to automatically generate moral dilemmas (MDS) based on value priorities. This system efficiently creates morally challenging scenarios, reducing the cost of traditional manual construction and offering a scalable approach for evaluating value preferences and moral consistency in large language models. **paper**: You can access the paper at this [link](https://arxiv.org/abs/2506.01495). **github**: You can access all the code from the paper at this [link](https://github.com/Beijing-AISI/C-VARC). **English version**: We are currently working on translating C-VARC into English. The file C-VARC(EN).jsonl provides 10,000 rules that have already been translated into English. The translation was conducted using DeepL Translator, which is widely recognized in academic contexts for its high accuracy and low ambiguity in technical and scholarly text translation.
数据集:C-VARC 支持语言:中文、英文 许可证:CC BY 4.0(知识共享署名4.0国际许可协议) 任务类别:文本生成、多项选择 多语言属性:单语言 样本规模:10万<样本量<100万 标注创建者:专家标注、机器生成 源数据集:Social Chemistry 101、道德诚信语料库(Moral Integrity Corpus)、Flames 官方名称:中文价值规则语料库(C-VARC) 标签:中文价值观、伦理道德、道德困境、大语言模型对齐(LLM-alignment)、文化对齐 配置项: - 默认配置: 数据文件: - 拆分集:c_varc_zh,文件路径:C-VARC.jsonl - 拆分集:c_varc_en,文件路径:C-VARC(EN).jsonl 本仓库包含与论文《C-VARC:面向大语言模型(Large Language Model)价值对齐的大规模中文价值规则语料库》相关的全部数据。  我们提出了一套基于中华核心价值观的三层价值分类框架,涵盖3大维度、12项核心价值与50项衍生价值。依托大语言模型辅助与人工核验,我们构建了规模超25万条规则的大规模、精细化高质量价值语料库。经实验验证,该语料库可为大语言模型的大规模自动化价值评估提供数据支撑。 主要贡献: 1. **首个大规模精细化中文价值规则语料库(C-VARC)的构建**:基于社会主义核心价值观,我们搭建了覆盖国家、社会与个人三个层面的本土化价值分类框架,包含12项核心价值与50项衍生价值。依托该框架,我们构建了首个大规模中文价值规则语料库(C-VARC),内含超25万条高质量人工标注的规范性规则,填补了该领域的重要空白。 2. **C-VARC生成引导优势与跨模型适用性的系统性验证**:我们验证了C-VARC在引导12项核心价值的场景生成方面的有效性。定量分析表明,经C-VARC引导生成的场景在t-SNE(t分布随机邻域嵌入)空间中聚类更紧凑、边界更清晰;在“法治”与“文明”类别中,场景多样性得到显著提升。在6项伦理主题的测试中,7款主流大语言模型选择C-VARC生成选项的比例超过70%,且与5位中文标注者的一致性评分超过0.87,证实了C-VARC强劲的引导能力与清晰的中文价值观表达能力。 3. **提出规则驱动的大规模道德困境生成方法**:依托C-VARC,我们提出了一种基于价值优先级的道德困境(Moral Dilemmas, MDS)自动生成方法。该系统可高效生成具有道德挑战性的场景,降低了传统人工构建的成本,并为评估大语言模型的价值偏好与道德一致性提供了可扩展的技术路径。 **论文**:可通过此[链接](https://arxiv.org/abs/2506.01495)获取论文全文。 **GitHub仓库**:可通过此[链接](https://github.com/Beijing-AISI/C-VARC)获取论文配套的全部代码。 **英文版本**:我们目前正在推进C-VARC的英文译制工作。文件C-VARC(EN).jsonl包含已完成英译的1万条规则。本次翻译采用DeepL翻译器,其在学术场景下的技术与学术文本翻译准确率高、歧义性低,已获得学术界广泛认可。




