french-civil-code-augmented
收藏资源简介:
法国民法典数据集(清理与增强版)是一个经过精心整理和增强的法国民法典版本,专门设计用于训练语言模型处理法语法律术语和推理任务。该数据集包含两个核心组成部分:一是清理后的法律条款,即从法国民法典中提取的原始条款,经过过滤移除了已废弃内容和无效链接;二是合成的问答对,由大型语言模型(Mistral/GPT-4)生成,基于特定法律条款模拟法律咨询场景。数据集采用JSONL格式,遵循ChatML模板,与Qwen、Mistral等现代架构兼容。主要数据字段为“text”,包含对话内容或法律条款的纯文本表示,结构化对话以<|im_start|>user和<|im_end|>标记开始。数据预处理包括移除仅包含URL、行政占位符或“已废除”提及而无实质内容的“不相关”条款,并为关键条款生成了展示法律如何应用于现实生活场景(如继承、财产权、义务)的合成情景。该数据集旨在扩展分词器的领域特定语料库,并微调模型以理解法语法律文本的正式和古老结构。数据集基于官方法国民法典(Legifrance),许可证为Apache-2.0,语言为法语,标签涉及法律和民法典领域。
The French Civil Code Dataset (Cleaned and Enhanced Version) is a meticulously curated and enhanced version of the French Civil Code, specifically designed for training language models to handle French legal terminology and reasoning tasks. The dataset consists of two core components: one is the cleaned legal provisions, i.e., original provisions extracted from the French Civil Code, filtered to remove obsolete content and invalid links; the other is synthetic question-answer pairs generated by large language models (Mistral/GPT-4) based on specific legal provisions to simulate legal consultation scenarios. The dataset is in JSONL format, following the ChatML template, and is compatible with modern architectures like Qwen and Mistral. The main data field is text, containing plain-text representations of dialogue content or legal provisions, with structured dialogues starting with <|im_start|>user and <|im_end|> tags. Data preprocessing includes removing irrelevant provisions that contain only URLs, administrative placeholders, or abrogated mentions without substantive content, and generating synthetic scenarios for key provisions to demonstrate how laws apply to real-life situations (such as inheritance, property rights, obligations). This dataset aims to expand the domain-specific corpus for tokenizers and fine-tune models to understand the formal and archaic structures of French legal texts. The dataset is based on the official French Civil Code (Legifrance), licensed under Apache-2.0, in the French language, with tags covering legal and civil code domains.




