flwrlabs/fed-legal
收藏资源简介:
Fed-Legal是一个仅限英语的五库联合监督微调(SFT)数据集,用于法律语言任务。它将选定的法律基准数据集转换为共享的聊天式模式,适用于指令调整和联合学习实验。每个库对应不同的法律任务家族或数据集来源。数据集通过设计是非独立同分布的:每个`client_id`代表不同的法律任务分布,而不是一个同质语料库的随机分片。数据集构建自LexGLUE/LEDGAR(法律条款分类)、LexGLUE/CaseHOLD(多选案例持有选择)、LexGLUE/Unfair-ToS(消费者服务条款不公平性分类)、LexGLUE/SCOTUS(美国最高法院问题领域分类)和LegalBench/Contract NLI(合同自然语言推理和法律推理任务)。所有示例都标准化为带有系统提示、用户指令和助手回答的聊天消息。输出数据集包含通过连接每个库的分区创建的全局`train`、`valid`和`test`分割。
Fed-Legal is an English-only, five-silo, federated supervised fine-tuning (SFT) dataset for legal language tasks. It converts selected legal benchmark datasets into a shared chat-style schema suitable for instruction tuning and federated learning experiments. Each silo corresponds to a different legal task family or dataset source. The dataset is intentionally non-IID by construction: each `client_id` represents a different legal task distribution rather than a random shard of one homogeneous corpus. The dataset is built from LexGLUE/LEDGAR (legal provision classification), LexGLUE/CaseHOLD (multiple-choice case holding selection), LexGLUE/Unfair-ToS (consumer terms-of-service unfairness classification), LexGLUE/SCOTUS (U.S. Supreme Court issue-area classification), and LegalBench/Contract NLI (contract natural language inference and legal reasoning tasks). All examples are normalized into chat messages with a system prompt, user instruction, and assistant answer. The output dataset contains global `train`, `valid`, and `test` splits created by concatenating each clients per-silo split.




