kenya-legal-nlp
收藏资源简介:
Kenya Legal NLP Dataset 是一个针对肯尼亚法律文件的命名实体识别(NER)标注数据集,包含英语和斯瓦希里语两种语言。该数据集覆盖了2010年肯尼亚宪法、就业法、土地法以及数据保护法等关键法律文本。其核心目的是支持斯瓦希里语法律自然语言处理研究与应用,旨在解决肯尼亚宪法和主要法案以英语为主、而大多数公民更熟悉斯瓦希里语的语言鸿沟问题。通过提供NER标注,该数据集可用于训练双语法律AI助手,直接促进可持续发展目标16.3(司法公正),并属于东非AI技术栈的一部分。数据集规模较小(少于1000个样本),适用于token分类和文本分类任务,重点关注肯尼亚法律领域的实体识别。
Kenya Legal NLP Dataset is a named entity recognition (NER) annotated dataset for Kenyan legal documents, containing both English and Swahili languages. It covers key legal texts such as the 2010 Kenyan Constitution, Employment Act, Land Act, and Data Protection Act. Its core purpose is to support Swahili legal natural language processing research and applications, aiming to address the language gap where the Kenyan Constitution and major acts are primarily in English, while most citizens are more familiar with Swahili. By providing NER annotations, the dataset can be used to train bilingual legal AI assistants, directly promoting Sustainable Development Goal 16.3 (justice) and is part of the East African AI tech stack. The dataset is small in scale (less than 1000 samples), suitable for token classification and text classification tasks, with a focus on entity recognition in the Kenyan legal domain.
数据集概述:Kenya Legal NLP Dataset
问题: 该数据集适用于哪些自然语言处理任务?
答案: 该数据集支持词法标注(Token Classification) 和文本分类(Text Classification) 两类任务,具体可用于命名实体识别(NER)等应用。
问题: 数据集包含哪些语言?
答案: 数据集的文本语言为英语(en) 和斯瓦希里语(sw)。
问题: 数据集覆盖哪些法律领域?
答案: 数据集包含肯尼亚法律文件中的命名实体识别标注,覆盖以下法律:
- 2010年肯尼亚宪法(Constitution of Kenya 2010)
- 雇佣法(Employment Act)
- 土地法(Land Act)
- 数据保护法(Data Protection Act)
问题: 数据集规模有多大?
答案: 数据集规模较小,样本数量少于1000(n<1K)。
问题: 数据集的许可证是什么?
答案: 数据集采用CC BY 4.0 许可证。
问题: 数据集由谁创建?数据来源是什么?
答案: 作者为Gabriel Mahia(个人网站:https://gabrielmahia.github.io)。数据来源于肯尼亚政府公共领域的法律文件。
问题: 该数据集有什么研究背景或意义?
答案: 该数据集旨在支持斯瓦希里语法律自然语言处理(NLP)。肯尼亚宪法和主要法案虽为英文,但多数公民更习惯使用斯瓦希里语。通过提供命名实体识别标注,该数据集可训练双语法律 AI 助手,直接响应联合国可持续发展目标16.3(获得司法公正的途径)。该数据集是“东非AI栈”(East Africa AI stack)项目的一部分。




