ipfs_azerbaijan_laws
收藏资源简介:
该数据集是阿塞拜疆官方立法文本的研究快照,从 e-qanun.az 门户网站收集而来,旨在为法律文本检索、信息提取等自然语言处理任务提供研究用途的数据支持。数据集包含两个子集:laws 和 articles。laws 子集包含 2667 条法律/文书,每条记录包含唯一标识符、标题、全文、日期、状态、其他标识符、条款计数和来源 URL 等字段。articles 子集包含 2460 条条款/章节,每条记录包含对应法律标识符、条款编号、标题、文本和来源 URL。数据语言为阿塞拜疆语,管辖区域为阿塞拜疆。请注意,该数据集并非官方整合版本,不构成法律建议,官方公报或政府门户网站为权威来源。数据集许可证为 az-eqanun,适用于非商业研究用途。
This dataset is a research snapshot of official legislative texts of Azerbaijan, collected from the e-qanun.az portal, aimed at providing data support for research purposes in natural language processing tasks such as legal text retrieval and information extraction. It contains two subsets: laws and articles. The laws subset includes 2,667 laws/documents, each with fields such as unique identifier, title, full text, date, status, other identifiers, article count, and source URL. The articles subset includes 2,460 articles/sections, each with fields such as corresponding law identifier, article number, title, text, and source URL. The data language is Azerbaijani, and the jurisdiction is Azerbaijan. Please note that this dataset is not an official consolidated version, does not constitute legal advice, and the official gazette or government portal is the authoritative source. The dataset is licensed under az-eqanun, suitable for non-commercial research purposes.
数据集概述
阿塞拜疆法律数据集是一个用于文本检索任务的法律文本数据集,语言为阿塞拜疆语(az),涵盖阿塞拜疆司法管辖区的官方立法文本。
基本信息
| 项目 | 详情 |
|---|---|
| 数据集名称 | Laws of Azerbaijan |
| 快照日期 | 2026-09-07 |
| 数据规模 | 1K-10K 条记录 |
| 语言 | 阿塞拜疆语(az) |
| 许可证 | other(自定义许可证:az-eqanun) |
| 任务类别 | 文本检索(text-retrieval) |
| 标签 | 法律、阿塞拜疆、官方文本 |
数据规模与覆盖范围
- 法律/文书数量: 3027 件
- 条款数量: 2801 条
- 覆盖范围: 基于目录支持,属于不完整覆盖
- 来源: e-qanun.az 网站的 downloadDetailPdf 功能
数据内容
数据集包含两个配置(config):
-
laws 配置(默认配置,文件:
data/laws.parquet)- 每行对应一件法律文书
- 包含字段:ID、标题、全文、日期、状态、标识符、条款数量、来源URL
-
articles 配置(文件:
data/articles.parquet)- 每行对应一个条款或章节
- 包含字段:法律ID(law_id)、条款编号、标题、文本、来源URL
使用说明与限制
- 非官方整合版本: 该数据集并非官方整合文本,官方公报/政府门户网站优先于本语料库
- 不构成法律建议: 数据集的打包、元数据和采集脚本仅供研究使用
- 采集辅助文件: 包含采集脚本(如
scrapers/collect_az.py、scrapers/common.py、scrapers/world_lib.py、scrapers/archive_fallbacks.py)用于构建快照及获取存档回退(Wayback / Common Crawl CDX)
许可与转载
官方政府/公报文本的使用需遵循来源门户网站的公共部门条款;该数据集的打包、元数据和采集脚本仅供研究使用。




