遇见数据集

Data accountability databases from collecting Antarctic Treaty Consultative Meeting documents

收藏
Zenodo2026-06-11 更新2026-05-26 收录
官方服务:

资源简介:

We present a consolidated dataset of Antarctic Treaty Consultative Meeting documents spanning 1961 to 2024 (ATCM I to ATCM XLVI). Our dataset drew primarily from the Antarctic Treaty Secretariat Database (ATSD, https://www.ats.aq/), with gaps in the ATSD's Working Paper collection filled from the Antarctic Treaty Documents Database (ATADD https://www.utas.edu.au/library-resources/atadd). We extracted machine-readable text from all documents using a combination of direct text extraction and OCR using a multimodal large language model for scanned documents. At each stage of our process, we documented the decisions we made and stored intermediate results to enhance reproducibility and provide data accountability. All HTTP responses from both databases were stored in a persistent cache, which will also allow future users to detect when the source datasets change; and the full OCR pipeline was also stored.

本研究构建了一套覆盖1961年至2024年的《南极条约》协商会议(Antarctic Treaty Consultative Meeting,ATCM)文件整合数据集,涵盖第1次至第46次协商会议(ATCM I至ATCM XLVI)的全部相关文件。 本数据集主要源自《南极条约》秘书处数据库(Antarctic Treaty Secretariat Database,ATSD,https://www.ats.aq/),针对该数据库工作文件馆藏存在的内容缺口,我们通过《南极条约文件数据库》(Antarctic Treaty Documents Database,ATADD,https://www.utas.edu.au/library-resources/atadd)补充了缺失内容。 针对所有文件,我们结合直接文本提取与针对扫描文件采用多模态大语言模型完成光学字符识别(Optical Character Recognition,OCR)两种技术路径,提取得到可机器读取的文本内容。 在数据处理的各个阶段,我们均对所做出的决策进行了详细记录,并留存了全部中间结果,以提升研究的可重复性,同时保障数据的可追溯性。我们将两个数据库的所有HTTP响应结果存储于持久化缓存中,该缓存不仅可留存历史数据,还可帮助后续使用者识别源数据集的变更情况;完整的OCR处理流水线也一并留存。

提供机构:
Zenodo
创建时间:
2026-05-21
二维码
社区交流群
二维码
科研交流群
商业服务