遇见数据集

Monsoon Project - Database Structure and Datasets

收藏
Zenodo2025-04-05 更新2026-05-26 收录
官方服务:

资源简介:

This dataset presents annotated historical texts from the Monsoon project. https://www.monsoon.uevora.pt/ It includes: - Full transcriptions of the texts in structured JSON format- Named entity annotations (people, places, organizations) in CSV format. For more information about the categories of the entities consult: https://www.monsoon.uevora.pt/annotatedentities#! - Event annotations with token positions and associated context- Geographic entity data, including coordinates and polygons- A PostgreSQL schema file for database reconstruction The dataset was prepared using Django and PostgreSQL, and supports applications in natural language processing, digital humanities, and historical research. All content is provided under a Creative Commons Attribution 4.0 International (CC BY 4.0) license.

本数据集收录了Monsoon项目(Monsoon project)的标注历史文本,项目官方网址为https://www.monsoon.uevora.pt/。 数据集包含以下内容: - 采用结构化JSON格式存储的完整文本转录数据; - CSV格式的命名实体标注信息,涵盖人物、地点、组织机构三类实体,实体类别详情可查阅网址:https://www.monsoon.uevora.pt/annotatedentities#!; - 带有Token位置及关联上下文的事件标注数据; - 包含坐标与多边形信息的地理实体数据集; - 用于数据库重建的PostgreSQL架构文件。 本数据集基于Django与PostgreSQL构建,可支持自然语言处理、数字人文及历史研究等领域的应用。 所有内容均采用知识共享署名4.0国际(CC BY 4.0)许可协议发布。

提供机构:
Zenodo
创建时间:
2025-04-04
二维码
社区交流群
二维码
科研交流群
商业服务