SINAI/ALIA-es-cultural-heritage
收藏资源简介:
ALIA西班牙文化与遗产语料库是一个开放获取的数据资源,它汇编和组织了大规模的西班牙语文化遗产文档集合。该语料库整合了遗产清单、专业期刊、档案记录、机构出版物以及关于物质和非物质遗产的描述性资源。它旨在为数字人文、文化机构、档案管理员、历史学家、语言学家和人工智能从业者提供一个同质化且可重复使用的文本基础。其广度支持文档探索和计算工作流,如语义检索、主题发现、术语提取和语言模型适应。处理后的语料库目前包含236,314个实例和939,315,404个标记,分布在100个源数据集中。这种规模和异质性使其成为构建和评估专注于西班牙文化遗产叙事、历史论述和机构文档的NLP系统的重要资源。
The ALIA Spanish Cultural and Heritage Corpus is an open-access data resource that compiles and organizes a large-scale collection of cultural heritage documents in Spanish. It integrates heritage inventories, specialized journals, archival records, institutional publications, and descriptive resources about tangible and intangible heritage. The corpus was designed to provide a homogeneous and reusable textual base for researchers in digital humanities, cultural institutions, archivists, historians, linguists, and AI practitioners. Its breadth supports both documentary exploration and computational workflows such as semantic retrieval, topic discovery, terminology extraction, and language model adaptation. The processed version of the corpus currently includes 236,314 instances and 939,315,404 tokens, distributed across 100 source datasets. This scale and heterogeneity make it an important resource for building and evaluating NLP systems focused on cultural heritage narratives, historical discourse, and institutional documentation in Spanish.




