CO.RA.PAN Sample Corpus (Public)
收藏资源简介:
The CO.RA.PAN Sample Corpus (Public) provides a small, fully annotated subset of the CO.RA.PAN corpus (Corpus Radiofónico Panhispánico). CO.RA.PAN investigates pluricentric standard Spanish based on radio news broadcasts from almost all Spanish-speaking countries. The full corpus comprises time-aligned audio, transcripts, and rich linguistic annotation (approx. 1.4 million words). This public Sample Corpus contains only JSON transcript files, not audio. Each JSON file mirrors the structure of the corresponding Full Corpus transcript and includes token-level linguistic annotation (POS, lemma, dependencies, morphology) as well as time alignment to the original audio. Time alignment is preserved through millisecond-based utterance and token offsets, even though audio files are not distributed here. The Sample Corpus is intended for testing, teaching, demonstrations, reproducibility, and methodological documentation. It enables users to work with the full annotation schema and data structure without requiring access to the restricted Full Corpus. CO.RA.PAN Resources and DOIs CO.RA.PAN Full Corpus (Restricted)DOI: https://doi.org/10.5281/zenodo.15360942 CO.RA.PAN Sample Corpus (Public)DOI: https://doi.org/10.5281/zenodo.15378479 CO.RA.PAN Metadata (Public)DOI: https://doi.org/10.5281/zenodo.17843469 CO.RA.PAN Web ApplicationDOI: https://doi.org/10.5281/zenodo.17834023 Web Application Access The CO.RA.PAN Web App is available at:https://corapan.online.uni-marburg.de The Web App provides authenticated access to the restricted Full Corpus and implements search, exploration, visualization, and metadata services. Source code and documentation of the Web App:https://github.com/FTacke/corapan-webapp Project Overview Further documentation and related digital humanities projects can be found at:https://hispanistica.online.uni-marburg.de/ Contents of this Sample Corpus • JSON transcript files for selected recordings• One JSON Schema file documenting the structure of a CO.RA.PAN transcript• README file describing annotation, alignment, and field semantics The Sample Corpus does not include audio files. File names referencing MP3 recordings indicate the corresponding audio file names found only in the restricted Full Corpus. Annotation Details All transcripts in this Sample Corpus include: • Tokenization, sentence segmentation, and utterance structure• POS tags, lemmas, and dependency relations• Universal Dependencies–style morphological features• Time-aligned token spans (start_ms, end_ms)• Time-aligned utterance spans (utt_start_ms, utt_end_ms)• Speaker metadata (speaker_type, sex, mode, discourse category)• Annotation metadata describing pipeline version, spaCy model, and transcript hash Licensing and Access This Sample Corpus is released under the license specified in each transcript (typically CC BY-NC 4.0).The Full Corpus remains restricted due to copyright constraints on the audio material. Users are asked to cite all relevant DOIs (Sample Corpus, Full Corpus, Metadata, Web Application) when using CO.RA.PAN.
CO.RA.PAN 公开样本语料库(Public)是CO.RA.PAN 泛伊比利亚广播语料库(Corpus Radiofónico Panhispánico,以下简称完整语料库)的小型全标注子集。CO.RA.PAN 泛伊比利亚广播语料库以几乎所有西班牙语国家的广播新闻节目为数据源,旨在研究多中心标准西班牙语。完整语料库包含时间对齐音频、转写文本与丰富的语言标注内容,总词量约140万。 本公开样本语料库仅包含JSON格式转写文件,不提供音频文件。每个JSON文件均与对应完整语料库的转写文本结构一致,包含Token级语言标注(包括词性标注POS、词元lemma、依存关系dependencies与词形特征morphology)以及与原始音频的时间对齐信息。尽管本次发布未包含音频文件,但通过基于毫秒级的话语与Token偏移量,保留了时间对齐信息。 本样本语料库旨在用于测试、教学、演示、可复现性研究与方法学文档编写,使用者可在无需获取受限完整语料库的前提下,使用完整的标注体系与数据结构开展研究。 ## CO.RA.PAN 相关资源与数字对象标识符(DOI) CO.RA.PAN 受限完整语料库(Restricted)DOI:https://doi.org/10.5281/zenodo.15360942 CO.RA.PAN 公开样本语料库(Public)DOI:https://doi.org/10.5281/zenodo.15378479 CO.RA.PAN 公开元数据DOI:https://doi.org/10.5281/zenodo.17843469 CO.RA.PAN 网页应用程序DOI:https://doi.org/10.5281/zenodo.17834023 ## 网页应用程序访问 CO.RA.PAN 网页应用程序的访问地址为:https://corapan.online.uni-marburg.de 该网页应用程序支持对受限完整语料库的授权访问,并提供检索、探索、可视化与元数据服务。 该网页应用程序的源代码与文档地址:https://github.com/FTacke/corapan-webapp ## 项目概览 更多文档与相关数字人文项目可访问:https://hispanistica.online.uni-marburg.de/ ## 本样本语料库包含内容 • 所选录音的JSON格式转写文件 • 一份用于说明CO.RA.PAN转写文本结构的JSON Schema文件 • 一份说明标注规则、对齐方式与字段语义的README文件 本样本语料库不包含音频文件。文件名中提及的MP3录音仅能在受限完整语料库中找到对应的音频文件。 ## 标注详情 本样本语料库中的所有转写文本均包含以下内容: • Token分词、分句与话语结构 • 词性标注(POS)、词元(lemma)与依存关系(dependencies) • 通用依存句法(Universal Dependencies)风格的词形特征 • 时间对齐的Token片段(start_ms、end_ms) • 时间对齐的话语片段(utt_start_ms、utt_end_ms) • 说话者元数据(speaker_type、sex、mode、discourse category) • 标注元数据,包括处理流程版本、spaCy模型与转写文本哈希值 ## 授权与访问 本样本语料库的授权协议见各转写文本文件(通常为知识共享署名-非商业性使用4.0国际许可协议(CC BY-NC 4.0))。由于音频素材受版权限制,完整语料库仍处于受限访问状态。 使用者在使用CO.RA.PAN相关资源时,需引用所有相关的数字对象标识符(DOI),包括样本语料库、完整语料库、元数据与网页应用程序对应的DOI。



