遇见数据集

CO.RA.PAN Full Corpus (Restricted)

收藏
Zenodo2026-02-26 更新2026-05-26 收录
官方服务:

资源简介:

This record contains the complete CO.RA.PAN (Corpus Radiofónico Panhispánico) broadcast corpus, organized into country-specific ZIP archives with audio files and linguistically annotated JSON transcripts. Due to copyright and broadcasting rights, all media files and transcripts are distributed under restricted access and cannot be shared openly. Users may request access directly through Zenodo. A separate public record containing the open CO.RA.PAN metadata is available here:<INSERT DOI OF PUBLIC METADATA RECORD> The public metadata collection includes: corapan_recordings.tsv – complete tabular metadata for all recordings corapan_recordings.json – machine-readable metadata corapan_corpus_metadata.json – documentation of variables, corpus structure, and annotation conventions These metadata files enable FAIR-compliant reuse, support corpus exploration, allow bibliographic referencing and reproducibility of sampling procedures, and facilitate scientific work without requiring access to restricted audio or transcripts. Contents of this record Each {COUNTRYCODE}.zip archive contains: Audio recordings in mp3 format (mp3-files/) Annotated JSON transcripts (json-transcripts/) All ZIP archives were generated using the internal script "zenodo_corpus_zip.py", which automatically tracks timestamps and file changes to ensure reproducible versioning. Versioning Each version of this record represents a coherent snapshot of the full corpus at a specific point in time. Updates may include newly added recordings, corrected or extended transcripts, and improvements to preprocessing and linguistic annotation. Related datasets CO.RA.PAN Full Corpus (Restricted): https://doi.org/10.5281/zenodo.15360942 CO.RA.PAN Sample Corpus (Public): https://doi.org/10.5281/zenodo.15378479 CO.RA.PAN Metadata (Public): https://doi.org/10.5281/zenodo.17843469 CO.RA.PAN Web Application and Code: https://doi.org/10.5281/zenodo.17834023 Annotation details Each JSON transcript contains: tokenization, sentence segmentation POS tags, lemmas, and morphological features dependency relations automatic categorization of verbal tense and related featuresAll annotations are generated using spaCy (model: es_dep_news_trf), followed by project-specific quality control steps. Legal and access information The restricted status of this record is due to copyright and broadcasting limitations. Only short audio-text extracts may be displayed publicly under scientific quotation rules and text-and-data-mining provisions of EU Directive 2019/790 and the German UrhG (§51, §60d, §44b). Redistribution or reuse of the full recordings and transcripts is not permitted. Access requests can be submitted directly through Zenodo. For scientific inquiries or technical questions, please contact the CO.RA.PAN project team.

提供机构:
Zenodo
创建时间:
2025-05-07
二维码
社区交流群
二维码
科研交流群
商业服务