遇见数据集

CO.RA.PAN Sample Corpus (Public)

收藏
Zenodo2025-12-07 更新2026-05-26 收录
官方服务:

资源简介:

The CO.RA.PAN Sample Corpus (Public) provides a small, fully annotated subset of the CO.RA.PAN corpus (Corpus Radiofónico Panhispánico). CO.RA.PAN investigates pluricentric standard Spanish based on radio news broadcasts from almost all Spanish-speaking countries. The full corpus comprises time-aligned audio, transcripts, and rich linguistic annotation (approx. 1.4 million words). This public Sample Corpus contains only JSON transcript files, not audio. Each JSON file mirrors the structure of the corresponding Full Corpus transcript and includes token-level linguistic annotation (POS, lemma, dependencies, morphology) as well as time alignment to the original audio. Time alignment is preserved through millisecond-based utterance and token offsets, even though audio files are not distributed here. The Sample Corpus is intended for testing, teaching, demonstrations, reproducibility, and methodological documentation. It enables users to work with the full annotation schema and data structure without requiring access to the restricted Full Corpus. CO.RA.PAN Resources and DOIs CO.RA.PAN Full Corpus (Restricted)DOI: https://doi.org/10.5281/zenodo.15360942 CO.RA.PAN Sample Corpus (Public)DOI: https://doi.org/10.5281/zenodo.15378479 CO.RA.PAN Metadata (Public)DOI: https://doi.org/10.5281/zenodo.17843469 CO.RA.PAN Web ApplicationDOI: https://doi.org/10.5281/zenodo.17834023 Web Application Access The CO.RA.PAN Web App is available at:https://corapan.online.uni-marburg.de The Web App provides authenticated access to the restricted Full Corpus and implements search, exploration, visualization, and metadata services. Source code and documentation of the Web App:https://github.com/FTacke/corapan-webapp Project Overview Further documentation and related digital humanities projects can be found at:https://hispanistica.online.uni-marburg.de/ Contents of this Sample Corpus • JSON transcript files for selected recordings• One JSON Schema file documenting the structure of a CO.RA.PAN transcript• README file describing annotation, alignment, and field semantics The Sample Corpus does not include audio files. File names referencing MP3 recordings indicate the corresponding audio file names found only in the restricted Full Corpus. Annotation Details All transcripts in this Sample Corpus include: • Tokenization, sentence segmentation, and utterance structure• POS tags, lemmas, and dependency relations• Universal Dependencies–style morphological features• Time-aligned token spans (start_ms, end_ms)• Time-aligned utterance spans (utt_start_ms, utt_end_ms)• Speaker metadata (speaker_type, sex, mode, discourse category)• Annotation metadata describing pipeline version, spaCy model, and transcript hash Licensing and Access This Sample Corpus is released under the license specified in each transcript (typically CC BY-NC 4.0).The Full Corpus remains restricted due to copyright constraints on the audio material. Users are asked to cite all relevant DOIs (Sample Corpus, Full Corpus, Metadata, Web Application) when using CO.RA.PAN.

提供机构:
Zenodo
创建时间:
2025-05-10
二维码
社区交流群
二维码
科研交流群
商业服务