NAQA Meeting Corpus
收藏资源简介:
The NAQA Meeting Corpus comprises transcripts of 69 official agency meetings recorded and published on NAQA's YouTube channel between October 12, 2020 and December 23, 2025. This period encompasses the agency's operational maturation, the COVID-19 pandemic, and Russia's full-scale invasion - providing a natural laboratory for examining discourse dynamics under varying conditions. Meetings were identified through systematic review of channel uploads. Selection criteria included: (1) official agency meetings (excluding press conferences, webinars, and other events); (2) availability of audio suitable for transcription; (3)~meeting duration exceeding 10 minutes. All identified meetings satisfying these criteria were included, eliminating selection bias. Two transcription methods were employed based on availability. For 33 videos (47.8%), YouTube's automatic captions provided initial transcripts. These were available for earlier meetings and those with clearer audio. For 36 videos (52.2%), OpenAI's Whisper large-v2 model generated transcripts. Whisper was selected for its strong multilingual capabilities and robust performance on Ukrainian speech. All transcripts underwent light normalization: removal of timestamp markers, standardization of whitespace, and correction of obvious segmentation errors. No substantive content was altered. The corpus exhibits characteristics typical of institutional spoken discourse: relatively low overall TTR reflecting repeated procedural language, but substantial vocabulary size indicating domain-specific terminology. The Heaps’ Law coefficient (β = 0.609) falls within expected ranges for specialized corpora, indicating sublinear vocabulary growth as the corpus expands. Table 1. NAQA Meeting Corpus summary statistics Metric Value Total meetings 69 Date range October 2020 – December 2025 Total duration (hours) 152.8 Total tokens 903,133 Filtered tokens (content words) 666,024 Unique word forms (types) 59,164 Overall type-token ratio (TTR) 0.066 Mean TTR per document 0.307 (±0.088) Mean MTLD per document 270.9 (±90.8) Overall lexical density 0.737 Overall hapax ratio 0.597 Yule’s K 14.53 Heaps’ Law β 0.609 Table 2 presents the distribution of meetings and word counts by year. Table 2. Yearly distribution of corpus documents Year Meetings Total Words Duration (min) 2020 13 144,776 1,611 2021 12 145,451 1,600 2022 6 91,527 953 2023 14 162,104 1,579 2024 10 132,208 1,241 2025 14 227,067 2,187 Total 69 903,133 9,171 The reduced meeting count in 2022 reflects disruption from the Russian invasion, with NAQA suspending regular operations during the initial months of full-scale war.



