Corpus92: Russian-language YouTube Vlogs about Bulgaria for Corpus-Discourse Analysis of Acculturation Markers and Strategies
收藏资源简介:
This dataset contains the reproducible Corpus92 research package for a corpus-discourse study of acculturation markers and strategies in Russian-language YouTube vlogs about Bulgaria. The verified corpus release is Corpus92 v1.1.172 and comprises 92 included videos from 10 Russian-language YouTube channels, 92 verified Russian transcripts, 1,556.5 minutes of video material and 174,555 words. The current analytical layer is ContextAuto v2.5, a context-sensitive automatic coding procedure with manual calibration and validation. It is not a simple keyword-count model. The procedure distinguishes source/home/source-language comparison from internal comparisons within Bulgaria, separates substantive institutional/procedural or infrastructure-related context from incidental mentions, and identifies primary domains by the semantic centre of the video rather than by raw token frequency alone. Manual calibration and validation were conducted on the discussed R20_01–R20_15 cases; the calibration audit is included in the package. This audit is not Cohen’s kappa. Inter-coder reliability is planned separately on independent human coding before final adjudication. Full third-party YouTube video/audio files and full transcript redistribution are not included in the open dataset due to copyright, platform-related and research-ethics considerations. The CC BY 4.0 license applies only to the author-created metadata, quality-control registers, coding tables, calibrated analytical coding, aggregated metrics, lexicons, methodological documentation and reproducibility materials included in this deposit.



