Dataset of Corpus Analysis: Strategic Narrative of the US on the Russo-Ukrainian War (2014-2025)
收藏资源简介:
Corpus of Official US Texts Regarding the Russian-Ukrainian War (2014–2025) and Frequency Analysis Results of Lemmas and Phrases.The dataset has been prepared for academic use and reproducibility of results. Upon publication, the dataset should be cited using the DOI of the Zenodo record.This set of files contains quantitative data collected using Sketch Engine based on the official public communications of the US executive branch and selected additional statements by US presidents on the Russian-Ukrainian war. The materials cover three periods of the war: the hybrid period, the period of full-scale war during the presidency of J. Biden, and the period of full-scale war during the presidency of D. Trump. The data include: metadata of all texts used in the research; lemma counts for each of the three periods: total and split into sub-corpora by specific periods and contexts; collocation (n-gram) counts for each of the three periods: total and split into sub-corpora by specific periods and contexts. As part of the procedure for fixing lexical indicators, differentiated linguo-statistical thresholds were applied, which are necessary to limit the impact of statistical noise on qualitative discourse analysis: 1) a lemma occurrence threshold of ≥1; 2) for counting n-grams in macro-corpora, a threshold of ≥ 5 was set, while for sub-corpora it was set to ≥ 2 (due to its significantly smaller volume compared to other macro-corpora, a threshold of ≥ 2 was also established for the macro-corpus of D. Trump's presidency period). More details about the procedure are in the Readme file.



