Morphology of Czech fake news from the viewpoint of quantitative linguistics.
收藏资源简介:
The dataset contains quantitative linguistic data describing the morphological profiles of five types of Czech news texts — credible, manipulative, misleading, partially credible, and unclassifiable. It consists of csv files, 3 python scripts and two figures (png). CSV files include values for the researched texts for a set of morphological indicators, such as the relative frequencies of parts of speech, grammatical cases, pronoun types, and deverbative adjectives. These variables serve as input for Principal Component Analysis (PCA). The accompanying figures show (1) the results of the permutation test for the first two principal components, and (2) the PC1–PC2 biplot with centroids and 1SD ellipses of data dispersion. The dataset is intended for quantitative linguistic and stylometric research, especially for analyzing and visualizing morphological differentiation among news types. Python scripts are published under licence CC-BY-4.0 Data files are published under licence CC-BY-NC-SA-4.0



