Dataset: Corpus-Based Lexical Complexity in Czech: New Approaches and Their Application to News Credibility Analysis
收藏资源简介:
The dataset contains quantitative linguistic data describing the lexical profiles of five types of Czech news texts — credible, manipulative, misleading, partially credible, and unclassifiable. It consists of 3 XLSX files, 3 TXT files, and 1 MD file. The files include a list of lemmata with the corresponding Normalized Index of Lexical Complexity (NICL) values for each lemma; a list of lemmatized bigrams with the corresponding NICL values for each bigram; lexical resources of lemmata characterizing journalistic, fiction, and academic language, which are the basis for the PUB_SCORE, FIC_SCORE, and ACAD_SCORE indices; a list of affixoids on which the OID index is based; and the raw results for each text in the Verifee corpus for all seven lexical indices.



