Uncanny Semantics - Dataset
收藏资源简介:
This dataset contains the curated results of a computational linguistics script used in Wegerhoff (2025). The script clusters lemmas into semantic categories based on a cosine similarity threshold of 0.7, using a corpus of 325 German academic texts in the field of linguistics—both AI-generated and human-authored. Multiple runs were conducted: individual runs for each POS category (nouns, verbs, adjectives, adverbs), as well as combined runs (e.g., adjectives + adverbs, nouns + verbs, and one run including all categories simultaneously). How to Read This Dataset The heatmap files display the 100 most frequent semantic categories for each tested AI model in each run. Note: The value for a category like “analytisch” does not reflect the frequency of the lemma analytisch itself, but the frequency of all lemmas whose cosine similarity to analytisch is 0.7 or higher. The *_members.txt files list the lemmas assigned to each semantic category for each run. The .xlsx files contain the complete numeric results of each run and can serve as a foundation for further statistical analysis. Script RepositoryGitHub: https://github.com/DayJay1992/SemanticAIAnalysis/



