遇见数据集

Uncanny Semantics - Dataset

收藏
Zenodo2025-07-02 更新2026-05-26 收录
官方服务:

资源简介:

This dataset contains the curated results of a computational linguistics script used in Wegerhoff (2025). The script clusters lemmas into semantic categories based on a cosine similarity threshold of 0.7, using a corpus of 325 German academic texts in the field of linguistics—both AI-generated and human-authored. Multiple runs were conducted: individual runs for each POS category (nouns, verbs, adjectives, adverbs), as well as combined runs (e.g., adjectives + adverbs, nouns + verbs, and one run including all categories simultaneously). How to Read This Dataset The heatmap files display the 100 most frequent semantic categories for each tested AI model in each run. Note: The value for a category like “analytisch” does not reflect the frequency of the lemma analytisch itself, but the frequency of all lemmas whose cosine similarity to analytisch is 0.7 or higher. The *_members.txt files list the lemmas assigned to each semantic category for each run. The .xlsx files contain the complete numeric results of each run and can serve as a foundation for further statistical analysis. Script RepositoryGitHub: https://github.com/DayJay1992/SemanticAIAnalysis/

提供机构:
Zenodo
创建时间:
2025-07-02
二维码
社区交流群
二维码
科研交流群
商业服务