遇见数据集

Psycholinguistic LIWC and n-gram counts in a corpus of 1145 English and Dutch novels

收藏
Mendeley Data2026-04-18 收录
官方服务:

资源简介:

This dataset consists of CSV files with word counts in several corpora: - 694 English language novels from male and female authors classified by authors' sexual orientation (heterosexual, bisexual, homosexual) - 401 bestselling Dutch language novels - 50 novels nominated for Dutch literary prizes Each corpus comes with: - LIWC counts; this file also includes the available metadata for each novel. The English data was created with LIWC 2015. The Dutch data was created with the validated translation of LIWC 2001. - Word counts (unigrams) and bigram counts per novel. All text has been converted to lowercase. Contractions are tokenized into separate tokens, e.g., can't => ca n't Two restrictions are applied: - only unigrams or bigrams that occur in at least 10 texts are retained - only the 100k most frequent are retained - Overall word counts and bigram counts; i.e., the sum across all novels. All files are encoded in UTF-8. The word counts were extracted with the countngrams.py script.

本数据集包含涵盖多个语料库的词频统计CSV文件: - 694部英语小说,作者涵盖男女作者,且按作者性取向(异性恋、双性恋、同性恋)完成分类; - 401部畅销荷兰语小说; - 50部获荷兰文学奖提名的小说。 每个语料库均包含以下内容: - 语言学调查与词数统计(Linguistic Inquiry and Word Count, LIWC)统计文件:该文件同时收录每部小说的可用元数据。其中英语语料基于LIWC 2015生成,荷兰语语料则基于经验证的LIWC 2001翻译版本生成。 - 单部小说的单字元(unigrams)词频与二元组(bigrams)频次数值。所有文本均已转换为小写形式,缩略词将被拆分为独立的Token,例如can't拆分为ca n't。 本次统计共应用两项筛选规则: - 仅保留至少在10部小说中出现过的单字元或二元组; - 仅保留出现频率最高的前100,000个单元。 - 全局词频与二元组频次数值,即所有小说的统计总和。 所有文件均采用UTF-8编码。 本次词频统计工作由countngrams.py脚本完成。

创建时间:
2020-11-04
二维码
社区交流群
二维码
科研交流群
商业服务