遇见数据集

Using Sentiment Analysis to investigate Sandro Abruzzese's literary works

收藏
Mendeley Data2026-04-18 收录
官方服务:

资源简介:

The dataset was created to support a comparative geo-affective analysis of Sandro Abruzzese’s literary works, with a focus on Casa per casa and Niente da vedere. Starting from digitized and normalized texts (typographic cleaning, diacritic and punctuation harmonization, and segmentation), place names (toponyms) were extracted and their occurrences were counted by work and in the aggregated corpus (ALL). Each toponym was then enriched with geographic information (latitude/longitude coordinates, when available) and with a semantic–affective profile computed through SenticNet, including polarity and the Hourglass dimensions (attitude, introspection, sensitivity, temper). In parallel, toponyms were embedded in AffectiveSpace and summarized via Principal Component Analysis (PCA) to obtain a compact 2D representation (PC1–PC2) suitable for cross-text comparison. Finally, clustering was performed on the geo-affective feature set (k = 5), generating cluster labels and interpretable descriptors for each group. The resulting dataset is delivered in a tabular format (CSV), ready for statistical exploration, visualization, and GIS integration.

本数据集旨在支撑桑德罗·阿布鲁泽塞(Sandro Abruzzese)文学作品的比较地理情感分析,研究重点聚焦于《Casa per casa》与《Niente da vedere》两部作品。首先对经数字化处理与标准化的文本(涵盖排版清理、变音符号与标点统一、文本分段)开展预处理,提取其中的地名(toponym),并统计各作品及总语料库(ALL)中地名的出现频次。随后,为每个地名补充地理信息(若可获取则包含经纬度坐标),并通过SenticNet计算其语义情感特征,包括情感极性与沙漏维度的四项指标:态度、内省性、敏感性与性情。与此同时,将所有地名嵌入情感空间(AffectiveSpace),并通过主成分分析(Principal Component Analysis,PCA)进行降维汇总,得到适用于跨文本比较的紧凑二维表征(PC1–PC2)。最后,基于地理情感特征集(聚类数k=5)开展聚类分析,为每个聚类生成簇标签与可解释性描述信息。最终生成的数据集以表格格式(CSV)交付,可直接用于统计探索、可视化及地理信息系统(Geographic Information System,GIS)集成。

创建时间:
2026-02-04
二维码
社区交流群
二维码
科研交流群
商业服务