遇见数据集

sentiwordnet_it 1.0

收藏
Zenodo2026-06-17 更新2026-05-26 收录
官方服务:

资源简介:

This repository contains a sentiment lexicon for Italian, based on SentiWordNet 3.0 (Baccianella, Esuli, and Sebastiani 2010; Esuli [2019] 2025) and MultiWordNet (Pianta, Bentivogli, and Girardi 2002). Unlike previous resources—SentiWordNet, which provides sentiment scores without Italian lexical coverage, and MultiWordNet, which offers Italian synsets without sentiment annotation—this dataset bridges the two by mapping Italian lexical entries to sentiment scores in a ready-to-use CSV format. This integration enables direct use in sentiment analysis and other NLP applications for Italian, filling a gap in existing resources. The included files, in the data/ folder are: swn_it.csv: A dataset of 35,001 Italian synsets with polarity scores, POS, synset, offset, English synset lemmas, and gloss (in English). swn_it_tidy.csv: A tidy (one token per row) dataset of 41,725 lemmas, with polarity scores. It is designed for use in R. It also contains a folder with examples in R, and scripts to use and manipulate the datasets: examples-R/: custom_dataset.R: Create a custom tidy dataset from the original one, for treating duplicate entries differently. example.R: Examples of how to use the dataset for sentiment analysis on a sample text. uso.md: Instructions for using the dataset in R (in Italian), referred to in example.R.

提供机构:
Zenodo
创建时间:
2025-10-02
二维码
社区交流群
二维码
科研交流群
商业服务