遇见数据集

bermaneh/codeswitching-sentiment-bias-results-v1

收藏
Hugging Face2026-04-28 更新2026-05-03 收录
官方服务:

资源简介:

该数据集包含了一个实验的结果,该实验研究了在3,483条真实的双语推文(来自SemEval 2020 Task 9语料库)中进行英语和西班牙语之间的单字替换时,情感分析模型预测的变化。实验假设NLP模型编码了语言依赖的偏见,通过替换单词来测量模型情感预测的偏移。使用了`cardiffnlp/twitter-roberta-base-sentiment-latest`模型和SHAP解释性方法。数据集详细记录了原始句子、扰动后的句子、替换的单词、情感标签和得分变化等信息。

This dataset contains the results of an experiment investigating the impact of single-word English-Spanish code-switching on sentiment analysis model predictions in 3,483 real bilingual tweets (from the SemEval 2020 Task 9 corpus). The experiment hypothesizes that NLP models encode language-dependent biases, and swapping words between English and Spanish measurably shifts the models sentiment prediction. The `cardiffnlp/twitter-roberta-base-sentiment-latest` model and SHAP explainability method were used. The dataset includes detailed records of original sentences, perturbed sentences, swapped words, sentiment labels, and score changes.

提供机构:
bermaneh
二维码
社区交流群
二维码
科研交流群
商业服务