bermaneh/codeswitching-sentiment-bias-exp5-neighborhood-v1
收藏资源简介:
该数据集名为codeswitching-sentiment-bias-exp5-neighborhood-v1,主要用于研究语言边界处的单词或在推文中少数语言中的单词是否获得了不成比例的高SHAP重要性分数,以测试模型是否将代码转换的语言功能内化为强调。数据集包含46217个词位记录,来自3360个句子,输入数据来自LinCE Spa-Eng推文的SHAP值和LinCE lid标签。数据集包含多个变量和列,如neighbor_config(孤立/左边界/右边界/内部)、is_minority(单词是否在推文的少数语言中)、sentence_id(句子ID)、word_position(单词位置)、word(单词文本)、lang(语言标签)等,用于详细分析语言边界和少数语言对SHAP值的影响。
The dataset is named codeswitching-sentiment-bias-exp5-neighborhood-v1 and is primarily used to investigate whether words at language boundaries or in the tweets minority language receive disproportionately higher SHAP importance scores, testing whether models have internalized the linguistic function of code-switching as emphasis. The dataset contains 46,217 word-position records from 3,360 sentences, with input data from LinCE Spa-Eng tweets SHAP values and LinCE lid tags. It includes multiple variables and columns such as neighbor_config (isolated/left_boundary/right_boundary/interior), is_minority (whether the word is in the tweets minority language), sentence_id (sentence ID), word_position (word position), word (word text), lang (language tag), etc., for detailed analysis of the impact of language boundaries and minority language on SHAP values.




