遇见数据集

bermaneh/codeswitching-sentiment-bias-exp2-results-v1

收藏
Hugging Face2026-04-28 更新2026-05-03 收录
官方服务:

资源简介:

该数据集是实验2的完整结果,涉及掩码标记语言预测。实验通过掩码处理代码转换句子中的关键词,并使用Qwen2.5-7B-Instruct模型来填充空白,测量模型在完成代码转换句子时的语言偏好。数据集包含多个列,如句子ID、原始句子、掩码句子、真实单词、真实语言等,并提供了详细的实验结果和指标。

Full results of Experiment 2: Masked Token Language Prediction. The experiment masks the top SHAP-contributing word in code-switched sentences and uses the Qwen2.5-7B-Instruct model to fill in the blank, measuring the models language preference when completing code-switched sentences. The dataset includes various columns such as sentence_id, original_sentence, masked_sentence, ground_truth_word, ground_truth_language, etc., and provides detailed experimental results and metrics.

提供机构:
bermaneh
二维码
社区交流群
二维码
科研交流群
商业服务