BanglaBias
收藏资源简介:
BanglaBias是一个包含200篇政治意义显著且高度争议的孟加拉语新闻文章的基准数据集,这些文章被标记为政府倾向、政府批评和中立立场。该数据集为评估大型语言模型(LLMs)提供了诊断分析。数据集的创建过程包括从多个新闻来源和博客收集政治上有争议的事件,然后由三位母语为孟加拉语的人对这些文章进行标注。BanglaBias旨在解决孟加拉语新闻中政治立场检测的挑战,并为低资源环境中的LLM性能改进提供见解。
BanglaBias is a benchmark dataset containing 200 politically significant and highly controversial Bengali news articles, which are labeled with three political stances: government-aligned, government-critical, and neutral. This dataset provides diagnostic analysis for evaluating large language models (LLMs). The dataset was created by collecting politically controversial events from multiple news sources and blogs, followed by annotation conducted by three native Bengali speakers. BanglaBias aims to address the challenges of political stance detection in Bengali news, and offer insights for improving LLM performance in low-resource settings.
- 1Read Between the Lines: A Benchmark for Uncovering Political Bias in Bangla News Articles达卡大学,马里兰大学巴尔的摩分校,孟加拉语LLM,玛哈希国际大学,Cisco Systems,Unityflow AI · 2025年



