遇见数据集

BanglaAbuseMeme: A Dataset for Bengali Abusive Meme Classification

收藏
NIAID Data Ecosystem2026-05-01 收录
数据链接:
官方服务:

资源简介:

The dramatic increase in the use of social media platforms for information sharing has also driven a significant rise in online abuse. A simple yet effective method of targeting individuals or communities involves creating memes, often combining an image with a brief piece of text overlaid on top. These harmful elements are widely used and pose a threat to online safety. Therefore, developing efficient models for detecting and flagging abusive memes is essential. This challenge becomes even more demanding in a low-resource setting, such as Bengali memes (i.e., images with Bengali text embedded in them), due to the absence of benchmark datasets on which AI models can be trained. This work addresses this gap by creating a Bengali meme dataset of 4K data points. This upload includes the labeled Bengali meme dataset obtained from the web, described in the paper 'BanglaAbuseMeme: A Dataset for Bengali Abusive Meme Classification.'

随着社交媒体平台用于信息分享的使用量激增,网络滥用行为也显著增多。针对个人或群体的一种简单却高效的攻击手段,便是制作梗图(memes)——通常是将图像与叠加其上的简短文本相结合。此类有害内容被广泛传播,对网络安全构成威胁。因此,研发能够检测并标记攻击性梗图的高效模型至关重要。而在低资源场景下,这类检测挑战的难度会进一步提升——例如孟加拉语梗图(即嵌入孟加拉语文本的图像),由于目前缺乏可供AI模型训练使用的基准数据集。本研究通过构建包含4000个数据点的孟加拉语梗图数据集,填补了这一研究空白。 本次上传的内容包含从网络获取的带标注孟加拉语梗图数据集,相关研究细节详见论文《BanglaAbuseMeme: A Dataset for Bengali Abusive Meme Classification》。

创建时间:
2023-10-26
二维码
社区交流群
二维码
科研交流群
商业服务