ANCHOLIK-NER
收藏资源简介:
ANCHOLIK-NER是一个面向孟加拉区域方言命名词义识别的语言多样性格式数据集,覆盖了锡尔赫特、吉大港和巴里萨尔三个地区的方言变化。该数据集由约10443个句子组成,每个地区约3481个句子,数据来源于两个公开可用的数据集以及通过网页抓取的各种在线报纸和文章。数据集使用BIO标注方案进行高质量的专业标注,分为各自独立的子集,并以CSV格式提供,每个条目包含文本数据和识别的命名实体及其相应注释。
ANCHOLIK-NER is a linguistically diverse formatted dataset for named entity recognition (NER) targeting regional Bengali dialects, covering dialectal variations across three regions: Sylhet, Chittagong, and Barisal. Comprising approximately 10,443 sentences in total, with roughly 3,481 sentences per region, the dataset is sourced from two publicly available datasets and various online newspapers and articles collected via web scraping. The dataset employs the BIO annotation scheme for high-quality professional annotation, is divided into independent subsets, and is provided in CSV format, where each entry contains the text data, recognized named entities and their corresponding annotations.

- 1ANCHOLIK-NER: A Benchmark Dataset for Bangla Regional Named Entity Recognition阿萨努拉大学科学与技术学院计算机科学与工程系 · 2025年



