81melody/algerian-realestate-ner-dataset
收藏资源简介:
这是一个专门用于命名实体识别的数据集,源自阿尔及利亚数字房地产市场Facebook群组中的复杂现实场景。数据集包含13个标记实体和7138个训练示例,涵盖了阿尔及利亚达尔贾方言(Darja)、阿拉伯语、法语和阿拉伯数字(Arabizi)之间的代码转换。该数据集通过提供标注语料库,填补了低资源方言NLP的空白,能够训练最先进的房地产信息提取模型。数据集分为标准的训练集、验证集和测试集,并采用了27标签的BIO(Begin, Inside, Outside)模式,基于13个核心实体类型。隐私保护方面,电话号码等个人信息被匿名化处理。数据集适用于房地产领域的信息提取任务,支持多种语言和方言的混合使用。
This is a specialized Named Entity Recognition dataset extracted from the complex reality of the Algerian digital real-estate market in Facebook groups, it contains 13 labeled entity and 7138 training example. Real estate advertisements in Algeria (found on Facebook groups) are unstructured, noisy and bloated with code-switching between Algerian Darja (dialect), Arabizi, Standard Arabic, and French. This dataset bridges the gap for low resource dialectal NLP by providing a labeled corpus capable of training state-of-the-art real-estate information extraction models. The dataset is split into standard Training, Validation, and Test sets and uses a 27 label BIO (Begin, Inside, Outside) schema based on 13 core entity types essential for real-estate information extraction. Privacy measures include anonymization of Personally Identifiable Information. The dataset is designed for real estate domain information extraction tasks, supporting multiple languages and dialects.




