遇见数据集

Annotated Marathi Word Sense Disambiguation Dataset

收藏
Mendeley Data2026-09-09 收录
官方服务:

资源简介:

This dataset is a large-scale sense-annotated resource for Word Sense Disambiguation (WSD) in Marathi, a morphologically rich and relatively low-resource Indic language. It was developed to address the limited availability of reusable, sense-annotated Marathi corpora for supervised WSD research. The final dataset contains 50,242 binary sentence–gloss pair records, covering 16 ambiguous Marathi words and 50 distinct senses across different grammatical categories, including nouns, verbs, adjectives, adverbs, and postpositions. The target words and their candidate senses were derived from the Marathi WordNet. Naturally occurring Marathi sentences were collected from established Marathi news portals and blogs, including Loksatta, Esakal, Lokmat, ABP Live Marathi, BBC Marathi, TV9 Marathi, News18 Marathi, Maharashtra Times, Mitraho, Marathi Spandan, and Maayboli. The collected text was cleaned, sentence-segmented, filtered based on sentence length, and duplicate sentences were removed to obtain a suitable corpus for WSD. Since Marathi exhibits substantial inflectional variation, the Stanford Stanza Marathi lemmatizer was used to identify target words in their different morphological forms. A prefix-based matching strategy was additionally applied when lemmatization failed for particular cases. For each word-sense pair, seed sentences were generated using the corresponding Marathi WordNet gloss and manually reviewed for clarity, grammatical correctness, and sense consistency. These verified examples were subsequently used as reference sentences for constructing sense representations. Candidate sentences were assigned to senses using MuRIL token-level embeddings and cosine similarity. A similarity threshold of 0.65 was used for the primary assignment, while a relaxed threshold of 0.60 was used for controlled padding of underrepresented senses. The dataset follows a binary sentence–gloss pair formulation. For every sentence containing an ambiguous target word, the sentence is paired with each candidate sense gloss. The correct sentence–sense pairing receives label 1, whereas incorrect pairings receive label 0. This structure makes the dataset suitable for cross-encoder WSD models and also allows conversion into a conventional multi-class classification format. Each record contains seven fields: sentence, target_word, lemma, pos, sense_id, gloss, and label. The dataset can be used for Marathi WSD benchmarking, supervised learning, transformer-based cross-encoder research, lexical-semantic analysis, contextual representation learning, and NLP research for low-resource Indic languages.

本数据集是面向马拉地语的词义消歧(Word Sense Disambiguation,WSD)任务的大规模词义标注资源,马拉地语是一种形态丰富且资源相对匮乏的印度语支语言。本数据集的构建旨在解决监督式WSD研究中可复用的马拉地语词义标注语料库稀缺的问题。最终数据集包含50242条二元句-释义对记录,涵盖16个歧义马拉地语词汇,以及名词、动词、形容词、副词和后置词等不同语法类别的50个独立词义。目标词汇及其候选词义均源自马拉地语词网。自然真实的马拉地语句本采集自成熟的马拉地语新闻门户与博客平台,包括Loksatta、Esakal、Lokmat、ABP Live Marathi、BBC Marathi、TV9 Marathi、News18 Marathi、Maharashtra Times、Mitraho、Marathi Spandan以及Maayboli。采集到的文本经过清洗、分句、基于句长过滤,并移除重复句本,以构建适配WSD任务的语料库。 鉴于马拉地语存在显著的屈折变化,研究人员使用斯坦福Stanza马拉地语词形还原器来识别不同形态形式下的目标词汇;对于部分词形还原失败的案例,额外采用了基于前缀的匹配策略。针对每个词-义对,研究人员利用对应的马拉地语词网释义生成种子句本,并对其进行人工审核,确保表述清晰、语法正确且词义一致。经审核通过的示例随后被用作构建词义表征的参考句本。 候选句本通过MuRIL的Token级嵌入与余弦相似度进行词义分配:主分配阶段采用0.65的相似度阈值,而为了对覆盖不足的词义进行可控填充,则使用0.60的宽松阈值。本数据集采用二元句-释义对的构建形式:对于每一条包含歧义目标词汇的句本,将其与每个候选词义释义进行配对,正确的句-义配对标注为1,错误配对则标注为0。该结构既适配跨编码器WSD模型,也可转换为常规的多分类分类格式。 每条记录包含7个字段:句本、目标词汇、词元、词性、词义ID、释义以及标签。本数据集可用于马拉地语WSD基准测试、监督学习、基于Transformer的跨编码器研究、词汇语义分析、上下文表征学习,以及面向低资源印度语支语言的自然语言处理研究。

创建时间:
2026-09-06
二维码
社区交流群
二维码
科研交流群
商业服务