L3Cube-IndicHeadline-ID
收藏资源简介:
L3Cube-IndicHeadline-ID是一个用于低资源印度语言新闻标题识别和语义评估的数据集,涵盖了十种低资源印度语言:马拉地语、印地语、泰米尔语、古吉拉特语、奥里亚语、卡纳达语、马拉雅拉姆语、旁遮普语、泰卢固语和孟加拉语。每种语言包括20,000篇新闻文章,每篇文章配以四种标题变体:原始标题、语义相似版本、词汇相似版本和无关版本。数据集用于评估模型根据文章和标题之间的相似性选择正确标题的能力,为低资源印度语言的自然语言处理提供了宝贵的资源。
L3Cube-IndicHeadline-ID is a dataset for news headline identification and semantic evaluation in low-resource Indian languages, covering ten low-resource Indian languages: Marathi, Hindi, Tamil, Gujarati, Odia, Kannada, Malayalam, Punjabi, Telugu, and Bengali. Each language includes 20,000 news articles, with each article paired with four headline variants: the original headline, a semantically similar version, a lexically similar version, and an irrelevant version. This dataset is used to evaluate a model's ability to select the correct headline based on the similarity between the article and its corresponding headline, serving as a valuable resource for natural language processing (NLP) in low-resource Indian languages.



