遇见数据集

SindhiLanguageorg/Sindhi-Open-Lexicon

收藏
Hugging Face2026-05-14 更新2026-05-31 收录
官方服务:

资源简介:

这是一个为信德语(Sindhi)准备的开发者就绪和AI就绪的主词典数据集,是可用于AI和NLP的最大结构化信德语数据集之一。它包含223,342个词条,这些词条来源于多个信德语词典和术语来源,并已标准化为CSV、JSONL和SQLite格式。信德语是一种历史丰富但在人工智能领域资源匮乏的语言,该数据集旨在帮助改变这一状况。

This is a developer-ready and AI-ready master lexical dataset for the Sindhi language — one of the largest structured Sindhi datasets available for AI and NLP use. It contains 223,342 lexical entries drawn from multiple Sindhi dictionary and terminology sources, normalized into CSV, JSONL, and SQLite formats. Sindhi is a historically rich but low-resource language in AI; this dataset aims to help change that.

提供机构:
SindhiLanguageorg
二维码
社区交流群
二维码
科研交流群
商业服务