遇见数据集

indicvoices-cleaned

收藏
魔搭社区2026-04-28 更新2026-08-09 收录
官方服务:

资源简介:

# IndicVoices Cleaned **IndicVoices Cleaned** is a curated repository containing cleaned transcripts derived from the IndicVoices dataset. We start with normalized verbatim text from IndicVoices and further process it using Google's Gemini model to produce grammatically correct and fluent sentences. These cleaned sentences are intended for use in creating high-quality datasets for: - Machine translation - Transliteration - Other multilingual NLP tasks ## Supported Languages ### Primary Languages We offer robust support for the following 15 major Indian languages: > Assamese, Bengali, Gujarati, Hindi, Kannada, Maithili, Malayalam, Marathi, Nepali, Odia, Punjabi, Sanskrit, Tamil, Telugu, Urdu ### Low-Resource Languages We also provide preliminary support for the following 7 low-resource Indian languages: > Bodo, Dogri, Kashmiri, Konkani, Manipuri, Santali, Sindhi

提供机构:
maas
创建时间:
2026-03-09
二维码
社区交流群
二维码
科研交流群
商业服务