IndicSQuAD
收藏资源简介:
IndicSQuAD是一个全面的多语言问答数据集,涵盖了九种主要的印度语言,系统地从SQuAD数据集中衍生而来。该数据集由L3Cube Labs创建,旨在解决印度语言资源匮乏的问题,为低资源语言模型的研究和开发提供坚实的基础。IndicSQuAD包括每个语言的广泛训练、验证和测试集,总共有超过15万个问答对,确保了语言的高保真度和准确答案跨度对齐。该数据集的创建过程涉及从英语SQuAD数据集中翻译和调整,以适应印度语言的独特特征,如形态变化、句法差异等。IndicSQuAD的应用领域包括信息检索、教育、医疗保健服务、治理应用以及人工智能驱动的客户支持系统,旨在为印度语言使用者提供更好的知识获取渠道,减少数字不平等现象。
IndicSQuAD is a comprehensive multilingual question answering (QA) dataset that covers nine major Indian languages and is systematically derived from the SQuAD dataset. This dataset was developed by L3Cube Labs to address the shortage of language resources for Indian languages and establish a solid foundation for research and development of low-resource language models. IndicSQuAD provides extensive training, validation, and test sets for each of the covered languages, boasting a total of over 150,000 question-answer pairs, which ensures high linguistic fidelity and precise alignment of answer spans. The creation of this dataset involved translating and adapting the English SQuAD dataset to cater to the unique linguistic features of Indian languages, including morphological variations, syntactic differences, and other similar characteristics. Application domains of IndicSQuAD encompass information retrieval, education, healthcare services, governance applications, and AI-powered customer support systems, with the aim of enhancing knowledge access for users of Indian languages and mitigating digital inequality.
L3Cube-IndicNLP 数据集概述
项目简介
- L3Cube-IndicNLP项目旨在为印度语言改进NLP资源。
- 包含10种印度语言的单语BERT模型。
- 提供单语和多语言(跨语言)Sentence BERT模型。
- 这些模型在下游任务中提供最先进的结果。
单语BERT模型
- 详细论文:https://arxiv.org/abs/2211.11418
- 包含以下语言的BERT模型:
- 马拉地语(Marathi BERT)
- 印地语(Hindi BERT)
- Dev BERT(印地语+马拉地语)
- 卡纳达语(Kannada BERT)
- 泰卢固语(Telugu BERT)
- 马拉雅拉姆语(Malayalam BERT)
- 泰米尔语(Tamil BERT)
- 古吉拉特语(Gujarati BERT)
- 奥里亚语(Oriya BERT)
- 孟加拉语(Bengali BERT)
- 旁遮普语(Punjabi BERT)
- 阿萨姆语(Assamese BERT)
印度语言Sentence BERT模型
- 详细论文:https://arxiv.org/abs/2304.11434
- 包含以下语言的相似度模型和Sentence BERT模型:
- 马拉地语(Marathi Similarity, Marathi SBERT)
- 印地语(Hindi Similarity, Hindi SBERT)
- 卡纳达语(Kannada Similarity, Kannada SBERT)
- 泰卢固语(Telugu Similarity, Telugu SBERT)
- 马拉雅拉姆语(Malayalam Similarity, Malayalam SBERT)
- 泰米尔语(Tamil Similarity, Tamil SBERT)
- 古吉拉特语(Gujarati Similarity, Gujarati SBERT)
- 奥里亚语(Oriya Similarity, Oriya SBERT)
- 孟加拉语(Bengali Similarity, Bengali SBERT)
- 旁遮普语(Punjabi Similarity, Punjabi SBERT)
- 印度语言(多语言)(Indic Similarity, Indic SBERT)
许可证
- 所有资源均采用知识共享署名-非商业性使用-相同方式共享4.0国际许可协议(CC BY-NC-SA 4.0)。
- 数据集仅供研究用途。
引用
bibtex @article{joshi2022l3cube_hind, title={L3Cube-HindBERT and DevBERT: Pre-Trained BERT Transformer models for Devanagari based Hindi and Marathi Languages}, author={Joshi, Raviraj}, journal={arXiv preprint arXiv:2211.11418}, year={2022} }
bibtex @article{deode2023l3cube, title={L3Cube-IndicSBERT: A simple approach for learning cross-lingual sentence representations using multilingual BERT}, author={Deode, Samruddhi and Gadre, Janhavi and Kajale, Aditi and Joshi, Ananya and Joshi, Raviraj}, journal={arXiv preprint arXiv:2304.11434}, year={2023} }
相关出版物
- Joshi, Raviraj. "L3Cube-HindBERT and DevBERT: Pre-Trained BERT Transformer models for Devanagari based Hindi and Marathi Languages." arXiv preprint arXiv:2211.11418 (2022).
- Deode, Samruddhi, et al. "L3Cube-IndicSBERT: A simple approach for learning cross-lingual sentence representations using multilingual BERT." arXiv preprint arXiv:2304.11434 (2023).
- Mirashi, Aishwarya, et al. "L3Cube-IndicNews: News-based Short Text and Long Document Classification Datasets in Indic Languages." arXiv preprint arXiv:2401.02254 (2024).

- 1L3Cube-IndicQuest: A Benchmark Questing Answering Dataset for Evaluating Knowledge of LLMs in Indic ContextL3Cube Labs, Pune · 2024年



