InfoBayAI/Marathi-STEM-Educational-Text-Corpus
收藏资源简介:
该数据集是一个大规模马拉地语STEM教科书数据集,包含173本书和781万个单词,旨在支持马拉地语科学理解、问题解决和概念学习的高级NLP系统和AI模型的开发和训练。它是大规模多语言教育语料库的一部分,该语料库涵盖超过260亿单词和39000多本书,涉及15种语言,用于多语言知识学习、推理和教育理解的NLP系统和AI模型训练。数据集具体包括书籍数量、单词数、模态为马拉地语、类型为教育/STEM、数据来源为精选学术教科书和教育材料、数据性质为真实世界和精选数据、内容涵盖科学概念、问题解决问题、答案和解释。关键用例包括STEM内容中的命名实体识别(NER)、科学文本摘要和理解、自动辅导和教育助手、STEM知识检索系统以及模型评估和基准测试。数据集的价值包括支持马拉地语STEM学科学习、提高AI模型的分析推理和问题解决能力、支持多语言和领域特定的NLP系统、帮助构建AI驱动的教育平台,以及增强LLM在STEM领域的准确性和可靠性。
This dataset is a large-scale collection of Marathi STEM textbook data, containing 173 books and 7.81 million words, designed to support the development and training of advanced NLP systems and AI models for scientific understanding, problem-solving, and concept learning in Marathi. It is part of a large-scale multilingual educational corpus containing over 2.6+ billion words across 5,000+ subjects, supported by interwoven images for deeper contextual understanding. It includes 39,000+ books across 15 languages, designed to support the development and training of advanced NLP systems and AI models for multilingual knowledge learning, reasoning, and educational understanding.




