InfoBayAI/Sanskrit-Non-STEM-Educational-Text-Corpus
收藏资源简介:
这是一个大规模梵语非STEM教科书数据集,包含141本书和805万个单词,旨在支持高级NLP系统和AI模型在梵语语言理解、推理和古典知识学习方面的开发和训练。数据集涵盖人文、社会科学、文学、历史、地理、公民学和教育解释等内容,数据源为精选的学术教科书和教育材料,具有真实世界和精选数据的性质。它是更大规模多语言教育语料库的一部分,该语料库包含超过26亿单词、5000多个主题和15种语言的39000多本书,支持多语言知识学习、推理和教育理解的NLP系统开发。关键用例包括非STEM内容中的命名实体识别、教育文本摘要和理解、自动化辅导和教育助手、知识检索系统以及模型评估和基准测试。该数据集的价值在于支持梵语非STEM学科学习、提升AI模型的推理和理解能力、支持多语言和领域特定NLP系统、帮助构建AI驱动的教育平台,并增强LLM在非STEM领域的准确性和可靠性。
This dataset is a large-scale collection of Sanskrit Non-STEM textbook data, containing 141 books and 8.05 million words, designed to support the development and training of advanced NLP systems and AI models for language understanding, reasoning, and classical knowledge learning in Sanskrit. It is part of a large-scale multilingual educational corpus containing over 2.6+ billion words across 5,000+ subjects, supported by interwoven images for deeper contextual understanding, and includes 39,000+ books across 15 languages, designed to support multilingual knowledge learning, reasoning, and educational understanding. The dataset covers humanities, social sciences, literature, history, geography, civics, and educational explanations, with data sourced from curated academic textbooks and educational material, and is of real-world and curated nature. Key use cases include Named Entity Recognition (NER) in Non-STEM content, educational text summarization and comprehension, automated tutoring and educational assistants, knowledge retrieval systems, and model evaluation and benchmarking. Its value lies in enabling learning of Non-STEM subjects in Sanskrit, improving reasoning and comprehension capabilities in AI models, supporting multilingual and domain-specific NLP systems, helping build AI-powered educational platforms, and enhancing accuracy and reliability of LLMs in Non-STEM domains.




