InfoBayAI/Malayalam-Non-STEM-Educational-Text-Corpus
收藏资源简介:
该数据集是一个大规模马拉雅拉姆语非STEM教科书数据集合,包含149本书和560万个单词,旨在支持马拉雅拉姆语语言理解、推理和常识学习的高级NLP系统和AI模型的开发与训练。完整数据集概览显示,它是大规模多语言教育语料库的一部分,涵盖超过26亿单词、5000多个主题、15种语言的39000多本书,并包含交织的图像以支持更深层次的上下文理解,专为多语言知识学习、推理和教育理解的高级NLP系统与AI模型的开发与训练而设计。数据集规格包括:书籍数量为149本,单词数为560万,模态为马拉雅拉姆语,类型为教育/非STEM,数据来源为精选的学术教科书和教育材料,数据性质为真实世界和精选数据,内容涵盖人文、社会科学、文学、历史、地理、公民教育和教育解释。关键用例包括非STEM内容的命名实体识别、教育文本摘要与理解、自动化辅导与教育助手、知识检索系统以及模型评估与基准测试。数据集价值包括支持马拉雅拉姆语非STEM学科学习、提升AI模型的推理与理解能力、支持多语言和领域特定NLP系统、帮助构建AI驱动的教育平台以及增强LLM在非STEM领域的准确性和可靠性。
This dataset is a large-scale collection of Malayalam Non-STEM textbook data, containing 149 books and 5.60 million words, designed to support the development and training of advanced NLP systems and AI models for language understanding, reasoning, and general knowledge learning in Malayalam. Full Dataset Overview: This dataset is part of a large-scale multilingual educational corpus containing over 2.6+ billion words across 5,000+ subjects, supported by interwoven images for deeper contextual understanding. It includes 39,000+ books across 15 languages, designed to support the development and training of advanced NLP systems and AI models for multilingual knowledge learning, reasoning, and educational understanding.




