InfoBayAI/Punjabi-Non-STEM-Educational-Text-Corpus
收藏资源简介:
该数据集是一个大规模旁遮普语非STEM教科书数据集合,包含270本书和2063万个单词,旨在支持高级NLP系统和AI模型的开发与训练,用于旁遮普语的语言理解、推理和常识学习。它是更大规模多语言教育语料库的一部分,该语料库涵盖超过26亿单词、5000多个主题和15种语言的39000多本书,并包含交织图像以提供更深层次的上下文理解。数据集内容涉及人文、社会科学、文学、历史、地理、公民学和教育解释等领域,数据来源于精选的学术教科书和教育材料,具有真实世界和精选性质。关键应用包括非STEM内容中的命名实体识别、教育文本摘要与理解、自动辅导与教育助手、知识检索系统以及模型评估与基准测试。数据集的价值在于支持旁遮普语非STEM学科学习、提升AI模型的推理和理解能力、促进多语言和领域特定NLP系统发展、助力构建AI驱动的教育平台,并增强大语言模型在非STEM领域的准确性和可靠性。
This is a large-scale Punjabi non-STEM textbook dataset, comprising 270 books and 20.63 million words. It is designed to support the development and training of advanced NLP systems and AI models for Punjabi language understanding, reasoning and commonsense learning. It is part of a larger multilingual educational corpus, which contains over 2.6 billion words, more than 5000 topics, over 39,000 books across 15 languages, and interleaved images to provide deeper contextual understanding. The dataset covers fields including humanities, social sciences, literature, history, geography, civics and educational explanations. The data is sourced from curated academic textbooks and educational materials, featuring real-world and curated attributes. Key applications include Named Entity Recognition (NER) in non-STEM content, educational text summarization and understanding, automated tutoring and educational assistants, knowledge retrieval systems, and model evaluation and benchmarking. The value of this dataset lies in supporting non-STEM subject learning in Punjabi, enhancing the reasoning and understanding capabilities of AI models, promoting the development of multilingual and domain-specific NLP systems, facilitating the construction of AI-driven educational platforms, and improving the accuracy and reliability of Large Language Models (LLMs) in non-STEM fields.




