STEM_TextBook_English
收藏资源简介:
The full corpus is curated across multiple STEM disciplines and structured for use in LLM training, evaluation, and instruction tuning (SFT/RLHF). This sample represents the structure and quality of the larger dataset. Dataset composition (full corpus): Text corpus: 1.6B+ words of curated STEM and Non-STEM educational content across 22,000+ textbooks in 7 languages (English, Hindi, Arabic, Bahasa, Tamil, Telugu, Kannada) Question–Answer pairs: 6.5M+ high-quality Q&A pairs of STEM and Non-STEM content in English, Arabic, Hindi, and Indic languages Video data: 100K+ hours of STEM videos and 30K+ hours of UGC Audio data: 821K+ hours of podcasts and dual-channel call center data Medical datasets: 30M+ files including clinical and diagnostic data such as CT scans, MRI, X-ray, pathology, EHRs, USG reports, and echo reports This repository includes: A small preview subset of the STEM English TextBook data Flat, viewer-friendly schema for inspection Parquet files suitable for benchmarking and evaluation Purpose of this dataset: Dataset preview and validation Model evaluation and experimentation Schema and format inspection before full-scale access Note: This repository contains sample data only. Access to the complete dataset is available separately under appropriate licensing or partnership terms.For detailed information and domain-specific requirements, please reach out to vipul.mishra@infobay.ai



