遇见数据集

DataOrigin/ncert-lectures-india

收藏
Hugging Face2026-04-06 更新2026-04-12 收录
官方服务:

资源简介:

--- license: other task_categories: - audio-classification - automatic-speech-recognition language: - en - hi - bn - ta - te - ml - mr - or - as - pa tags: - ncert, - education, - india, - upsc, - humanities, - recorded-lectures, - government-exams, - science, - long-form-audio pretty_name: NCERT Lectures India size_categories: - 1K<n<10K # NCERT Lectures India ## Dataset Description A large-scale collection of recorded lectures covering NCERT curriculum across Science and Humanities streams, specifically designed for government exam preparation including UPSC, SSC, and State PSC examinations. Produced by Prepp, India's largest government exam preparation platform, operated by Collegedunia Web Private Limited. ## Dataset Summary - **Total duration:** 5,100 hours of recorded lectures - **Content type:** Structured curriculum lectures covering NCERT Science and Humanities - **Exam relevance:** UPSC Civil Services, SSC, State PSC, Railways, Banking, and all major government competitive examinations - **Curriculum:** Full NCERT coverage — Classes 6 through 12, Science and Humanities streams - **Chapters available:** History Ch-1, Ch-2, Ch-3 (samples) - **Languages:** Hindi and English primary; regional language variants available - **Format:** Audio/Video with structured chapter-by-chapter delivery ## Sample Data Three sample lectures are available in this repository: - History Chapter 1 — The Early Societies - History Chapter 2 — Early Economies and Empires - History Chapter 3 — Political and Economic History ## Key Features - **NCERT-aligned:** Follows official NCERT curriculum chapter structure exactly — high-value signal for Indian education AI models - **Government exam optimised:** Content structured specifically for competitive exam preparation — includes emphasis patterns, important facts, and exam-relevant framing that general educational content lacks - **Long-form audio:** 5,100 hours of continuous structured lecture audio — one of the largest Indic long-form educational audio datasets available - **Subject breadth:** Covers History, Geography, Polity, Economics, Science, and Environment across multiple NCERT grades - **High information density:** Unlike casual educational content, government exam lectures are dense with factual content — valuable for knowledge-intensive LLM training ## Intended Uses - Training automatic speech recognition (ASR) models for Hindi and Indic languages in educational domain - Long-form audio understanding model development - Knowledge-intensive question answering model training - Indian history, geography, and polity domain model fine-tuning - Government exam preparation AI development - Curriculum-structured audio dataset for educational AI research ## Why This Dataset Is Unique Government exam preparation content is structurally different from general educational content. It is optimised for information retention, fact density, and conceptual clarity under exam conditions. No comparable large-scale structured dataset exists for Indian government exam preparation content in audio format. ## Data Collection and Rights All content is proprietary, produced by Prepp's in-house faculty team of government exam subject matter experts. Content is curriculum-mapped to NCERT chapter structure and ethically sourced under work-for-hire agreements. Full dataset licensing is available for commercial AI training purposes. ## Licensing and Commercial Access This repository contains sample data only. The full dataset of 5,100 hours of NCERT lecture recordings is available for commercial AI training licensing. **For licensing inquiries contact:** Ankit Dubey — Head of AI Data Partnerships, Collegedunia ankit.dubey@collegedunia.com ## Dataset Curator [Collegedunia Web Private Limited](https://collegedunia.com) | [Prepp](https://prepp.in) Gurugram, Haryana, India

--- license: 其他 task_categories: - 音频分类 - 自动语音识别 language: - 英语 - 印地语 - 孟加拉语 - 泰米尔语 - 泰卢固语 - 马拉雅拉姆语 - 马拉地语 - 奥里亚语 - 阿萨姆语 - 旁遮普语 tags: - ncert - 教育 - 印度 - UPSC - 人文 - 录制讲座 - 公务员考试 - 科学 - 长音频 pretty_name: 印度NCERT讲座 size_categories: - 1K<n<10K # 印度NCERT讲座合集 ## 数据集描述 本数据集为大规模录制讲座合集,覆盖科学与人文方向的NCERT(印度国家教育研究与培训委员会,National Council of Educational Research and Training)课程体系,专为UPSC、SSC、邦级PSC等公务员考试备考打造。数据集由印度最大公务员考试备考平台Prepp出品,其运营方为Collegedunia Web Private Limited。 ## 数据集摘要 - **总时长**:5100小时录制讲座 - **内容类型**:覆盖NCERT科学与人文方向的结构化课程讲座 - **考试适配性**:适配UPSC公务员考试、SSC、邦级PSC、铁路、银行及所有主流公务员竞争性考试 - **课程范围**:完整覆盖6至12年级科学与人文方向的NCERT课程 - **可获取章节**:历史第1、2、3章(样本) - **语言支持**:以印地语和英语为主要语言,同时提供区域语言版本 - **格式规格**:音视频格式,按章节结构化授课 ## 样本数据 本仓库提供三份样本讲座: - 历史第1章——早期社会 - 历史第2章——早期经济与帝国 - 历史第3章——政治与经济史 ## 核心优势 - **贴合NCERT标准**:严格遵循官方NCERT课程章节结构,对印度教育领域AI模型而言具备高价值参考意义 - **公务员考试优化**:内容专为竞争性考试备考设计,涵盖通用教育内容所不具备的重点标记模式、核心知识点及考试导向的内容架构 - **长音频资源**:5100小时连续结构化讲座音频,为当前可获取的规模最大的印度语族长格式教育音频数据集之一 - **学科覆盖广度**:覆盖多NCERT年级的历史、地理、政治、经济、科学与环境学科 - **高信息密度**:与泛娱乐教育内容不同,公务员考试讲座蕴含大量事实性知识点,对知识密集型大语言模型(LLM,Large Language Model)训练极具价值 ## 预期用途 - 教育领域印地语及印度语族自动语音识别(ASR,Automatic Speech Recognition)模型的训练 - 长音频理解模型研发 - 知识密集型问答模型训练 - 印度历史、地理与政治领域模型的微调 - 公务员考试备考AI研发 - 面向教育AI研究的课程结构化音频数据集 ## 数据集独特性 公务员考试备考内容与通用教育内容在结构上存在显著差异,其针对考试场景下的信息留存、知识点密度及概念清晰度进行了专门优化。目前尚无同规模、同结构的印度公务员考试备考音频格式数据集。 ## 数据采集与版权 所有内容均为专有资产,由Prepp内部的公务员考试学科专家师资团队制作。内容已与NCERT章节结构完成课程映射,并通过雇佣创作协议实现合规获取。完整数据集可授权用于商业AI训练场景。 ## 授权与商业获取 本仓库仅包含样本数据。完整的5100小时NCERT讲座录制数据集可授权用于商业AI训练。 **如需咨询授权事宜,请联系:** Ankit Dubey — Collegedunia AI数据合作主管 ankit.dubey@collegedunia.com ## 数据集运营方 [Collegedunia Web Private Limited](https://collegedunia.com) | [Prepp](https://prepp.in) 印度哈里亚纳邦古鲁格拉姆

提供机构:
DataOrigin
二维码
社区交流群
二维码
科研交流群
商业服务