遇见数据集

shinigamiRaj/4VedasEnglish

收藏
Hugging Face2026-05-27 更新2026-05-31 收录
官方服务:

资源简介:

--- license: other license_name: public-domain license_link: LICENSE language: - en - sa - hi tags: - Vedas - Hindu - History - India - RigVeda - SamaVeda - YajurVeda - AtharvaVeda size_categories: - 10K<n<100K data_prepared_from: - Ralph T.H. Griffith's Rig Veda, Sama Veda, White Yajur Veda, and Atharva Veda English translations - Arthur Berriedale Keith's Black Yajur Veda English translation Chunk_Overlap: - 150 characters Chunk_Size: - Full Hymn / Anuvaka / Verse granularity --- # 📖 Four Vedas Corpus (Rig Veda, Sama Veda, Yajur Veda, Atharva Veda) **Languages:** English (en) **Dataset size:** ~10K < n < ~100K records (chunks and verses combined) --- ## Dataset Summary This corpus is a comprehensive, highly structured dataset comprising the four primary pillars of ancient Indian literature: the **Rig Veda**, **Sama Veda**, **Yajur Veda**, and **Atharva Veda**. It integrates multiple authoritative translations, offering Sanskrit source verses alongside English and Hindi interpretations. The dataset is formatted for continuous pre-training, fine-tuning, and Retrieval-Augmented Generation (RAG) applications in Indology, computational linguistics, historical analysis, and ancient language model development. --- ## 🏛️ Corpus Composition & Authorship The dataset brings together works from highly respected translators and scholars: 1. **Rig Veda**: - **Translator**: **Ralph T.H. Griffith** (1896). - **Language**: English (en). - **Structure**: 10 Mandalas (Books), organized by Suktas (Hymns) and verses. 2. **Sama Veda**: - **Translators**: - **Ralph T.H. Griffith** (1895) - English translation. - **PDF Manuscripts** - High-resolution digital renderings. - **Languages**: English (en) | Sanskrit (sa) | Hindi (hi). - **Structure**: Part I (Mula/Decades) & Part II (Hymns). 3. **Yajur Veda**: - **Translators**: - **Black Yajur Veda** (Taittiriya Samhita): **Arthur Berriedale Keith** (1914) - English. - **White Yajur Veda** (Vajasaneya Samhita): **Ralph T.H. Griffith** (1899) - English. - **Languages**: English (en). - **Structure**: Kandas & Prapathakas (Black) | 40 Books & Verses (White). 4. **Atharva Veda**: - **Translators**: - **Ralph T.H. Griffith** (1895) - English. - **Languages**: English (en). - **Structure**: 20 Books, Suktas, and Verses. --- ## ⚙️ Chunking & Preprocessing Strategies Two specialized pipelines are used to prepare the Vedic data for LLM ingestion: ### 1. Document Page Chunking (PDF pipeline) Used for the bilingual translation documents (e.g., Sharma's Atharva Veda and PDF versions of Sama Veda) to ensure balanced context sizes. | Parameter | Value | Description | |-----------|-------|-------------| | **Chunk Overlap** | 150 characters | Ensures smooth transitions and preserves semantic boundaries | | **Chunk Size** | 4 chunks per page | Optimizes sequence length for embedding and context retrieval | | **Strategy** | Page-level overlapping splits | Preserves adjacent translation pairs (Sanskrit-Hindi) | ### 2. Semantic Structural Scraping (Web pipeline) Used for the web-scraped books to preserve original ritualistic, musical, and narrative divisions: * **Rig Veda & Atharva Veda**: Chunked at the **Hymn (Sukta)** level. This encapsulates full structural themes as single database rows. * **Sama Veda**: Chunked at the **Decade/Hymn** level, aligning with Vedic musical metres. * **Black Yajur Veda**: Chunked at the **Anuvaka** level, maintaining ritual instruction blocks. * **White Yajur Veda**: Chunked at the individual **Verse** level with high-resolution metadata tags. --- ## 📁 Data Structure Dataset records follow a unified JSON schema containing the text payload alongside hierarchical metadata: ### Scraped Structural Record Example (JSONL) ```json { "text": "BLACK YAJUR VEDA (TAITTIRIYA SAMHITA)\nKANDA I\nPRAPATHAKA I\nANUVAKA i. 1. 1.\nThe New and Full Moon Sacrifices\n\na For food thee, for strength thee!\nb Ye are winds, ye are approachers.\nc Let the god Savitr impel you to the most excellent offering...", "source": "sacred-texts.com", "collection": "Yajur Veda", "sub_collection": "Black Yajur Veda", "translator": "Arthur Berriedale Keith", "kanda": 1, "prapathaka": 1, "anuvaka": "i. 1. 1.", "title": "The New and Full Moon Sacrifices" } ``` --- ## 🚀 Purpose & Applications * **Ancient Text Pre-Training**: Ideal for injecting Vedic domain knowledge into deep learning models (e.g., VedaGPT). * **Bilingual RAG (Retrieval-Augmented Generation)**: High-quality translations allow cross-lingual semantic searches between English, Hindi, and Sanskrit. * **Linguistic & Digital Humanities**: Facilitates statistical studies of ancient Sanskrit syntax, meter, and translation structures.

This corpus is a comprehensive, highly structured dataset comprising the four primary pillars of ancient Indian literature: the Rig Veda, Sama Veda, Yajur Veda, and Atharva Veda. It integrates multiple authoritative translations, offering Sanskrit source verses alongside English and Hindi interpretations. The dataset is formatted for continuous pre-training, fine-tuning, and Retrieval-Augmented Generation (RAG) applications in Indology, computational linguistics, historical analysis, and ancient language model development.

提供机构:
shinigamiRaj
搜集汇总
数据集介绍
shinigamiRaj/4VedasEnglish 数据集图片
构建方式
该数据集以四部吠陀本集为核心,整合了多位权威学者的英译与印地语译文,构建过程融合了两种互补策略:针对网络来源文本,采用语义结构爬取,按颂歌、赞歌、阿努瓦卡或诗节等原始仪式与韵律单元进行切分,保留文本固有的宗教与音乐层级;对于双语PDF文献,则实施页面级重叠分块,设定150字符重叠以维持语义边界与翻译对完整性。所有记录均映射至统一的JSON架构,携带译者、卷次、章节等层级元数据,从而在尊重吠陀文献原生结构的同时,满足大语言模型对上下文连续性与检索粒度的需求。
特点
该数据集在内容与结构上均体现出高度的学术严谨性。它涵盖《梨俱吠陀》《娑摩吠陀》《夜柔吠陀》与《阿闼婆吠陀》四部根本典籍,包含英语、梵语及印地语多语种文本,来源涉及Griffith与Keith等经典译本。数据规模介于一万至十万条记录之间,以颂歌、诗节或阿努瓦卡为基本单元,并配套完整的译者信息、篇章编号与标题等元数据。这种细粒度的结构化设计既保留了吠陀文献的仪式性与音乐性划分,又通过重叠分块缓解了文本割裂,为跨语言检索与历史语言学研究提供了高质量语料。
使用方法
该数据集适用于吠陀领域知识预训练、微调及检索增强生成等任务。研究者可借助其层级元数据,按特定吠陀、卷次或译者进行筛选与切片,以开展古梵语语法、诗律及翻译结构的计量分析。在双语RAG场景中,英、梵、印地语并行文本支持跨语言语义检索,便于构建问答系统或数字人文工具。使用时应遵循公共领域许可,引用时注明原始译者与数据来源。加载时可直接解析JSONL记录,利用text字段作为输入,collection、kanda等字段作为过滤或条件标签,实现高效的数据管道集成。
背景与挑战
背景概述
四吠陀作为古印度文明的核心经典,其文本的数字化与结构化长期面临多语种、多译本与多层级编排的复杂性问题。4VedasEnglish数据集由研究人员整合Ralph T.H. Griffith与Arthur Berriedale Keith等权威英译本,于2020年代构建,涵盖《梨俱吠陀》《娑摩吠陀》《夜柔吠陀》及《阿闼婆吠陀》的梵文原典与英印译文。该数据集旨在支撑古印度学、计算语言学及历史分析中的预训练、微调与检索增强生成任务,为古语言模型注入吠陀领域知识,推动跨语种语义检索与数字人文研究。
当前挑战
该数据集所应对的领域问题在于古印度文献的语义结构保存与跨语言对齐:吠陀文本具有仪式性、韵律性与多层级组织(如曼陀罗、赞歌、诗节),传统数字化方法易破坏其内在边界。构建过程中的挑战包括多源译本版权与转写差异的整合、梵文天城体与拉丁转写的准确映射、PDF文档与网页抓取在分块策略上的协调,以及150字符重叠与语义单元粒度间的平衡。此外,保持译文与原文的对应关系、避免分块截断语义,亦对预处理流程提出严苛要求。
常用场景
经典使用场景
在古印度宗教文献的计算语言学研究领域,四吠陀数据集凭借其多语言对齐与结构化分层特性,成为预训练古代文本语言模型的经典语料。该数据集最典型的使用场景在于:以《梨俱吠陀》《娑摩吠陀》《夜柔吠陀》及《阿闼婆吠陀》的英译与梵文原典为输入,通过完整颂诗、阿努瓦卡及诗节粒度的分块策略,驱动掩码语言建模与因果语言建模任务,从而令模型习得吠陀梵语的句法韵律与仪式语境下的语义表征。
实际应用
在实际应用层面,该数据集支撑了面向吠陀文献的检索增强生成系统,使开发者能够构建跨英语、印地语与梵语的语义搜索引擎,服务于学术研究与文化教育。同时,其结构化元数据为仪式指令的自动化抽取、颂诗韵律的统计建模以及古代语言教学工具的研发提供了数据基础。博物馆与数字档案馆亦可借助该语料实现吠陀手稿的语义标注与智能问答。
衍生相关工作
围绕该数据集,已衍生出若干经典工作,包括以吠陀领域知识注入为导向的预训练模型VedaGPT,以及基于双语对齐的跨语言检索基准。此外,部分研究利用其分层结构开展梵语诗律的计量分析,并构建了面向《夜柔吠陀》仪式文本的细粒度分类器。这些工作共同拓展了古代文献在深度学习与数字人文中的应用边界。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务