om-ashish-soni/vivechan-spritual-text-dataset-v3
收藏资源简介:
--- dataset_info: features: - name: text dtype: string splits: - name: train num_bytes: 23671713 num_examples: 80780 download_size: 12397166 dataset_size: 23671713 configs: - config_name: default data_files: - split: train path: data/train-* license: apache-2.0 task_categories: - text-generation - text2text-generation - question-answering - text-retrieval language: - en size_categories: - 10K<n<100K --- # Vivechan - Spiritual Text Dataset ## Description The Vivechan - Spiritual Text Dataset is an open and public collection of textual data extracted from significant spiritual texts, curated to support discussions, inquiries, doubts, and Q&A sessions within the realm of spirituality. This dataset provides valuable content from the following revered sources: - Shrimad Bhagwat Mahapurana - Shripad Shri Vallabha Charitramrutam - Shiv Mahapurana Sankshipt - Valmiki Ramayan - Vachanamrutam - Shikshapatri - Shree Sai Charitra - Devi Mahatmaya (Chandipath) - Eknathi Bhagwat - Shri Dattapurana - Shri Gurucharitra - Shrimad Bhagwad Gita - Bhagwad Gita ## Dataset Information - **Features**: - **text**: Each example consists of a string containing textual excerpts from the mentioned sources. - **Splits**: - **Train**: 80,780 examples - **Download Size**: 12,397,166 bytes - **Dataset Size**: 23,671,713 bytes ## Task Categories The dataset is designed to facilitate the following tasks: - **Text Retrieval**: Retrieve relevant passages based on user queries or specified topics. - **Text-to-Text Generation**: Generate responses or elaborate on queries based on input text. - **Text-to-Speech**: Convert textual data into speech for auditory presentation. ## Usage This dataset, Vivechan - Spiritual Text Dataset, is openly available and can be utilized to train or fine-tune Language Models (LLMs), existing AI models, or develop new models for various applications within the realm of spirituality and spiritual texts. ## Language The dataset is available in English (en). ## Size Categories The dataset falls within the size category of 10K < n < 100K, making it suitable for training or fine-tuning LLMs and other AI models. ## License This dataset is released under the Apache License 2.0, enabling open usage, modification, and distribution. ## Citation If you use this dataset in your work, please cite it as: [Insert citation details here] ## Acknowledgements We express our gratitude to the original sources of the texts included in this dataset: - Shrimad Bhagwat Mahapurana - Shripad Shri Vallabha Charitramrutam - Shiv Mahapurana Sankshipt - Valmiki Ramayan - Vachanamrutam - Shikshapatri - Shree Sai Charitra - Devi Mahatmaya (Chandipath) - Eknathi Bhagwat - Shri Dattapurana - Shri Gurucharitra - Shrimad Bhagwad Gita - Bhagwad Gita
数据集信息: 特征: - 名称:text 数据类型:字符串 数据集划分: - 名称:训练集(train) 字节数:23671713 样本数:80780 下载大小:12397166 数据集总大小:23671713 配置项: - 配置名称:default 数据文件: - 划分:训练集 路径:data/train-* 许可证:Apache-2.0 任务类别: - 文本生成 - 文本到文本生成 - 问答 - 文本检索 语言: - 英语(en) 规模类别: - 10K<n<100K # Vivechan——灵修文本数据集(Vivechan - Spiritual Text Dataset) ## 描述 Vivechan——灵修文本数据集是一套开放公开的文本数据集,提取自多部重要的灵修典籍,经整理后可用于灵修领域的讨论、咨询、答疑及问答活动。本数据集收录了以下经典灵修文本中的宝贵内容: - Shrimad Bhagwat Mahapurana - Shripad Shri Vallabha Charitramrutam - Shiv Mahapurana Sankshipt - Valmiki Ramayan - Vachanamrutam - Shikshapatri - Shree Sai Charitra - Devi Mahatmaya (Chandipath) - Eknathi Bhagwat - Shri Dattapurana - Shri Gurucharitra - Shrimad Bhagwad Gita - Bhagwad Gita ## 数据集信息 - **特征**: - **text**:每条样本均为包含上述来源文本节选的字符串。 - **划分**: - **训练集**:共80780条样本 - 下载大小:12397166字节 - 数据集总大小:23671713字节 ## 任务类别 本数据集旨在支持以下任务类型: - **文本检索**:基于用户查询或指定主题检索相关段落。 - **文本到文本生成**:基于输入文本生成回应内容,或对查询进行详细阐述。 - **文本到语音**:将文本数据转换为语音以实现听觉呈现。 ## 使用方式 本数据集Vivechan——灵修文本数据集可公开获取,可用于训练或微调大语言模型(Large Language Model,简称LLM)、现有AI模型,或针对灵修及灵修文本领域的各类应用开发全新模型。 ## 语言 本数据集采用英语(en)编写。 ## 规模类别 本数据集规模处于10K < n < 100K区间,适配大语言模型及其他AI模型的训练与微调需求。 ## 许可证 本数据集采用Apache许可证2.0(Apache License 2.0)发布,允许公开使用、修改与分发。 ## 引用 若您在研究工作中使用本数据集,请按以下格式引用: [此处插入引用详情] ## 致谢 我们谨向本数据集收录文本的原始来源致以诚挚谢意: - Shrimad Bhagwat Mahapurana - Shripad Shri Vallabha Charitramrutam - Shiv Mahapurana Sankshipt - Valmiki Ramayan - Vachanamrutam - Shikshapatri - Shree Sai Charitra - Devi Mahatmaya (Chandipath) - Eknathi Bhagwat - Shri Dattapurana - Shri Gurucharitra - Shrimad Bhagwad Gita - Bhagwad Gita
数据集概述
基本信息
- 名称: Vivechan - Spiritual Text Dataset
- 语言: 英语 (en)
- 许可证: Apache License 2.0
- 大小: 10K < n < 100K
数据集特征
- 特征:
- text: 字符串类型,包含来自多个宗教文本的摘录。
数据集分割
- 训练集: 80,780个示例,大小为23,671,713字节。
下载与数据集大小
- 下载大小: 12,397,166字节
- 数据集大小: 23,671,713字节
任务类别
- 文本检索: 根据用户查询或指定主题检索相关段落。
- 文本到文本生成: 根据输入文本生成响应或扩展查询。
- 文本到语音: 将文本数据转换为语音进行听觉展示。
使用场景
- 用于训练或微调语言模型(LLMs),现有AI模型,或开发新的模型,适用于宗教和宗教文本领域的各种应用。




