agri-slm-corpus-3
收藏资源简介:
Agriculture SLM Corpus v3 是一个专为印度农业领域设计的预训练语料库,旨在支持3亿参数规模的特定语言模型(SLM)训练。数据集包含两个核心文件:corpus_india_downsampled.jsonl 汇集了印度相关的多源文本数据,包括来自PubMed、EuropePMC、Wikipedia、扩展语料和印度农业研究委员会(ICAR)等的内容,涵盖约20.1万篇文档,总计约12.44亿词元;krishikosh.jsonl 则专门收录ICAR KrishiKosh平台上的学术资源,包括学位论文、研究论文、书籍和报告,涵盖约5.5万篇文档,总计约13.85亿词元。数据以JSONL格式组织,每个样本包含文本内容、来源(如krishikosh)、领域(固定为agriculture)、子领域(如theses、research_papers、books)、语言(固定为英语)、词数统计、原始URL和标题等结构化字段。该语料库适用于农业领域的自然语言处理预训练任务,特别聚焦于印度农业背景的知识建模与应用开发。
Agriculture SLM Corpus v3 is a pre-training corpus specifically designed for the Indian agricultural domain, aimed at supporting the training of specific language models (SLMs) with 300 million parameters. The dataset consists of two core files: corpus_india_downsampled.jsonl aggregates multi-source text data related to India, including content from PubMed, EuropePMC, Wikipedia, extended corpora, and the Indian Council of Agricultural Research (ICAR), covering approximately 201,000 documents with a total of about 1.244 billion tokens; krishikosh.jsonl specifically collects academic resources from the ICAR KrishiKosh platform, including theses, research papers, books, and reports, covering around 55,000 documents with a total of about 1.385 billion tokens. The data is organized in JSONL format, with each sample containing structured fields such as text content, source (e.g., krishikosh), domain (fixed as agriculture), subdomain (e.g., theses, research_papers, books), language (fixed as English), token count statistics, original URL, and title. This corpus is suitable for natural language processing pre-training tasks in the agricultural field, with a particular focus on knowledge modeling and application development in the context of Indian agriculture.
数据集概述
数据集名称:Agriculture SLM Corpus v3
计划用途:为印度农业领域一个3亿参数的小语言模型(SLM)提供预训练语料
许可协议:CC-BY-4.0
语言:英语
标签:农业、印度、耕作、自然语言处理、预训练
数据集规模:1M < 样本数 < 10M
数据集内容
数据集包含两个主要文件:
| 文件名 | 描述 | 文档数 | Token数 |
|---|---|---|---|
| corpus_india_downsampled.jsonl | 以印度为重点的下采样基础语料库,涵盖PubMed、EuropePMC、Wikipedia、扩展资料、ICAR等来源 | 约201K | 约1,244M |
| krishikosh.jsonl | ICAR KrishiKosh资料(论文、文章、书籍、报告),通过DSpace 7采集 | 约55K | 约1,385M |
数据模式
每条数据记录包含以下字段:
- text:文本内容
- source:数据来源(如 krishikosh 或其他)
- domain:领域(固定为 agriculture)
- subdomain:子领域(如 theses, research_papers, books 等)
- language:语言(固定为 en)
- word_count:词数
- url:来源链接
- title:标题





