indro-web-data
收藏资源简介:
Indro-Veda数据集是一个经过整理的大规模结构化网络数据存储库,属于Indro-Veda AI研究计划的一部分,旨在开发具有语言多样性和网络规模索引的下一代信息检索系统和大型语言模型。数据集体积约为400GB(实时增长),主要来源为公共领域存档Common Crawl,数据格式为JSONLines(.jsonl),优化用于分布式训练。处理引擎为Indro Titan V10(分布式工业分片)。该数据集适用于索引优化(研究低延迟检索)、语言建模(构建多语言AI训练的强大语料库)和数据工程(研究P2P浏览器数据摄取的效率)。数据集遵循Apache License 2.0许可,所有数据摄取均遵循robots.txt协议和Common Crawl的使用条款,严格用于学术研究和非商业AI开发。
The Indro-Veda Dataset is a curated large-scale structured web data repository, part of the Indro-Veda AI Research Program. It aims to develop next-generation information retrieval systems and large language models (LLMs) with linguistic diversity and web-scale indexing. The dataset has a total size of approximately 400 GB and is growing in real time. Its primary data source is the public-domain archive Common Crawl, with the data format being JSONLines (.jsonl), optimized for distributed training. The processing engine employed is Indro Titan V10 (distributed industrial sharding). This dataset supports three main research directions: index optimization (research on low-latency retrieval), language modeling (building robust corpora for multilingual AI training), and data engineering (researching the efficiency of P2P browser data ingestion). The dataset is licensed under Apache License 2.0. All data ingestion operations comply with the robots.txt protocol and the terms of use of Common Crawl, and it is strictly restricted to academic research and non-commercial AI development.
Indro-web-data 数据集概述
数据集基本信息
- 名称: Indro-web-data
- 别名: Indro-Veda Dataset
- 许可证: Apache License 2.0
- 任务类别: 文本生成、填充掩码
- 语言: 英语、印地语
- 标签: 网络爬取、indro-veda、大规模、研究
- 规模类别: 100B < n < 1T
数据集内容与架构
- 项目动机: 作为 Indro-Veda AI 研究计划的一部分,旨在开发专注于语言多样性和网络规模索引的下一代信息检索系统和大型语言模型。
- 数据来源: 主要源自 Common Crawl(公共领域存档)。
- 数据格式: JSONLines (.jsonl),针对分布式训练进行了优化。
- 处理引擎: Indro Titan V10(分布式工业分片)。
- 数据量: 约 400 GB(实时增长的语料库)。
研究目标与用途
- 索引优化: 研究 Indro 搜索引擎的低延迟检索。
- 语言建模: 为多语言 AI 训练构建强大的语料库。
- 数据工程: 研究 P2P 浏览器数据摄取的效率。
- 使用范围: 严格用于学术研究和非商业性 AI 开发。
合规性与伦理
- 来源完整性: 所有数据摄取均遵循
robots.txt协议和 Common Crawl 的使用条款。 - 机构领导: Indro AI Research Lab
- 状态: 处于主动摄取阶段




