Dragondata-Finance-Corpus
收藏资源简介:
Dragondata Finance Corpus(DD-CORPUS-001)是一个由Dragondata构建的开源金融语料库,适用于大型语言模型预训练和金融AI。数据来源于100%免费、许可宽松的公开来源(如公共领域、CC BY 4.0、MIT、开放获取),经过策划、去重、质量过滤和分词处理,格式专业且标准化。覆盖25个金融相关领域,包括SEC EDGAR文件、美联储、IMF与世界银行、OECD与国际组织、WTO与贸易、金融新闻、股票市场数据、财报电话会议、加密货币、外汇与商品、固定收益与债券、衍生品与期权、另类数据与ESG、美国银行监管、欧盟金融监管、税法与会计、保险与精算、金融研究论文、经济学工作论文、央行研究、个人理财、公司金融、量化金融、房地产金融等。估计总token数约为313B,每个分片包含134M个token,以JSONL格式存储,每行包含tokens(整数列表)和num_tokens(整数)字段。分词器采用bigcode/starcoder2-3b(词汇表大小49,152)。适用于文本生成和填充掩码任务,采用Apache-2.0许可证。
Dragondata Finance Corpus (DD-CORPUS-001) is an open-source financial corpus built by Dragondata, designed for large language model pretraining and financial AI. The data is compiled from 100% free, permissively licensed public sources (e.g., public domain, CC BY 4.0, MIT, open access), and has been curated, deduplicated, quality-filtered, and tokenized into a professional and standardized format. It covers 25 financial domains including SEC EDGAR filings, Federal Reserve, IMF & World Bank, OECD & International Organizations, WTO & Trade, financial news, stock market data, earnings calls, cryptocurrency, forex & commodities, fixed income & bonds, derivatives & options, alternative data & ESG, US banking regulation, EU financial regulation, tax law & accounting, insurance & actuarial, financial research papers, economics working papers, central bank research, personal finance, corporate finance, quantitative finance, real estate finance, etc. The estimated total number of tokens is approximately 313B, with each shard containing 134M tokens stored in JSONL format. Each line contains tokens (list of integers) and num_tokens (integer). The tokenizer is bigcode/starcoder2-3b (vocabulary size 49,152). It is suitable for tasks such as text generation and fill-mask, and is released under the Apache-2.0 license.
DragonData Finance Corpus 数据集概述
基本信息
| 属性 | 详情 |
|---|---|
| 数据集名称 | DragonData Finance Corpus (DD-CORPUS-001) |
| 所有者 | Dragon Limited |
| 品牌 | DragonData (DD) — Datasets Source Center |
| Hugging Face 组织 | dragonlimited |
| 许可证 | Apache-2.0 |
| 语言 | 英语 |
| 任务类型 | 文本生成、掩码填充 |
| 标签 | 金融、语料库、预训练、LLM、金融AI、SEC EDGAR |
| 最后更新 | 2026-08-30 |
数据集简介
这是一个精心策划、已分词的多领域金融语料库,专门用于金融语言模型的预训练。该语料库完全由100%免费、许可宽松的公开来源编译而成。Dragon Limited 拥有数据的策展权——包括搜索、选择、存储、处理和分词,将原始公开数据转化为专业、标准化、可直接使用的语料库,并面向公众开放用于训练 AI 大语言模型等用途。
核心特征
- 25个精选领域:涵盖证券、宏观、监管、市场和量化金融
- 数据规模:预计总计 3130亿 tokens,所有领域均采用 134M-token 分片
- 数据格式:JSONL 格式,每行包含
{"tokens": [...], "num_tokens": N} - 分词器:
bigcode/starcoder2-3b(词汇量 49,152) - 许可证政策:每个来源均验证为宽松许可(公共领域 / CC BY / MIT / 开放获取)
- 溯源:完整的按领域来源文档;已去重和质量过滤
25个领域详情
| # | 领域 | 预计Tokens | 许可证 | 状态 |
|---|---|---|---|---|
| 01 | SEC EDGAR 文件 | 770亿 | 公共领域 | 🔄 构建中 |
| 02 | 美联储 | 110亿 | 公共领域 | 🔄 构建中 |
| 03 | IMF 与世界银行 | 170亿 | CC BY 4.0 | ⏳ 待处理 |
| 04 | OECD 与国际组织 | 150亿 | CC BY 4.0 | ⏳ 待处理 |
| 05 | WTO 与贸易 | 50亿 | 开放获取 | ⏳ 待处理 |
| 06 | 金融新闻 | 670亿 | 混合 | ⏳ 待处理 |
| 07 | 股票市场数据 | 370亿 | MIT / CC | 🔄 构建中 |
| 08 | 财报电话会议 | 80亿 | MIT / 开放 | ⏳ 待处理 |
| 09 | 金融语料库 | 740亿 | 混合 | ⏳ 待处理 |
| 10 | 加密货币 | 50亿 | CC BY-SA 4.0 | ⏳ 待处理 |
| 11 | 外汇与大宗商品 | 50亿 | 公共领域 | ⏳ 待处理 |
| 12 | 固定收益与债券 | 50亿 | 公共领域 | ⏳ 待处理 |
| 13 | 衍生品与期权 | 30亿 | 混合开放 | ⏳ 待处理 |
| 14 | 另类与ESG | 40亿 | 混合开放 | ⏳ 待处理 |
| 15 | 美国银行监管 | 50亿 | 公共领域 | ⏳ 待处理 |
| 16 | 欧盟金融监管 | 50亿 | CC BY 4.0 | ⏳ 待处理 |
| 17 | 税法与会计 | 40亿 | 公共领域 | ⏳ 待处理 |
| 18 | 保险与精算 | 30亿 | 公共领域 | ⏳ 待处理 |
| 19 | 金融研究论文 | 50亿 | CC BY 4.0 | ⏳ 待处理 |
| 20 | 经济学工作论文 | 50亿 | CC BY 4.0 | ⏳ 待处理 |
| 21 | 央行研究 | 50亿 | CC BY / 开放 | ⏳ 待处理 |
| 22 | 个人理财 | 40亿 | 合理使用 / CC | ⏳ 待处理 |
| 23 | 公司金融 | 50亿 | 公共领域 | ⏳ 待处理 |
| 24 | 量化金融 | 30亿 | CC BY / arXiv | ⏳ 待处理 |
| 25 | 房地产金融 | 30亿 | 公共领域 | ⏳ 待处理 |
仓库结构
DragonData-Finance-Corpus/ ├── README.md ← 数据集卡片 ├── LICENSE.md ← 许可条款 ├── DATASET_CARD.md ← 详细数据集文档 ├── DOMAINS.json ← 机器可读的25领域分类 ├── schema.json ← token/分片 schema 规范 ├── manifest.json ← 实时分片注册表 └── domain_XX_<slug>/ ← 每个领域一个文件夹 └── train-00000.jsonl ← 134M-token 分片
文件名约定:{split}-{index:05d}.{ext},索引连续无间隙。
使用示例
分片为 JSONL 格式,每行包含 {"tokens": [int, ...], "num_tokens": N}。使用 bigcode/starcoder2-3b 分词器解码:
python import json from transformers import AutoTokenizer
tok = AutoTokenizer.from_pretrained("bigcode/starcoder2-3b", trust_remote_code=True) with open("domain_01_sec_edgar/train-00000.jsonl") as f: shard = json.loads(f.readline()) text = tok.decode(shard["tokens"])
许可与伦理
- 所有来源均为公共领域或宽松许可(CC BY 4.0、MIT、开放获取)
- 不包含受版权保护的门控或仅限非商业使用的来源
- 每个领域保留完整的许可归属
- 数据集根据 LICENSE.md 中的条款供研究和公开使用




