semanta-dataset-suite
收藏资源简介:
Semanta数据集套件是一个面向多个行业垂直领域的大规模合成表格数据集,由Semanta“世界智能操作系统”生成,旨在证明其能够为广泛行业生成结构化、场景覆盖的合成世界。数据集覆盖11个行业领域:银行、人工智能、对冲基金、保险、制药、医疗、制造、安全、科学、快消品和零售。每个行业数据集包含100,000行和100列,总数据规模为1,100,000行×100列,共计1.1亿个单元格。所有数据均为完全合成生成,未使用任何客户数据或外部模型提供商数据。数据集整体质量得分为0.965,场景覆盖得分为1.000,每个行业都有独立的质量评估指标(范围0.946-0.981)。该数据集适用于机器学习、深度学习、神经网络和量化实验,特别设计用于场景鲁棒性测试、漂移原型设计以及模型训练前的验证。通过提供行业特定的模拟场景(如信用风险、交易异常、流动性模拟、提示复杂性评估、因子制度变化、索赔频率、试验信号、患者流、缺陷检测、攻击强度、假设强度、促销弹性、客户流失等),使组织能够在历史证据不足或现实成本过高的情况下,提前测试和训练其模型系统。数据集配套提供机器可读的列合约和语义类型模式文件、统计质量和覆盖度指标文件、可重复性证据谱系文件以及发布清单文件。
The Semanta dataset suite is a large-scale synthetic tabular dataset for multiple industry verticals. It is generated by Semantas World Intelligent Operating System to demonstrate its ability to produce structured, scenario-covered synthetic worlds across a wide range of industries, rather than a single demo dataset. The dataset includes 11 industry domains: banking, AI, hedge funds, insurance, pharmaceuticals, healthcare, manufacturing, security, science, FMCG, and retail. Each industry dataset contains 100,000 rows and 100 columns, with a total data scale of 1,100,000 rows × 100 columns, amounting to 110 million cells. All data is fully synthetic, with no use of customer data or external model provider data. The overall dataset quality score is 0.965, and the scenario coverage score is 1.000, with independent quality assessment metrics for each industry (ranging from 0.946 to 0.981). The dataset is suitable for machine learning, deep learning, neural networks, and quantitative experiments, specifically designed for scenario robustness testing, drift prototyping, and pre-training model validation. By providing industry-specific simulation scenarios (such as credit risk, transaction anomalies, liquidity simulations, prompt complexity assessment, factor regime changes, claim frequency, trial signals, patient flow, defect detection, attack intensity, hypothesis strength, promotion elasticity, customer churn, etc.), it enables organizations to test and train their model systems in advance when historical evidence is insufficient or real-world costs are prohibitive. The dataset is accompanied by machine-readable column contracts and semantic type schema files, statistical quality and coverage metric files, reproducibility evidence lineage files, and release manifest files.
Semanta Dataset Suite 数据集概述
基本信息
- 数据集名称: Semanta Dataset Suite - Industry Vertical Synthetic Worlds
- 许可证: MIT
- 语言: 英语
- 数据集大小: 1M < n < 10M
- 任务类别: 表格分类、表格回归、时间序列预测
- 标签: 合成数据、世界原生数据、表格数据、银行、人工智能、对冲基金、保险、制药、医疗、制造、安全、科学、快消品、零售
数据集规模
- 仓库中发布的合成数据总行数:2,370,621
- 行业垂直套件行数:1,100,000
- 涵盖行业数:11
- 每个行业列数:100
- 总单元格数:110,000,000
- 综合质量评分:0.965
- 场景覆盖评分:1.000
数据来源与特性
- 纯合成数据,不使用任何客户数据
- 不涉及外部模型提供商
- 为机器/深度学习、神经网络、量化实验、场景鲁棒性、漂移原型及StarForge/Gamma训练交接设计
可下载内容
| 内容 | 说明 |
|---|---|
data/industry_vertical_suite/<行业>/*.csv.gz |
按行业划分的训练就绪合成表格数据世界 |
schemas/industry_vertical_suite/*.json |
机器可读的字段合约与语义类型 |
metrics/industry_vertical_suite/*_quality_metrics.json |
各垂直领域统计质量与覆盖度指标 |
metrics/industry_vertical_suite/suite_scorecard.json |
执行层面的套件就绪度与质量评分卡 |
semanta/industry_vertical_suite/lineage.json |
可复现性与源策略证据 |
semanta/industry_vertical_suite/publication_manifest.json |
发布清单、文件及声明边界 |
行业数据集详情
1. 银行 (Banks)
- 行数: 100,000 | 列数: 100 | 质量评分: 0.968
- 主要用例: 信用风险、交易异常、流动性与操作风险模拟
- 数据文件:
data/industry_vertical_suite/banks/semanta_banks_world_native_synthetic_100000x100.csv.gz
2. 人工智能 (AI)
- 行数: 100,000 | 列数: 100 | 质量评分: 0.973
- 主要用例: 提示复杂度、评估难度、幻觉风险与工具使用场景生成
- 数据文件:
data/industry_vertical_suite/ai/semanta_ai_world_native_synthetic_100000x100.csv.gz
3. 对冲基金 (Hedge Funds)
- 行数: 100,000 | 列数: 100 | 质量评分: 0.978
- 主要用例: 因子体制、回撤、流动性与阿尔法衰减场景生成
- 数据文件:
data/industry_vertical_suite/hedge_funds/semanta_hedge_funds_world_native_synthetic_100000x100.csv.gz
4. 保险 (Insurance)
- 行数: 100,000 | 列数: 100 | 质量评分: 0.977
- 主要用例: 理赔频率、损失严重性、准备金充足率与巨灾尾部模拟
- 数据文件:
data/industry_vertical_suite/insurance/semanta_insurance_world_native_synthetic_100000x100.csv.gz
5. 制药 (Pharma)
- 行数: 100,000 | 列数: 100 | 质量评分: 0.961
- 主要用例: 试验信号、安全性、监管延迟与上市准备场景生成
- 数据文件:
data/industry_vertical_suite/pharma/semanta_pharma_world_native_synthetic_100000x100.csv.gz
6. 医疗 (Medicine)
- 行数: 100,000 | 列数: 100 | 质量评分: 0.967
- 主要用例: 患者流动、诊断不确定性、治疗反应与容量风险模拟
- 数据文件:
data/industry_vertical_suite/medicine/semanta_medicine_world_native_synthetic_100000x100.csv.gz
7. 制造 (Manufacturing)
- 行数: 100,000 | 列数: 100 | 质量评分: 0.957
- 主要用例: 缺陷、停机、传感器漂移、吞吐量与供应商风险模拟
- 数据文件:
data/industry_vertical_suite/manufacturing/semanta_manufacturing_world_native_synthetic_100000x100.csv.gz
8. 安全 (Security)
- 行数: 100,000 | 列数: 100 | 质量评分: 0.981
- 主要用例: 攻击强度、告警噪声、响应延迟与资产关键性模拟
- 数据文件:
data/industry_vertical_suite/security/semanta_security_world_native_synthetic_100000x100.csv.gz
9. 科学 (Science)
- 行数: 100,000 | 列数: 100 | 质量评分: 0.946
- 主要用例: 假设强度、实验噪声、复现性与发现潜力模拟
- 数据文件:
data/industry_vertical_suite/science/semanta_science_world_native_synthetic_100000x100.csv.gz
10. 快消品 (FMCG)
- 行数: 100,000 | 列数: 100 | 质量评分: 0.952
- 主要用例: 促销、需求弹性、缺货、渠道转移与定价场景生成
- 数据文件:
data/industry_vertical_suite/fmcg/semanta_fmcg_world_native_synthetic_100000x100.csv.gz
11. 零售 (Retail)
- 行数: 100,000 | 列数: 100 | 质量评分: 0.957
- 主要用例: 购物篮、流失、客流、降价、库存与客户行为模拟
- 数据文件:
data/industry_vertical_suite/retail/semanta_retail_world_native_synthetic_100000x100.csv.gz
运行凭证
- 运行ID:
industry-suite-20260618-01 - 规范版本:
v3.7 Final - 生产烟雾门: 13/13 公共证明目标已通过




