compliments-reference-db
收藏资源简介:
Compliments Reference DB 是一个结构化的参考数据库,基于确定性、完全可追溯的管道构建,用于处理原始 Compliments 产品数据。该数据集包含 4,440 个产品,经过6个阶段的处理:数据质量基础、语义标准化、产品领域与分类及分组、变体分配、营养集成、以及 Nutri-Score 2023 与 Agribalyse 代理评分。最终输出包括产品分组映射(3,691 个产品组)、变体分配(4,299 个变体)、营养数据(每100克营养信息、清洁营养数据、产品-营养映射)和评分结果(产品分数、Agribalyse 映射、评分排除项)。设计原则强调无LLM、单一事实源、完全可追溯(通过 external_id 跨阶段链接)、无数据丢失、食品与非食品领域分离。数据集适用于产品数据标准化、营养评分计算、变体分析、产品分组研究等任务,可用于学术或研究用途。
Compliments Reference DB is a structured reference database built on a deterministic and fully traceable pipeline for processing raw Compliments product data. This dataset contains 4,440 products that have undergone six processing stages: data quality foundation, semantic standardization, product domain, classification and grouping, variant assignment, nutrition integration, and Nutri-Score 2023 and Agribalyse proxy scoring. The final outputs include product grouping mappings (3,691 product groups), variant assignments (4,299 variants), nutrition data (nutrition information per 100 grams, clean nutrition data, product-nutrition mappings) and scoring results (product scores, Agribalyse mappings, scoring exclusions). Its design principles emphasize no LLMs, single source of truth, full traceability (cross-stage linking via external_id), no data loss, and separation of food and non-food domains. This dataset is suitable for tasks such as product data standardization, nutrition score calculation, variant analysis, product grouping research, and other tasks, and can be used for academic or research purposes.
Compliments Reference DB 数据集详情
数据集概述
Compliments Reference DB 是一个基于 Sobeys Inc. 旗下 Compliments 和 Sensations 自有品牌构建的干净、可复现的产品参考数据库。该数据集通过 6 个确定性的处理阶段构建,完全不使用大语言模型(LLM),保证了全程可追溯性。
- 语言: 英语
- 许可证: MIT
- 权威数据来源: saraNour/compliments-brand
- 发布地址: saraNour/compliments-reference-db(公开)
管道概览(6 个阶段)
| 阶段 | 名称 | 输入数据 | 输出数据 | 关键指标 |
|---|---|---|---|---|
| 1 | 数据质量基础 | products.parquet (4,440 × 16) | phase1_output.parquet (4,440 × 14) | 100% 空值列被移除 |
| 2 | 语义归一化 | phase1_output (4,440 × 14) | phase2_output.parquet (4,440 × 34) | 新增 20 列 |
| 3 | 产品领域 + 分类 + 分组 | phase2_output (4,440 × 34) | reference_product_catalog.csv (3,691 组) | 19 个分类类别 |
| 4 | 变体分配 | phase3 + phase2 映射 | product_variant_mapping.parquet (4,440 × 11) | 4,299 个变体 |
| 5 | 营养数据整合 | nutrition.parquet + phase3 + phase4 | nutrition_per_100g.parquet | 100% 匹配率 |
| 6 | Nutri-Score + Agribalyse | phase5 输出 | phase6_product_scores.parquet (4,440) | 364 个获得评分 (8.2%) |
数据来源
权威输入数据
| 文件 | 行 × 列 | 来源 |
|---|---|---|
| products.parquet | 4,440 × 16 | saraNour/compliments-brand |
| nutrition.parquet | 4,440 × 21 | saraNour/compliments-brand |
外部参考资源
- Google 产品分类体系: 用于 GPC 映射
- Nutri-Score 2023 算法: 基于 Eurofins 引用的 Santé Publique France FAQ(2023年12月21日版)
- Agribalyse v3.2: CIQUAL 数据库,用于环境类别映射
关键处理细节
阶段 1 — 数据质量基础
- 数据清理:空白归一化、空字符串转 null、品牌和标题清洗
- UPC 验证:格式检查、重复检测、复用 UPC 分析
- 移除 100% 空值列(
size_per_unit、size_total)
阶段 2 — 语义归一化
- 品牌归一化:9 个原始品牌变体 → 2 个品牌 + 7 条产品线
- 提取 9 个布尔身份标志:有机、无麸质、无糖、无盐、无乳糖、无花生、植物基、低钠等
- 生成确定性身份哈希(identity_hash)
阶段 3 — 产品分类与分组
- 食品/非食品分类: food 2,814 (63.4%)、unknown 1,232 (27.7%)、non_food 394 (8.9%)
- 19 个分类类别: 涵盖乳制品、肉类海鲜、饮料、烘焙、冷冻、农产品、调味品、糖果、零食、意面大米、早餐、罐头、家居清洁、健康保健、个人护理、婴儿护理、宠物食品等
- 产品分组: 3,691 个唯一分组,303 个模糊案例
阶段 4 — 变体分配
- 基于包装规格/属性构建变体键(如
amt500.0|unitg、count20、nosize) - 生成 4,299 个唯一变体
阶段 5 — 营养数据整合
- 归一化方法:直接 100g(150 个)、按食用份量缩放(1,351 个)、未归一化(2,939 个)
- 营养质量状态:VALID 4,076 个、MISSING 364 个、SUSPICIOUS 0 个
- 归一化至每 100g/mL 的营养值(17 个营养字段)
阶段 6 — Nutri-Score 2023 和 Agribalyse 映射
- 符合评分条件的产品: 364 个(8.2%),需同时满足食品领域、营养数据有效、6 个必需字段完整
- Nutri-Score 等级分布: A 级 31 个 (8.5%)、B 级 42 个 (11.5%)、C 级 135 个 (37.1%)、D 级 83 个 (22.8%)、E 级 73 个 (20.1%)
- Agribalyse 映射置信度: HIGH 1,908 个、MEDIUM 885 个、LOW 1,310 个、NONE 337 个
数据文件结构
compliments-reference-db/ ├── phase1/ (源代码、输出、验证、统计) ├── phase2/ (源代码、输出、验证、统计) ├── phase3/ (参考产品目录、产品分组映射、分类映射、非食品产品表) ├── phase4/ (产品变体映射、参考产品变体表) ├── phase5/ (每100g营养数据、产品营养映射) └── phase6/ (产品评分、Agribalyse映射、评分排除表)
快速使用
python from huggingface_hub import hf_hub_download import pandas as pd
加载任意阶段输出
path = hf_hub_download("saraNour/compliments-reference-db", "phase6/outputs/phase6_product_scores.parquet", repo_type="dataset") df = pd.read_parquet(path)
主要输出表
- reference_product_catalog.csv: 3,691 个参考产品组
- product_group_mapping.csv: 4,440 个产品到组的映射
- product_variant_mapping.parquet: 4,440 个产品到变体的映射
- reference_product_variants.parquet: 4,299 个唯一变体
- nutrition_per_100g.parquet: 4,440 个产品的每 100g 营养数据
- phase6_product_scores.parquet: 4,440 个产品的 Nutri-Score 评分
- phase6_agribalyse_mapping.parquet: 4,440 个产品的 Agribalyse 类别映射





