遇见数据集

Multi-source cosmetic ingredient prevalence dataset (87-concept INCI ontology)

收藏
Zenodo2026-08-10 更新2026-08-13 收录
官方服务:

资源简介:

A prevalence dataset of 87 cosmetic active ingredients drawn from two ecological niches: an EU-leaning crowd-sourced corpus (Open Beauty Facts, n=19,045) and North-American retail, spanning premium (Sephora, n=1,472) and multi-platform (e-commerce catalogue, n=6,600) channels, plus a US regulatory negative control (California Safe Cosmetics Program, n=34,502). A reusable 87-concept INCI ontology is applied by a rule-based matcher validated on a 50-product gold standard (precision 1.000, recall 0.996, F1 0.998). A three-tier validation framework comprises: (Tier 1) random-sample gold-standard audit of 54 concepts; (Tier 2) per-concept positive-decision audit and negative-boundary testing of 32 rare concepts; (Tier 3) one concept with zero corpus prevalence. Cross-source Spearman correlations (common-filter rho = 0.56-0.81), a robustness check removing the shared retail subset, and a regulatory negative control (rho near zero, non-significant) are included. The dataset also provides Eclat frequent-itemset formula skeletons, five figures (SVG/PNG), the 87-concept ontology (CSV), and reproducible fixed-seed Node.js analysis scripts. LICENSE: Most files (aggregate statistics, figures, scripts, ontology) are CC-BY 4.0. Four validation-sample JSON files containing verbatim per-product ingredient text derived from Open Beauty Facts (ODbL-1.0) are licensed ODbL-1.0/DbCL-1.0; see LICENSE.md for the per-file breakdown.

提供机构:
Zenodo
创建时间:
2026-08-10
二维码
社区交流群
二维码
科研交流群
商业服务