Multi-source cosmetic ingredient prevalence dataset (87-concept INCI ontology)
收藏资源简介:
A prevalence dataset of 87 cosmetic active ingredients drawn from two ecological niches: an EU-leaning crowd-sourced corpus (Open Beauty Facts, n=19,045) and North-American retail, spanning premium (Sephora, n=1,472) and multi-platform (e-commerce catalogue, n=6,600) channels, plus a US regulatory negative control (California Safe Cosmetics Program, n=34,502). A reusable 87-concept INCI ontology is applied by a rule-based matcher validated on a 50-product gold standard (precision 1.000, recall 0.996, F1 0.998). A three-tier validation framework comprises: (Tier 1) random-sample gold-standard audit of 54 concepts; (Tier 2) per-concept positive-decision audit and negative-boundary testing of 32 rare concepts; (Tier 3) one concept with zero corpus prevalence. Cross-source Spearman correlations (common-filter rho = 0.56-0.81), a robustness check removing the shared retail subset, and a regulatory negative control (rho near zero, non-significant) are included. The dataset also provides Eclat frequent-itemset formula skeletons, five figures (SVG/PNG), the 87-concept ontology (CSV), and reproducible fixed-seed Node.js analysis scripts. LICENSE: Most files (aggregate statistics, figures, scripts, ontology) are CC-BY 4.0. Four validation-sample JSON files containing verbatim per-product ingredient text derived from Open Beauty Facts (ODbL-1.0) are licensed ODbL-1.0/DbCL-1.0; see LICENSE.md for the per-file breakdown.



