thebeautyapi/beautyproducts
收藏资源简介:
该数据集名为“The Beauty API: 180k+ Skincare & Cosmetics Products Dataset (Sample)”,是一个专注于美容科技和市场研究的结构化样本数据集。它包含超过180,000个独特的护肤和化妆品产品记录,覆盖全球20,000多个品牌(如Paulas Choice、CeraVe、The Ordinary等)。数据集提供18GB的高分辨率产品图像(标准化白色背景),并以CSV和JSONL格式交付。核心功能包括:1) 成分智能与安全分析:每个产品都有完整的INCI成分分解、常见别名和功能分类(如抗氧化、舒缓),并提供研究支持的致痘性评分(0-5)和刺激性评分(0-5)在成分级别,以及专家安全分类(如超级明星、好东西等),支持无酒精、无香料、无精油等高级标签过滤。2) 机器学习与计算机视觉应用:作为预标记训练集,可用于品牌分析(训练多模态模型解码包装设计与成分质量之间的关联)、视觉搜索(构建“点即知”相机功能,绕过条形码需求)和零售审计(实时识别180,000+ SKU)。3) 电子商务增强:元数据旨在提升独立电商和零售的转化率并降低退货率,包括页面丰富(生成SEO友好的产品描述和成分词汇表)和信任信号(实施安全徽章和个性化肤质匹配引擎)。技术架构包括稳定产品ID、品牌、名称、描述、图像文件名、标签、简单成分列表和丰富成分对象等字段。数据集适用于非商业研究,商业用途需额外许可。
The dataset is titled The Beauty API: 180k+ Skincare & Cosmetics Products Dataset (Sample), a structured sample dataset focused on beauty-tech and market research. It contains over 180,000 unique skincare and cosmetics product records, covering more than 20,000 global brands (e.g., Paulas Choice, CeraVe, The Ordinary). The dataset provides 18GB of high-resolution product images (standardized white backgrounds) and is delivered in CSV and JSONL formats. Key capabilities include: 1) Ingredients Intelligence & Safety: Each product is parsed with full INCI breakdowns, common aliases, and functional classifications (e.g., antioxidant, soothing), along with research-backed comedogenic ratings (0–5) and irritancy scores (0–5) at the ingredient level, expert safety classifications (e.g., Superstars, Goodies), and advanced tagging for filtering alcohol-free, fragrance-free, and essential oil-free formulations. 2) Machine Learning & Computer Vision: It serves as a pre-labeled training set for branding analysis (training multimodal models to decode correlations between packaging design and ingredient quality), visual search (building Point-and-Know camera features that bypass barcodes), and retail auditing (powering automated shelf-scanners to identify 180,000+ SKUs in real-world conditions). 3) E-commerce Enrichment: Metadata is designed to boost conversion and lower return rates for indie e-commerce and retail, including page enrichment (generating SEO-rich product descriptions and ingredient glossaries) and trust signals (implementing safety badges and personalized skin-type matching engines). The technical schema includes fields such as stable product ID, brand, name, description, image filename, tags, simple ingredients list, and rich ingredient objects. It is licensed for non-commercial research, with commercial use requiring additional licensing.





