MAVE
收藏资源简介:
MAVE数据集是由谷歌研究团队创建的,旨在促进产品属性值提取研究。该数据集包含从亚马逊页面精选的220万种产品,覆盖1257个独特类别,拥有300万条属性值标注。MAVE数据集的四大优势包括:规模最大、多源信息捕获、属性值多样性以及提供具有挑战性的零样本测试集。此数据集不仅适用于属性值提取任务,还能有效应对零样本属性提取的挑战,广泛应用于电子商务领域的客户服务机器人、产品排名、检索和推荐系统等。
The MAVE dataset was developed by the Google Research team to facilitate research on product attribute value extraction. This dataset includes 2.2 million products curated from Amazon webpages, spanning 1,257 distinct categories, and features 3 million annotated attribute value entries. The four core advantages of the MAVE dataset are as follows: the largest scale, multi-source information capture, diverse attribute values, and provision of a challenging zero-shot test set. This dataset is not only suitable for attribute value extraction tasks, but also effectively addresses the challenges of zero-shot attribute extraction, and is widely applied in e-commerce scenarios including customer service robots, product ranking, retrieval, and recommendation systems.




