MATCHED
收藏资源简介:
MATCHED数据集是由马斯特里赫特大学的Law & Tech Lab创建,专门用于多模态作者归属研究,旨在打击人口贩卖。该数据集包含27,619条独特的文本描述和55,115张图片,来源于Backpage escort平台,覆盖美国七个城市的广告数据。数据集的创建过程涉及从多个地理区域收集数据,并通过多任务训练框架进行处理,以提高分类和检索性能。MATCHED数据集主要应用于执法机构(LEAs),帮助其通过多模态分析识别和验证广告发布者,从而打击人口贩卖网络。
The MATCHED dataset was developed by the Law & Tech Lab at Maastricht University, specifically designed for multimodal author attribution research with the core goal of combating human trafficking. This dataset comprises 27,619 unique text descriptions and 55,115 images, sourced from the Backpage escort platform, and covers advertising data from seven U.S. cities. The construction of the dataset involves collecting data from multiple geographic regions and processing it via a multi-task training framework to enhance classification and retrieval performance. The MATCHED dataset is primarily applied to Law Enforcement Agencies (LEAs), assisting them in identifying and verifying ad publishers through multimodal analysis so as to combat human trafficking networks.




