arsalan-anwari/zer-data
收藏资源简介:
荷兰执法实体解析基准数据集是一个用于实体解析和记录链接研究的合成基准数据集集合,基于真实的荷兰和欧盟/欧洲经济区执法数据模式建模。该数据集旨在解决荷兰执法数据中的特定挑战,如荷兰名字复杂性、多文化名字变体、跨申根身份问题、ANPR OCR错误、估计出生日期、电信身份变动和金融网络结构。数据集包含多种类型的记录,每种记录都有详细的模式、真实配对格式和注入的错误模式,以复制现实世界的挑战。数据集提供多种规模,适用于不同的使用场景,所有数据均为100%合成,以避免隐私或法律问题。
The Dutch Law Enforcement Entity Resolution Benchmark is a collection of synthetic benchmark datasets for entity resolution and record linkage research, modelled on real Dutch and EU/EEA law enforcement data schemas. The dataset is designed to address specific challenges in Dutch law enforcement data, such as Dutch name complexity, multicultural name variation, cross-Schengen identity issues, ANPR OCR errors, estimated dates of birth, telecom identity churn, and financial network structure. The dataset includes multiple types of records, each with detailed schemas, ground truth formats, and injected error patterns to replicate real-world challenges. The dataset is available in various sizes, each suitable for different use cases, and all data is 100% synthetic to avoid privacy or legal concerns.




