遇见数据集

Synthetic mobile device data

收藏
Mendeley Data2021-05-04 更新2026-04-09 收录
官方服务:

资源简介:

This repository contains two synthetic mobile device datasets, one for GPS location records ("input_case1_v2.csv") and the other for cellular location records ("input_case2_v2.csv"). The two datasets are stored in CSV files. In each CSV file, there are 12 data fields, explained in the "Dictionary.docx" file. The two datasets contain the location records for 582 individual mobile devices for a month. The GPS dataset ("input_case1_v2.csv") contains 668,939 location records, and the cellular dataset ("input_case2_v2.csv") contains 61,390 location records. Using the synthetic mobile data generation method developed by Chen et al. (2014), the two datasets are generated based on two real-world data sources. The first one is mobile app data, which comes from people using location-aware mobile apps. The mobile app data encompasses both GPS and cellular data, and covers the month of March in 2019 in the central Puget Sound region. It includes 582 individual mobile device users. The second data source is household travel survey data. It covers the month of March in 2017 in the central Puget Sound region, and includes 582 survey respondents. The 582 (mobile device) users and the 582 survey respondents are randomly linked. The visited locations in the household travel survey are viewed as the ground-truth stays. Four fields of information from the mobile app data are preserved in the synthetic location records: the number of location records, and the user ID (anonymized), timestamp of each location records, and location accuracy associated with a record. If the timestamp of a location record falls within the duration of a ground-truth stay, the location record will be associated to the stay. The latitudes and longitudes of synthetic location records are generated such that their spatial distribution is the same as that from the mobile app data for a given user on a given day. The spatial distribution is measured by the distance and angle from a location record to the corresponding stay. Methods to infer stays from the mobile data is described in Wang et al., (2019), which was developed using the method developed in (Chen et al. 2014). For synthetic location records not associated to any (ground-truth) stay, their locations are random deviates from locations evenly distributed on the straight line connecting the last and the next stays, as described in Chen et al. (2014).

本仓库包含两个合成移动设备数据集,其一为GPS位置记录数据集(文件名为"input_case1_v2.csv"),其二为蜂窝网络位置记录数据集(文件名为"input_case2_v2.csv")。两类数据集均以CSV格式存储,每份CSV文件包含12个数据字段,字段说明详见"Dictionary.docx"文件。这两个数据集均涵盖582台独立移动设备为期一个月的位置记录。其中GPS位置记录数据集("input_case1_v2.csv")共包含668,939条位置记录,蜂窝网络位置记录数据集("input_case2_v2.csv")共包含61,390条位置记录。这两类合成数据集采用Chen等人于2014年提出的合成移动数据生成方法,基于两类真实数据源构建而成。第一类数据源为移动应用数据,源自使用位置感知类移动应用的用户群体,涵盖2019年3月普吉特海湾中部地区的GPS与蜂窝网络位置数据,涉及582名独立移动设备用户。第二类数据源为家庭出行调查数据,覆盖2017年3月普吉特海湾中部地区,共纳入582名调查受访者。本次研究将582名移动设备用户与582名调查受访者进行随机匹配。家庭出行调查中记录的到访点位将作为真实停留基准(ground-truth stays)。合成位置记录保留了移动应用数据中的四类信息字段:位置记录编号、匿名化用户ID、每条位置记录的时间戳,以及关联至该记录的位置精度参数。若某条位置记录的时间戳落在某一真实停留基准的持续时段内,则该记录将关联至该停留点位。合成位置记录的经纬度生成逻辑为:针对特定用户在特定日期的位置记录,其空间分布与对应移动应用数据中的空间分布保持一致。空间分布通过位置记录至对应停留点位的距离与方位角进行衡量。从移动数据中推断停留点位的方法详见Wang等人2019年的研究,该方法基于Chen等人2014年提出的方法开发而来。对于未关联至任何真实停留基准的合成位置记录,其点位将按照Chen等人2014年提出的方法生成:在相邻两次停留点位的连线上均匀分布的点位中随机选取。

创建时间:
2021-05-04
二维码
社区交流群
二维码
科研交流群
商业服务