遇见数据集

CaDENCE: a large-scale call disconnection event dataset from consumer Android devices (Mar–Jun 2024)

收藏
Zenodo2026-07-18 更新2026-08-01 收录
官方服务:

资源简介:

CaDENCE is a curated event-level dataset of call disconnection records collected from consumer Android smartphones. The dataset was prepared to support studies on call termination outcomes, dropped-call characterization, radio access technology transitions, signal-quality conditions, device-side network diagnostics, and machine learning methods applied to event-level mobile network data. The source records were generated by event-driven device-level software instrumentation and were curated through filtering, standardization, pseudonymization, binary labeling, and release packaging. The dataset contains 125,532,358 records collected over 98 days, from 2024-03-09 to 2024-06-14. Each row corresponds to one call disconnection event and includes temporal attributes, pseudonymous device and software identifiers, mobile network codes, radio access technology indicators, radio band and channel information, signal-quality measurements (RSRP and RSRQ), service-state flags, and the binary label is_drop, which distinguishes dropped-call events from other non-drop call termination outcomes. Important note on sampling: the public release was produced using a daily capped export strategy that prioritized dropped-call records. As a result, the released class distribution is enriched for dropped-call analysis and should not be interpreted as the original call-outcome distribution in the source telemetry. Aggregate warehouse-versus-release counts are provided to document this effect. The warehouse-side values are author-provided aggregates derived from non-public source data, while the released-dataset values can be reproduced from the published Parquet partitions. The complete release is provided as a compressed archive with the following structure: CaDENCE_call_disconnection_dataset_release_v1_3/ ├── data_by_date/ ├── metadata/ │ ├── schema.json │ ├── missingness.csv │ ├── value_ranges.csv │ ├── summary_by_day.csv │ ├── summary_by_mcc.csv │ ├── calldrop_warehouse_vs_released.csv │ └── calldrop_warehouse_vs_released.png ├── sample/ │ ├── cadence_preview_sample.csv │ └── README.md ├── scripts/ │ ├── metadata_generation/ │ │ └── requirements.txt │ └── curation_etl_logic/ ├── README.md ├── CITATION.cff └── DATA_LICENSE.txt The preview CSV contains 9,624 released records and is provided for schema inspection and loading tests rather than statistical analysis. The metadata-generation scripts reproduce schema.json, missingness.csv, value_ranges.csv, summary_by_day.csv, and summary_by_mcc.csv from the released Parquet files. Version note: Version 1.3 updates the documentation and reproducibility materials without changing the released event-level Parquet data. This version corrects the documented encoding of day_of_week to 1 = Sunday and 7 = Saturday, reorganizes and documents the preview sample, adds Quick Start examples for pandas, Polars, and DuckDB, adds a dependency file for the metadata-generation scripts, introduces explicit caveats regarding mobility and vehicular-use analyses, clarifies the reproducibility limits of the warehouse-versus-release comparison, and further sanitizes the anonymized curation pseudocode. Funding/context: This dataset was produced within the PD&I SWPERFI Project, “Artificial Intelligence Techniques for Analysis and Optimization of Software Performance,” conducted in partnership with UFAM and Motorola Mobility under Agreement No. 004/2021 with ICOMP/UFAM and Brazilian Federal Law No. 8.387/1991.

提供机构:
Zenodo
创建时间:
2026-07-18
二维码
社区交流群
二维码
科研交流群
商业服务