Synthetic Benchmark Dataset for Streaming Multi-Source Integration (N=180)
收藏资源简介:
This dataset contains 180 synthetic benchmark test scenarios used to validate the adaptive credibility assessment framework presented in the paper "Streaming Multi-Source Integration Through Adaptive Credibility Assessment." The dataset includes:- 180 controlled test scenarios with known ground truth (test_cases.csv)- Simulated measurements from 4 providers (Google Analytics 4, Mixpanel, PostHog, Microsoft Clarity)- Categories: basic operation (60%), edge cases (20%), failure simulations (20%)- Sample server-side logs for validation (server_logs_sample.csv)- Python validation and data generation scripts Key Results:- Mean Absolute Error (MAE): 4.1% (weighted aggregation)- 67% error reduction vs. best single source (12.3% MAE)- Ground truth agreement: 95.2% The dataset enables researchers to reproduce paper results and validate multi-source integration approaches. Related Paper:"Streaming Multi-Source Integration Through Adaptive Credibility Assessment"Authors: Ramsha Mehreen, Renikunta Ramesh Journal: Knowledge-Based Systems (Elsevier) GitHub Repository: https://github.com/m24ds001/streaming-multi-experts



