遇见数据集

Meteorological data

收藏
Mendeley Data2026-04-18 收录
官方服务:

资源简介:

This dataset provides comprehensive meteorological time-series data for two major Indian cities, Jaipur and Kanpur, structured to facilitate robust climate analysis, environmental modeling, and machine learning research. The collection comprises four CSV files: two for each city, differentiated by the presence or absence of missing values. For Jaipur, both "with missing value" and "without missing value" datasets contain 575 records across six primary atmospheric variables: Dew point, Cloud cover, Precipitation, Temperature, Wind direction, and Wind speed. The Kanpur datasets follow a similar design but include an additional 'date\_time' column, representing concatenated timestamps, and extend to 575 records with seven columns in total. The inclusion of both raw (with missing values) and preprocessed (imputed, without missing values) variants enables users to benchmark data cleaning, imputation, and gap-filling algorithms, offering valuable real-world scenarios where sensor data is incomplete. All variables are provided in a consistent numerical format, with the imputed files utilizing float datatypes for improved compatibility with statistical and computational methods. The temporal granularity and spatial specificity of the data make it well-suited for urban climate studies, high-resolution forecasting, and analysis of atmospheric dynamics in Northern India. While the units for each variable (e.g., temperature, wind speed) follow standard meteorological conventions, users are advised to consult supplementary documentation or the data authors for precise measurement details, as direct unit annotation is not present within the files. The 'date\_time' field in the Kanpur datasets warrants special attention due to its composite format, and may require parsing for fine-grained temporal analysis. The data capture typical meteorological variation as observed by regional weather stations, including natural fluctuations, diurnal cycles, and seasonal transitions, while also realistically reflecting common challenges in environmental monitoring, such as sporadic missing values arising from sensor outages or data transmission errors. By providing both incomplete and fully-imputed datasets, this resource supports the development and evaluation of robust analytical pipelines in climate informatics and related disciplines. Researchers can leverage these datasets for tasks ranging from baseline climate characterization and anomaly detection to the training and validation of data-driven predictive models. Ultimately, the structured format, practical relevance, and dual-variant design position this dataset as a valuable open-access asset for atmospheric scientists, data engineers, and interdisciplinary researchers interested in advancing methods for meteorological data analysis, urban climatology, and environmental data science.

本数据集涵盖印度两大城市斋浦尔(Jaipur)与坎普尔(Kanpur)的完整气象时序数据,其结构化设计可支撑高质量气候分析、环境建模与机器学习研究。本数据集共包含4个CSV文件,每个城市对应2个文件,以是否存在缺失值作为区分依据。针对斋浦尔(Jaipur),含缺失值与无缺失值两个数据集均包含575条记录,涵盖6项核心大气变量:露点(Dew point)、云量(Cloud cover)、降水量(Precipitation)、气温(Temperature)、风向(Wind direction)与风速(Wind speed)。坎普尔(Kanpur)数据集的设计逻辑一致,但额外包含一列`date_time`(拼接时间戳),同样拥有575条记录,总计7列数据。同时提供原始(含缺失值)与预处理(插补后无缺失值)两种变体,可支持用户对数据清洗、缺失值插补与间隙填充算法进行性能基准测试,还原传感器数据缺失的真实应用场景。所有变量均采用统一的数值格式,其中插补后的数据集采用浮点(float)数据类型,以提升与统计及计算方法的兼容性。该数据的时间粒度与空间特异性使其非常适用于印度北部的城市气候研究、高分辨率预报以及大气动力学分析。尽管各变量的单位(如气温、风速)遵循标准气象惯例,但由于文件中未直接标注单位,建议用户查阅补充文档或联系数据作者以获取精确的测量细节。坎普尔数据集中的`date_time`字段采用复合格式,需特别注意,若要开展细粒度时间分析,可能需要对其进行格式解析。该数据集收录了区域气象站观测到的典型气象变化,包括自然波动、昼夜循环与季节更替,同时真实还原了环境监测中常见的挑战,例如因传感器故障或数据传输错误导致的零星缺失值。通过同时提供不完整与完全插补后的数据集,本资源可支撑气候信息学及相关领域中稳健分析流程的开发与评估。研究人员可利用该数据集完成多项任务,从基础气候特征刻画、异常检测,到数据驱动预测模型的训练与验证。最终,其结构化格式、实际应用价值与双变体设计,使本数据集成为大气科学家、数据工程师,以及致力于推进气象数据分析、城市气候学与环境数据科学方法的跨学科研究人员的宝贵开放获取资源。

创建时间:
2025-09-18
二维码
社区交流群
二维码
科研交流群
商业服务