Synthetic and Encoded Database of Dengue, Zika, Chikungunya, and Influenza From Literature for Machine Learning
收藏资源简介:
The data included in this database represent patients exhibiting symptomatology indicative of one of the aforementioned diseases. The Excel file is divided into four sheets, each corresponding to one pathology. All sheets contain the same number and type of columns, corresponding to the label (pathology) and the features (symptoms). The pathologies included—Dengue, Zika, Chikungunya, and Influenza—share signs and symptoms during their early symptomatic phase, the febrile stage. The first three are arboviral diseases, meaning viral infections transmitted through the bite of infected arthropods, primarily Aedes aegypti and Aedes albopictus. In addition to sharing vectors and certain symptoms, these diseases also share geographic distribution and even seasonality, as well as vector activity increases under specific climatic conditions. Laboratory tests are ultimately required to confirm infection; however, diagnosis for all three diseases is initially clinical. Influenza was included because a previous study hypothesized that adding a negative control to predictive models would enable more accurate evaluation of arboviral diseases. Thus, if a patient’s data did not match any arboviral infection, Influenza could serve as the alternative diagnosis. Influenza also shares several symptoms with the three arboviral diseases and is prone to epidemic outbreaks, making it a meaningful comparative dataset. The database consists of 22,379 Dengue records, 7,135 Zika records, 7,959 Chikungunya records, and 10,741 Influenza records. For Dengue, data were collected from patients diagnosed with Dengue without warning signs and Dengue with warning signs, but do not include severe Dengue cases. For Influenza, records include Influenza A, B, and seasonal Influenza. All values are binary. Each row represents an individual record, beginning with the label followed by values of 1 if the clinical sign was present and 0 if absent. No missing values or NaNs (Not a Number) are present. This structure allows for direct statistical analysis and integration into ML models without the need for imputation.



