Animal_Species_Synthetic_Dataset_Dirty_and_Enriched
收藏资源简介:
Synthetic Animal Dataset for Data Cleaning & Classification (4000 Rows) This dataset contains 4,000 synthetic animal records designed for educational and machine learning purposes. It was generated to simulate a realistic yet imperfect data environment, making it suitable for data cleaning, preprocessing, feature engineering, and classification tasks. The dataset includes both numerical and categorical attributes describing animal characteristics such as weight, body length, gender, age, speed (km/h), color, and biological class (e.g., mammal, reptile, bird, amphibian, fish, insect). It also contains geographic coordinates formatted in directional notation (e.g., 31.9539°N, 35.9106°E), observation dates, and unique animal identifiers. To reflect real-world data challenges, the dataset intentionally includes: Missing values (NULL entries) Inconsistent or previously misspelled categorical values (e.g., corrected gender labels such as "unknown") Mixed data types Synthetic yet statistically distributed numerical features The numerical attributes were generated based on statistical distributions to preserve realistic variability, while categorical attributes follow probability-based distributions. Although the data is synthetic and does not represent real biological measurements, it maintains logical structure and internal consistency suitable for experimentation. Intended Use Cases: Data cleaning and preprocessing practice Handling missing values and categorical encoding Exploratory data analysis (EDA) Supervised learning and classification modeling Feature engineering exercises Important Note:This is a fully synthetic dataset created for learning and experimentation. It should not be used for real-world biological or scientific conclusions.



