遇见数据集

SyntInfra-India: A Large-Scale Synthetic Dataset for AI-Driven University Infrastructure Analytics in Indian Higher Education

收藏
Zenodo2026-03-26 更新2026-05-26 收录
官方服务:

资源简介:

SyntInfra-India is a large-scale synthetic dataset developed to support advanced research in Artificial Intelligence (AI), Machine Learning (ML), and Data Science, with a specific focus on university infrastructure analysis in the Indian higher education ecosystem. The dataset simulates realistic, multi-dimensional data representing institutional characteristics, infrastructure capacity, student demographics, academic load, and environmental factors across universities and colleges in India. It is designed to enable predictive modeling, optimization, and policy analysis without involving any real-world sensitive or personal data. The dataset incorporates temporal dynamics (multi-year and semester-level data), allowing researchers to study infrastructure evolution, resource utilization trends, and the impact of policy interventions such as AI program adoption and digital transformation initiatives. 📊 Key Characteristics Large-scale, high-dimensional dataset (50M+ records, configurable) India-centric design inspired by AISHE, UGC, and NIRF trends Multi-modal features: institutional, infrastructural, academic, environmental Time-series component for longitudinal analysis Machine learning-ready with derived target variables Fully synthetic and ethically compliant 📂 Dataset Components The dataset includes the following major feature groups: Institutional Data Institution type (Central, State, Private, Deemed) Location category (Tier 1, Tier 2, Tier 3, Rural) Accreditation and ranking indicators (synthetic) Infrastructure Data Classrooms, laboratories, and library capacity Smart classroom penetration Digital infrastructure index IT Infrastructure Internet bandwidth Campus Wi-Fi coverage GPU/HPC availability Cloud adoption index Student & Academic Data Enrollment (UG, PG, PhD) Student-faculty ratio Gender distribution Dropout and placement rates Environmental Factors Urbanization index Industry presence score Cost of living index Derived Variables (ML Targets) Infrastructure Stress Index Classroom Utilization Rate Hostel Demand Gap Energy Efficiency Score Student Satisfaction Proxy Score ⚙️ Technical Specifications Format: CSV / Parquet Data Type: Tabular + Time-Series Features: 80–150 columns Records: Scalable (millions to 50M+) Missing Values: Controlled and realistic Noise Injection: Statistical perturbation for realism 🧠 Data Generation Methodology The dataset is generated using a hybrid synthetic data generation framework: Statistical modeling based on Indian higher education distributions Rule-based constraints to ensure logical consistency Correlation-aware feature engineering Advanced generative techniques (e.g., CTGAN, variational methods) This approach ensures that the dataset maintains realistic patterns while remaining fully synthetic. 🧪 Potential Applications Infrastructure demand forecasting Smart campus planning and optimization AI adoption impact analysis Resource allocation modeling Urban vs rural education disparity analysis Benchmarking machine learning models 🔐 Ethics and Data Compliance This dataset is entirely synthetic and does not contain any real personal, institutional, or sensitive data. No ethical approval is required for its use.

提供机构:
Zenodo
创建时间:
2026-03-26
二维码
社区交流群
二维码
科研交流群
商业服务