Awesome Public Datasets
收藏github2023-09-09 更新2024-05-31 收录
下载链接:
https://github.com/zhuwenxing/awesome-public-datasets
下载链接
链接失效反馈官方服务:
资源简介:
一个主题中心的高质量公开数据集列表,收集并整理自博客、问答和用户反馈。
A curated list of high-quality public datasets centered around specific themes, collected and organized from blogs, Q&A platforms, and user feedback.
创建时间:
2018-05-27
原始信息汇总
数据集概述
农业
- U.S. Department of Agricultures Nutrient Database
- 链接: https://www.ars.usda.gov/northeast-area/beltsville-md/beltsville-human-nutrition-research-center/nutrient-data-laboratory/docs/sr28-download-files/
- U.S. Department of Agricultures PLANTS Database
- 链接: http://www.plants.usda.gov/dl_all.html
生物学
- 1000 Genomes
- 链接: http://www.1000genomes.org/data
- American Gut (Microbiome Project)
- 链接: https://github.com/biocore/American-Gut
- Broad Bioimage Benchmark Collection (BBBC)
- 链接: https://www.broadinstitute.org/bbbc
- Broad Cancer Cell Line Encyclopedia (CCLE)
- 链接: http://www.broadinstitute.org/ccle/home
- Cell Image Library
- 链接: http://www.cellimagelibrary.org
- Complete Genomics Public Data
- 链接: http://www.completegenomics.com/public-data/69-genomes/
- EBI ArrayExpress
- 链接: http://www.ebi.ac.uk/arrayexpress/
- EBI Protein Data Bank in Europe
- 链接: http://www.ebi.ac.uk/pdbe/emdb/index.html/
- ENCODE project
- 链接: https://www.encodeproject.org
- Electron Microscopy Pilot Image Archive (EMPIAR)
- 链接: http://www.ebi.ac.uk/pdbe/emdb/empiar/
- Ensembl Genomes
- 链接: http://ensemblgenomes.org/info/genomes
- Gene Expression Omnibus (GEO)
- 链接: http://www.ncbi.nlm.nih.gov/geo/
- Gene Ontology (GO)
- 链接: http://geneontology.org/page/download-annotations
- Global Biotic Interactions (GloBI)
- 链接: https://github.com/jhpoelen/eol-globi-data/wiki#accessing-species-interaction-data
- Harvard Medical School (HMS) LINCS Project
- 链接: http://lincs.hms.harvard.edu
- Human Genome Diversity Project
- 链接: http://www.hagsc.org/hgdp/files.html
- Human Microbiome Project (HMP)
- 链接: http://www.hmpdacc.org/reference_genomes/reference_genomes.php
- ICOS PSP Benchmark
- 链接: http://ico2s.org/datasets/psp_benchmark.html
- International HapMap Project
- 链接: http://hapmap.ncbi.nlm.nih.gov/downloads/index.html.en
- Journal of Cell Biology DataViewer
- 链接: http://jcb-dataviewer.rupress.org
- KEGG
- 链接: http://www.genome.jp/kegg/
- MIT Cancer Genomics Data
- 链接: http://www.broadinstitute.org/cgi-bin/cancer/datasets.cgi
- NCBI Proteins
- 链接: http://www.ncbi.nlm.nih.gov/guide/proteins/#databases
- NCBI Taxonomy
- 链接: http://www.ncbi.nlm.nih.gov/taxonomy
- NCI Genomic Data Commons
- 链接: https://gdc-portal.nci.nih.gov
- NIH Microarray data
- 链接: http://bit.do/VVW6
- OpenSNP genotypes data
- 链接: https://opensnp.org/
- Pathguid - Protein-Protein Interactions Catalog
- 链接: http://www.pathguide.org/
- Protein Data Bank
- 链接: http://www.rcsb.org/
- Psychiatric Genomics Consortium
- 链接: https://www.med.unc.edu/pgc/downloads
- PubChem Project
- 链接: https://pubchem.ncbi.nlm.nih.gov/
- PubGene (now Coremine Medical)
- 链接: http://www.pubgene.org/
- Sanger Catalogue of Somatic Mutations in Cancer (COSMIC)
- 链接: http://cancer.sanger.ac.uk/cosmic
- Sanger Genomics of Drug Sensitivity in Cancer Project (GDSC)
- 链接: http://www.cancerrxgene.org/
- Sequence Read Archive(SRA)
- 链接: http://www.ncbi.nlm.nih.gov/Traces/sra/
- Stanford Microarray Data
- 链接: http://smd.stanford.edu/
- Stowers Institute Original Data Repository
- 链接: http://www.stowers.org/research/publications/odr
- Systems Science of Biological Dynamics (SSBD) Database
- 链接: http://ssbd.qbic.riken.jp
- The Cancer Genome Atlas (TCGA), available via Broad GDAC
- 链接: https://gdac.broadinstitute.org/
- The Catalogue of Life
- 链接: http://www.catalogueoflife.org/content/annual-checklist-archive
- The Personal Genome Project
- 链接: http://www.personalgenomes.org/
- UCSC Public Data
- 链接: http://hgdownload.soe.ucsc.edu/downloads.html
- UniGene
- 链接: http://www.ncbi.nlm.nih.gov/unigene
- Universal Protein Resource (UnitProt)
- 链接: http://www.uniprot.org/downloads
气候+天气
- Actuaries Climate Index
- 链接: http://actuariesclimateindex.org/data/
- Australian Weather
- 链接: http://www.bom.gov.au/climate/dwo/
- Aviation Weather Center - Consistent, timely and accurate weather [...]
- 链接: https://aviationweather.gov/adds/dataserver
- Brazilian Weather - Historical data (In Portuguese)
- 链接: http://sinda.crn2.inpe.br/PCD/SITE/novo/site/
- Canadian Meteorological Centre
- 链接: http://weather.gc.ca/grib/index_e.html
- Climate Data from UEA (updated monthly)
- 链接: https://crudata.uea.ac.uk/cru/data/temperature/#datter and ftp://ftp.cmdl.noaa.gov/
- European Climate Assessment & Dataset
- 链接: http://eca.knmi.nl/
- Global Climate Data Since 1929
- 链接: http://en.tutiempo.net/climate
- NASA Global Imagery Browse Services
- 链接: https://wiki.earthdata.nasa.gov/display/GIBS
- NOAA Bering Sea Climate
- 链接: http://www.beringclimate.noaa.gov/
- NOAA Climate Datasets
- 链接: http://www.ncdc.noaa.gov/data-access/quick-links
- NOAA Realtime Weather Models
- 链接: http://www.ncdc.noaa.gov/data-access/model-data/model-datasets/numerical-weather-prediction
- NOAA SURFRAD Meteorology and Radiation Datasets
- 链接: https://www.esrl.noaa.gov/gmd/grad/stardata.html
- The World Bank Open Data Resources for Climate Change
- 链接: http://data.worldbank.org/developers/climate-data-api
- UEA Climatic Research Unit
- 链接: http://www.cru.uea.ac.uk/data
- WU Historical Weather Worldwide
- 链接: https://www.wunderground.com/history/index.html
- WorldClim - Global Climate Data
- 链接: http://www.worldclim.org
复杂网络
- AMiner Citation Network Dataset
- 链接: http://aminer.org/citation
- CrossRef DOI URLs
- 链接: https://archive.org/details/doi-urls
- DBLP Citation dataset
- 链接: https://kdl.cs.umass.edu/display/public/DBLP
- DIMACS Road Networks Collection
- 链接: http://www.dis.uniroma1.it/challenge9/download.shtml
- NBER Patent Citations
- 链接: http://nber.org/patents/
- NIST complex networks data collection
- 链接: http://math.nist.gov/~RPozo/complex_datasets.html
- Network Repository with Interactive Exploratory Analysis Tools
- 链接: http://networkrepository.com/
- Protein-protein interaction network
- 链接: http://vlado.fmf.uni-lj.si/pub/networks/data/bio/Yeast/Yeast.htm
- PyPI and Maven Dependency Network
- 链接: https://ogirardot.wordpress.com/2013/01/31/sharing-pypimaven-dependency-data/
- Scopus Citation Database
- 链接: https://www.elsevier.com/solutions/scopus
- Small Network Data
- 链接: http://www-personal.umich.edu/~mejn/netdata/
- Stanford GraphBase
- 链接: http://www3.cs.stonybrook.edu/~algorith/implement/graphbase/implement.shtml
- Stanford Large Network Dataset Collection
- 链接: http://snap.stanford.edu/data/
- Stanford Longitudinal Network Data Sources
- 链接: http://stanford.edu/group/sonia/dataSources/index.html
- The Koblenz Network Collection
- 链接: http://konect.uni-koblenz.de/
- The Laboratory for Web Algorithmics (UNIMI)
- 链接: http://law.di.unimi.it/datasets.php
- The Nexus Network Repository
- 链接: http://nexus.igraph.org/
- UCI Network Data Repository
- 链接: https://networkdata.ics.uci.edu/resources.php
- UFL sparse matrix collection
- 链接: http://www.cise.ufl.edu/research/sparse/matrices/
- WSU Graph Database
- 链接: http://www.eecs.wsu.edu/mgd/gdb.html
计算机网络
- 3.5B Web Pages from CommonCrawl 2012
- 链接: http://www.bigdatanews.com/profiles/blogs/big-data-set-3-5-billion-web-pages-made-available-for-all-of-us
- 53.5B Web clicks of 100K users in Indiana Univ.
- 链接: http://cnets.indiana.edu/groups/nan/webtraffic/click-dataset/
- CAIDA Internet Datasets
- 链接: http://www.caida.org/data/overview/
- CRAWDAD Wireless datasets from Dartmouth Univ.
- 链接: https://crawdad.cs.dartmouth.edu/
- ClueWeb09 - 1B web pages
- 链接: http://lemurproject.org/clueweb09/
- ClueWeb12 - 733M web pages
- 链接: http://lemurproject.org/clueweb12/
- CommonCrawl Web Data over 7 years
- 链接: http://commoncrawl.org/the-data/get-started/
- Criteo click-through data
- 链接: http://labs.criteo.com/2015/03/criteo-releases-its-new-dataset/
- Internet-Wide Scan Data Repository
- 链接: https://scans.io/
- OONI: Open Observatory of Network Interference - Internet censorship data
- 链接: https://ooni.torproject.org/data/
- Open Mobile Data by MobiPerf
- 链接: https://console.developers.google.com/storage/openmobiledata_public/
- Rapid7 Sonar Internet Scans
- 链接: https://sonar.labs.rapid7.com/
- UCSD Network Telescope, IPv4 /8 net
- 链接: http://www.caida.org/projects/network_telescope/
数据挑战
- Bruteforce Database
- 链接: https://github.com/duyetdev/bruteforce-database
- Challenges in Machine Learning
- 链接: http://www.chalearn.org/
- CrowdANALYTIX dataX
- 链接: http://data.crowdanalytix.com
- D4D Challenge of Orange
- 链接: http://www.d4d.orange.com/en/home
- DrivenData Competitions for Social Good
- 链接: http://www.drivendata.org/
- ICWSM Data Challenge (since 2009)
- 链接: http://icwsm.cs.umbc.edu/
- KDD Cup by Tencent 2012
- 链接: http://www.kddcup2012.org/
- Kaggle Competition Data
- 链接: https://www.kaggle.com/
- Localytics Data Visualization Challenge
- 链接: https://github.com/localytics/data-viz-challenge
- Netflix Prize
- 链接: http://netflixprize.com/leaderboard.html
- Space Apps Challenge
- 链接: https://2015.spaceappschallenge.org/
- Telecom Italia Big Data Challenge
- 链接: https://dandelion.eu/datamine/open-big-data/
- TravisTorrent Dataset - MSR2017 Mining Challenge
- 链接: https://travistorrent.testroots.org/
- TunedIT - Data mining & machine learning data sets, algorithms, challenges
- 链接: http://tunedit.org/challenges/
- Yelp Dataset Challenge
- 链接: http://www.yelp.com/dataset_challenge
地球科学
- AQUASTAT - Global water resources and uses
- 链接: http://www.fao.org/nr/water/aquastat/data/query/index.html?lang=en
- BODC - marine data of ~22K vars
- 链接: https://www.bodc.ac.uk/data/
- EOSDIS - NASAs earth observing system data
- 链接: http://sedac.ciesin.columbia.edu/data/sets/browse
- Earth Models
- 链接: http://www.earthmodels.org/
- Integrated Marine Observing System (IMOS) - roughly 30TB of ocean measurements
- 链接: https://imos.aodn.org.au/
- Marinexplore - Open Oceanographic Data
- 链接: http://marinexplore.org/
- Smithsonian Institution Global Volcano and Eruption Database
- 链接: http://volcano.si.edu/
- USGS Earthquake Archives
- 链接: http://earthquake.usgs.gov/earthquakes/search/
经济学
- American Economic Association (AEA)
- 链接: https://www.aeaweb.org/resources/data
- EconData from UMD
- 链接: http://inforumweb.umd.edu/econdata/econdata.html
- Economic Freedom of the World Data
- 链接: http://www.freetheworld.com/datasets_efw.html
- Historical MacroEconomic Statistics
- 链接: http://www.historicalstatistics.org/
- **INFORUM - Inter
搜集汇总
数据集介绍

构建方式
Awesome Public Datasets 是一个高质量、主题导向的公共数据源集合,涵盖了多个领域的开放数据集。该数据集的构建方式主要依赖于从博客、问答平台以及用户反馈中收集和整理数据。通过自动化工具 `apd-core` 生成和维护,确保数据集的持续更新和准确性。数据集的内容经过严格筛选,确保其权威性和实用性。
特点
该数据集的特点在于其广泛的主题覆盖范围,涵盖了农业、生物学、气候与天气、复杂网络、计算机网络、数据挑战、地球科学、经济学、教育、能源、金融和地理信息系统等多个领域。每个数据集都经过标注,标明其可用性和状态(如“OK”表示可用,“FIXME”表示需要修复)。数据集的质量较高,大部分数据源为免费提供,部分数据集可能需要付费获取。
使用方法
用户可以通过访问 Awesome Public Datasets 的 GitHub 页面,浏览按主题分类的数据集列表。每个数据集都附有详细的描述和链接,用户可以直接访问原始数据源进行下载和使用。对于希望贡献新数据集的用户,项目提供了明确的贡献指南,建议通过 `apd-core` 工具提交数据,而非直接修改主仓库。该数据集适用于研究人员、数据科学家以及对开放数据感兴趣的开发者。
背景与挑战
背景概述
Awesome Public Datasets 是一个由社区驱动的公共数据集集合,旨在为研究人员、开发者和数据科学家提供高质量、主题广泛的开放数据资源。该数据集由 awesomedata 组织维护,涵盖了从农业、生物学到气候、经济等多个领域的数据集。其创建时间可追溯至2010年代初期,随着数据科学和机器学习的兴起,该数据集逐渐成为学术界和工业界的重要参考资源。通过整合来自博客、用户反馈和其他开放数据源的资源,Awesome Public Datasets 为全球的研究人员提供了一个便捷的数据获取平台,极大地推动了数据驱动的研究和创新。
当前挑战
尽管 Awesome Public Datasets 提供了丰富的数据资源,但其构建和维护仍面临诸多挑战。首先,数据集的多样性和广泛性使得数据质量难以统一,部分数据集可能存在格式不一致、数据缺失或更新不及时的问题。其次,随着数据量的增加,如何有效管理和分类这些数据集成为一个技术难题,尤其是在自动化生成和维护过程中,可能会出现数据重复或分类错误的情况。此外,部分数据集可能涉及版权或隐私问题,如何在开放共享与数据保护之间找到平衡,也是该平台需要持续解决的问题。最后,随着数据科学领域的快速发展,如何及时更新和扩展数据集以满足新兴研究需求,也是该平台面临的长期挑战。
常用场景
经典使用场景
Awesome Public Datasets 是一个广泛收集和整理高质量公共数据源的资源库,涵盖了从农业、生物学到气候、经济等多个领域。该数据集最经典的使用场景是为研究人员提供跨学科的数据支持,特别是在数据驱动的科学研究中,帮助学者快速获取所需的数据集,从而加速研究进程。例如,生物学家可以通过该数据集获取基因组数据,气候学家则可以访问全球气候数据,进行深入的分析和建模。
解决学术问题
Awesome Public Datasets 解决了学术研究中数据获取困难的问题。许多研究领域的数据分散在不同的平台和机构中,获取和整理这些数据往往耗费大量时间和精力。该数据集通过集中整理和分类,提供了一个便捷的入口,使得研究人员能够快速找到所需的数据,避免了数据获取的瓶颈。此外,该数据集还促进了跨学科研究,使得不同领域的研究者能够共享数据资源,推动了科学研究的协同发展。
衍生相关工作
Awesome Public Datasets 衍生了许多经典的研究工作。例如,基于该数据集中的基因组数据,研究人员开发了新的生物信息学工具和算法;利用气候数据,科学家们构建了全球气候模型,预测未来的气候变化趋势。此外,该数据集还促进了开源社区的发展,许多数据科学家和开发者基于这些数据开发了开源工具和平台,进一步推动了数据科学领域的创新和进步。
以上内容由遇见数据集搜集并总结生成



