awesome-public-datasets
收藏数据集概述
Agriculture
-
U.S. Department of Agricultures Nutrient Database
- URL: https://www.ars.usda.gov/northeast-area/beltsville-md/beltsville-human-nutrition-research-center/nutrient-data-laboratory/docs/sr28-download-files/
-
U.S. Department of Agricultures PLANTS Database
- URL: http://www.plants.usda.gov/dl_all.html
Biology
-
1000 Genomes
- URL: http://www.1000genomes.org/data
-
American Gut (Microbiome Project)
- URL: https://github.com/biocore/American-Gut
-
Broad Bioimage Benchmark Collection (BBBC)
- URL: https://www.broadinstitute.org/bbbc
-
Broad Cancer Cell Line Encyclopedia (CCLE)
- URL: http://www.broadinstitute.org/ccle/home
-
Cell Image Library
- URL: http://www.cellimagelibrary.org
-
Complete Genomics Public Data
- URL: http://www.completegenomics.com/public-data/69-genomes/
-
EBI ArrayExpress
- URL: http://www.ebi.ac.uk/arrayexpress/
-
EBI Protein Data Bank in Europe
- URL: http://www.ebi.ac.uk/pdbe/emdb/index.html/
-
ENCODE project
- URL: https://www.encodeproject.org
-
Electron Microscopy Pilot Image Archive (EMPIAR)
- URL: http://www.ebi.ac.uk/pdbe/emdb/empiar/
-
Ensembl Genomes
- URL: http://ensemblgenomes.org/info/genomes
-
Gene Expression Omnibus (GEO)
- URL: http://www.ncbi.nlm.nih.gov/geo/
-
Gene Ontology (GO)
- URL: http://geneontology.org/page/download-annotations
-
Global Biotic Interactions (GloBI)
- URL: https://github.com/jhpoelen/eol-globi-data/wiki#accessing-species-interaction-data
-
Harvard Medical School (HMS) LINCS Project
- URL: http://lincs.hms.harvard.edu
-
Human Genome Diversity Project
- URL: http://www.hagsc.org/hgdp/files.html
-
Human Microbiome Project (HMP)
- URL: http://www.hmpdacc.org/reference_genomes/reference_genomes.php
-
ICOS PSP Benchmark
- URL: http://ico2s.org/datasets/psp_benchmark.html
-
International HapMap Project
- URL: http://hapmap.ncbi.nlm.nih.gov/downloads/index.html.en
-
Journal of Cell Biology DataViewer
- URL: http://jcb-dataviewer.rupress.org
-
KEGG
- URL: http://www.genome.jp/kegg/
-
MIT Cancer Genomics Data
- URL: http://www.broadinstitute.org/cgi-bin/cancer/datasets.cgi
-
NCBI Proteins
- URL: http://www.ncbi.nlm.nih.gov/guide/proteins/#databases
-
NCBI Taxonomy
- URL: http://www.ncbi.nlm.nih.gov/taxonomy
-
NCI Genomic Data Commons
- URL: https://gdc-portal.nci.nih.gov
-
NIH Microarray data
- URL: http://bit.do/VVW6
-
OpenSNP genotypes data
- URL: https://opensnp.org/
-
Pathguid - Protein-Protein Interactions Catalog
- URL: http://www.pathguide.org/
-
Protein Data Bank
- URL: http://www.rcsb.org/
-
Psychiatric Genomics Consortium
- URL: https://www.med.unc.edu/pgc/downloads
-
PubChem Project
- URL: https://pubchem.ncbi.nlm.nih.gov/
-
PubGene (now Coremine Medical)
- URL: http://www.pubgene.org/
-
Sanger Catalogue of Somatic Mutations in Cancer (COSMIC)
- URL: http://cancer.sanger.ac.uk/cosmic
-
Sanger Genomics of Drug Sensitivity in Cancer Project (GDSC)
- URL: http://www.cancerrxgene.org/
-
Sequence Read Archive(SRA)
- URL: http://www.ncbi.nlm.nih.gov/Traces/sra/
-
Stanford Microarray Data
- URL: http://smd.stanford.edu/
-
Stowers Institute Original Data Repository
- URL: http://www.stowers.org/research/publications/odr
-
Systems Science of Biological Dynamics (SSBD) Database
- URL: http://ssbd.qbic.riken.jp
-
The Cancer Genome Atlas (TCGA), available via Broad GDAC
- URL: https://gdac.broadinstitute.org/
-
The Catalogue of Life
- URL: http://www.catalogueoflife.org/content/annual-checklist-archive
-
The Personal Genome Project
- URL: http://www.personalgenomes.org/
-
UCSC Public Data
- URL: http://hgdownload.soe.ucsc.edu/downloads.html
-
UniGene
- URL: http://www.ncbi.nlm.nih.gov/unigene
-
Universal Protein Resource (UnitProt)
- URL: http://www.uniprot.org/downloads
Climate+Weather
-
Actuaries Climate Index
- URL: http://actuariesclimateindex.org/data/
-
Australian Weather
- URL: http://www.bom.gov.au/climate/dwo/
-
Aviation Weather Center - Consistent, timely and accurate weather [...]
- URL: https://aviationweather.gov/adds/dataserver
-
Brazilian Weather - Historical data (In Portuguese)
- URL: http://sinda.crn2.inpe.br/PCD/SITE/novo/site/
-
Canadian Meteorological Centre
- URL: http://weather.gc.ca/grib/index_e.html
-
Climate Data from UEA (updated monthly)
- URL: https://crudata.uea.ac.uk/cru/data/temperature/#datter and ftp://ftp.cmdl.noaa.gov/
-
European Climate Assessment & Dataset
- URL: http://eca.knmi.nl/
-
Global Climate Data Since 1929
- URL: http://en.tutiempo.net/climate
-
NASA Global Imagery Browse Services
- URL: https://wiki.earthdata.nasa.gov/display/GIBS
-
NOAA Bering Sea Climate
- URL: http://www.beringclimate.noaa.gov/
-
NOAA Climate Datasets
- URL: http://www.ncdc.noaa.gov/data-access/quick-links
-
NOAA Realtime Weather Models
- URL: http://www.ncdc.noaa.gov/data-access/model-data/model-datasets/numerical-weather-prediction
-
NOAA SURFRAD Meteorology and Radiation Datasets
- URL: https://www.esrl.noaa.gov/gmd/grad/stardata.html
-
The World Bank Open Data Resources for Climate Change
- URL: http://data.worldbank.org/developers/climate-data-api
-
UEA Climatic Research Unit
- URL: http://www.cru.uea.ac.uk/data
-
WU Historical Weather Worldwide
- URL: https://www.wunderground.com/history/index.html
-
WorldClim - Global Climate Data
- URL: http://www.worldclim.org
ComplexNetworks
-
AMiner Citation Network Dataset
- URL: http://aminer.org/citation
-
CrossRef DOI URLs
- URL: https://archive.org/details/doi-urls
-
DBLP Citation dataset
- URL: https://kdl.cs.umass.edu/display/public/DBLP
-
DIMACS Road Networks Collection
- URL: http://www.dis.uniroma1.it/challenge9/download.shtml
-
NBER Patent Citations
- URL: http://nber.org/patents/
-
NIST complex networks data collection
- URL: http://math.nist.gov/~RPozo/complex_datasets.html
-
Network Repository with Interactive Exploratory Analysis Tools
- URL: http://networkrepository.com/
-
Protein-protein interaction network
- URL: http://vlado.fmf.uni-lj.si/pub/networks/data/bio/Yeast/Yeast.htm
-
PyPI and Maven Dependency Network
- URL: https://ogirardot.wordpress.com/2013/01/31/sharing-pypimaven-dependency-data/
-
Scopus Citation Database
- URL: https://www.elsevier.com/solutions/scopus
-
Small Network Data
- URL: http://www-personal.umich.edu/~mejn/netdata/
-
Stanford GraphBase
- URL: http://www3.cs.stonybrook.edu/~algorith/implement/graphbase/implement.shtml
-
Stanford Large Network Dataset Collection
- URL: http://snap.stanford.edu/data/
-
Stanford Longitudinal Network Data Sources
- URL: http://stanford.edu/group/sonia/dataSources/index.html
-
The Koblenz Network Collection
- URL: http://konect.uni-koblenz.de/
-
The Laboratory for Web Algorithmics (UNIMI)
- URL: http://law.di.unimi.it/datasets.php
-
The Nexus Network Repository
- URL: http://nexus.igraph.org/
-
UCI Network Data Repository
- URL: https://networkdata.ics.uci.edu/resources.php
-
UFL sparse matrix collection
- URL: http://www.cise.ufl.edu/research/sparse/matrices/
-
WSU Graph Database
- URL: http://www.eecs.wsu.edu/mgd/gdb.html
ComputerNetworks
-
3.5B Web Pages from CommonCrawl 2012
- URL: http://www.bigdatanews.com/profiles/blogs/big-data-set-3-5-billion-web-pages-made-available-for-all-of-us
-
53.5B Web clicks of 100K users in Indiana Univ.
- URL: http://cnets.indiana.edu/groups/nan/webtraffic/click-dataset/
-
CAIDA Internet Datasets
- URL: http://www.caida.org/data/overview/
-
CRAWDAD Wireless datasets from Dartmouth Univ.
- URL: https://crawdad.cs.dartmouth.edu/
-
ClueWeb09 - 1B web pages
- URL: http://lemurproject.org/clueweb09/
-
ClueWeb12 - 733M web pages
- URL: http://lemurproject.org/clueweb12/
-
CommonCrawl Web Data over 7 years
- URL: http://commoncrawl.org/the-data/get-started/
-
Criteo click-through data
- URL: http://labs.criteo.com/2015/03/criteo-releases-its-new-dataset/
-
Internet-Wide Scan Data Repository
- URL: https://scans.io/
-
OONI: Open Observatory of Network Interference - Internet censorship data
- URL: https://ooni.torproject.org/data/
-
Open Mobile Data by MobiPerf
- URL: https://console.developers.google.com/storage/openmobiledata_public/
-
Rapid7 Sonar Internet Scans
- URL: https://sonar.labs.rapid7.com/
-
UCSD Network Telescope, IPv4 /8 net
- URL: http://www.caida.org/projects/network_telescope/
DataChallenges
-
Bruteforce Database
- URL: https://github.com/duyetdev/bruteforce-database
-
Challenges in Machine Learning
- URL: http://www.chalearn.org/
-
CrowdANALYTIX dataX
- URL: http://data.crowdanalytix.com
-
D4D Challenge of Orange
- URL: http://www.d4d.orange.com/en/home
-
DrivenData Competitions for Social Good
- URL: http://www.drivendata.org/
-
ICWSM Data Challenge (since 2009)
- URL: http://icwsm.cs.umbc.edu/
-
KDD Cup by Tencent 2012
- URL: http://www.kddcup2012.org/
-
Kaggle Competition Data
- URL: https://www.kaggle.com/
-
Localytics Data Visualization Challenge
- URL: https://github.com/localytics/data-viz-challenge
-
Netflix Prize
- URL: http://netflixprize.com/leaderboard.html
-
Space Apps Challenge
- URL: https://2015.spaceappschallenge.org
-
Telecom Italia Big Data Challenge
- URL: https://dandelion.eu/datamine/open-big-data/
-
TravisTorrent Dataset - MSR2017 Mining Challenge
- URL: https://travistorrent.testroots.org/
-
TunedIT - Data mining & machine learning data sets, algorithms, challenges
- URL: http://tunedit.org/challenges/
-
Yelp Dataset Challenge
- URL: http://www.yelp.com/dataset_challenge
EarthScience
-
AQUASTAT - Global water resources and uses
- URL: http://www.fao.org/nr/water/aquastat/data/query/index.html?lang=en
-
BODC - marine data of ~22K vars
- URL: https://www.bodc.ac.uk/data/
-
EOSDIS - NASAs earth observing system data
- URL: http://sedac.ciesin.columbia.edu/data/sets/browse
-
Earth Models
- URL: http://www.earthmodels.org/
-
Integrated Marine Observing System (IMOS) - roughly 30TB of ocean measurements
- URL: https://imos.aodn.org.au/
-
Marinexplore - Open Oceanographic Data
- URL: http://marinexplore.org/
-
Smithsonian Institution Global Volcano and Eruption Database
- URL: http://volcano.si.edu/
-
USGS Earthquake Archives
- URL: http://earthquake.usgs.gov/earthquakes/search/
Economics
-
American Economic Association (AEA)
- URL: https://www.aeaweb.org/resources/data
-
EconData from UMD
- URL: http://inforumweb.umd.edu/econdata/econdata.html
-
Economic Freedom of the World Data
- URL: http://www.freetheworld.com/datasets_efw.html
-
Historical MacroEconomic Statistics
- URL: http://www.historicalstatistics.org/
-
INFORUM - Interindustry Forecasting at the University of Maryland
- URL:




