Awesome Public Datasets
收藏github2019-07-13 更新2024-05-31 收录
下载链接:
https://github.com/Abd-Elrazek/awesome-public-datasets
下载链接
链接失效反馈官方服务:
资源简介:
一个主题中心的高质量公开数据集列表,这些数据集来自公共领域。
A high-quality open dataset list from the public domain, centered around a specific theme.
创建时间:
2018-12-27
原始信息汇总
数据集概述
农业
- Hyperspectral benchmark dataset on soil moisture: https://doi.org/10.5281/zenodo.1227837
- U.S. Department of Agricultures Nutrient Database: https://www.ars.usda.gov/northeast-area/beltsville-md/beltsville-human-nutrition-research-center/nutrient-data-laboratory/docs/sr28-download-files/
- U.S. Department of Agricultures PLANTS Database: http://www.plants.usda.gov/dl_all.html
生物学
- 1000 Genomes: http://www.1000genomes.org/data
- American Gut (Microbiome Project): https://github.com/biocore/American-Gut
- Broad Bioimage Benchmark Collection (BBBC): https://www.broadinstitute.org/bbbc
- Broad Cancer Cell Line Encyclopedia (CCLE): http://www.broadinstitute.org/ccle/home
- Cell Image Library: http://www.cellimagelibrary.org
- Complete Genomics Public Data: http://www.completegenomics.com/public-data/69-genomes/
- EBI ArrayExpress: http://www.ebi.ac.uk/arrayexpress/
- EBI Protein Data Bank in Europe: http://www.ebi.ac.uk/pdbe/emdb/index.html/
- ENCODE project: https://www.encodeproject.org
- Electron Microscopy Pilot Image Archive (EMPIAR): http://www.ebi.ac.uk/pdbe/emdb/empiar/
- Ensembl Genomes: http://ensemblgenomes.org/info/genomes
- Gene Expression Omnibus (GEO): http://www.ncbi.nlm.nih.gov/geo/
- Gene Ontology (GO): http://geneontology.org/page/download-annotations
- Global Biotic Interactions (GloBI): https://github.com/jhpoelen/eol-globi-data/wiki#accessing-species-interaction-data
- Harvard Medical School (HMS) LINCS Project: http://lincs.hms.harvard.edu
- Human Genome Diversity Project: http://www.hagsc.org/hgdp/files.html
- Human Microbiome Project (HMP): http://www.hmpdacc.org/reference_genomes/reference_genomes.php
- ICOS PSP Benchmark: http://ico2s.org/datasets/psp_benchmark.html
- International HapMap Project: http://hapmap.ncbi.nlm.nih.gov/downloads/index.html.en
- Journal of Cell Biology DataViewer: http://jcb-dataviewer.rupress.org
- KEGG: http://www.genome.jp/kegg/
- MIT Cancer Genomics Data: http://www.broadinstitute.org/cgi-bin/cancer/datasets.cgi
- NCBI Proteins: http://www.ncbi.nlm.nih.gov/guide/proteins/#databases
- NCBI Taxonomy: http://www.ncbi.nlm.nih.gov/taxonomy
- NCI Genomic Data Commons: https://gdc.cancer.gov/access-data/gdc-data-portal
- NIH Microarray data: ftp://ftp.ncbi.nih.gov/pub/geo/DATA/supplementary/series/GSE6532/
- OpenSNP genotypes data: https://opensnp.org/
- Pathguid - Protein-Protein Interactions Catalog: http://www.pathguide.org/
- Protein Data Bank: http://www.rcsb.org/
- Psychiatric Genomics Consortium: https://www.med.unc.edu/pgc/downloads
- PubChem Project: https://pubchem.ncbi.nlm.nih.gov/
- PubGene (now Coremine Medical): https://www.coremine.com/
- Sanger Catalogue of Somatic Mutations in Cancer (COSMIC): http://cancer.sanger.ac.uk/cosmic
- Sanger Genomics of Drug Sensitivity in Cancer Project (GDSC): http://www.cancerrxgene.org/
- Sequence Read Archive(SRA): http://www.ncbi.nlm.nih.gov/Traces/sra/
- Stanford Microarray Data: http://smd.stanford.edu/
- Stowers Institute Original Data Repository: http://www.stowers.org/research/publications/odr
- Systems Science of Biological Dynamics (SSBD) Database: http://ssbd.qbic.riken.jp
- The Cancer Genome Atlas (TCGA), available via Broad GDAC: https://gdac.broadinstitute.org/
- The Catalogue of Life: http://www.catalogueoflife.org/content/annual-checklist-archive
- The Personal Genome Project: http://www.personalgenomes.org/
- UCSC Public Data: http://hgdownload.soe.ucsc.edu/downloads.html
- UniGene: http://www.ncbi.nlm.nih.gov/unigene
- Universal Protein Resource (UnitProt): http://www.uniprot.org/downloads
气候+天气
- Actuaries Climate Index: http://actuariesclimateindex.org/data/
- Australian Weather: http://www.bom.gov.au/climate/dwo/
- Aviation Weather Center - Consistent, timely and accurate weather: https://aviationweather.gov/adds/dataserver
- Brazilian Weather - Historical data (In Portuguese) - Data related to: http://sinda.crn.inpe.br/PCD/SITE/novo/site/historico/index.php
- Canadian Meteorological Centre: http://weather.gc.ca/grib/index_e.html
- Climate Data from UEA (updated monthly): http://www.cru.uea.ac.uk/data/
- European Climate Assessment & Dataset: http://eca.knmi.nl/
- Global Climate Data Since 1929: http://en.tutiempo.net/climate
- NASA Global Imagery Browse Services: https://wiki.earthdata.nasa.gov/display/GIBS
- NOAA Bering Sea Climate: http://www.beringclimate.noaa.gov/
- NOAA Climate Datasets: http://www.ncdc.noaa.gov/data-access/quick-links
- NOAA Realtime Weather Models: http://www.ncdc.noaa.gov/data-access/model-data/model-datasets/numerical-weather-prediction
- NOAA SURFRAD Meteorology and Radiation Datasets: https://www.esrl.noaa.gov/gmd/grad/stardata.html
- The World Bank Open Data Resources for Climate Change: http://data.worldbank.org/developers/climate-data-api
- UEA Climatic Research Unit: http://www.cru.uea.ac.uk/data
- WU Historical Weather Worldwide: https://www.wunderground.com/history/index.html
- WorldClim - Global Climate Data: http://www.worldclim.org
复杂网络
- AMiner Citation Network Dataset: http://aminer.org/citation
- CrossRef DOI URLs: https://archive.org/details/doi-urls
- DBLP Citation dataset: https://kdl.cs.umass.edu/display/public/DBLP
- DIMACS Road Networks Collection: http://www.dis.uniroma1.it/challenge9/download.shtml
- NBER Patent Citations: http://nber.org/patents/
- NIST complex networks data collection: http://math.nist.gov/~RPozo/complex_datasets.html
- Network Repository with Interactive Exploratory Analysis Tools: http://networkrepository.com/
- Protein-protein interaction network: http://vlado.fmf.uni-lj.si/pub/networks/data/bio/Yeast/Yeast.htm
- PyPI and Maven Dependency Network: https://ogirardot.wordpress.com/2013/01/31/sharing-pypimaven-dependency-data/
- Scopus Citation Database: https://www.elsevier.com/solutions/scopus
- Small Network Data: http://www-personal.umich.edu/~mejn/netdata/
- Stanford GraphBase: http://www3.cs.stonybrook.edu/~algorith/implement/graphbase/implement.shtml
- Stanford Large Network Dataset Collection: http://snap.stanford.edu/data/
- Stanford Longitudinal Network Data Sources: http://stanford.edu/group/sonia/dataSources/index.html
- The Koblenz Network Collection: http://konect.uni-koblenz.de/
- The Laboratory for Web Algorithmics (UNIMI): http://law.di.unimi.it/datasets.php
- UCI Network Data Repository: https://networkdata.ics.uci.edu/resources.php
- UFL sparse matrix collection: http://www.cise.ufl.edu/research/sparse/matrices/
- WSU Graph Database: http://www.eecs.wsu.edu/mgd/gdb.html
计算机网络
- 3.5B Web Pages from CommonCrawl 2012: http://www.bigdatanews.com/profiles/blogs/big-data-set-3-5-billion-web-pages-made-available-for-all-of-us
- 53.5B Web clicks of 100K users in Indiana Univ.: http://cnets.indiana.edu/groups/nan/webtraffic/click-dataset/
- CAIDA Internet Datasets: http://www.caida.org/data/overview/
- CRAWDAD Wireless datasets from Dartmouth Univ.: https://crawdad.cs.dartmouth.edu/
- ClueWeb09 - 1B web pages: http://lemurproject.org/clueweb09/
- ClueWeb12 - 733M web pages: http://lemurproject.org/clueweb12/
- CommonCrawl Web Data over 7 years: http://commoncrawl.org/the-data/get-started/
- Criteo click-through data: http://labs.criteo.com/2015/03/criteo-releases-its-new-dataset/
- Internet-Wide Scan Data Repository: https://scans.io/
- OONI: Open Observatory of Network Interference - Internet censorship data: https://ooni.torproject.org/data/
- Open Mobile Data by MobiPerf: https://console.developers.google.com/storage/openmobiledata_public/
- The Peer-to-Peer Trace Archive - Real-world measurements play a key role: http://p2pta.ewi.tudelft.nl/
- Rapid7 Sonar Internet Scans: https://sonar.labs.rapid7.com/
- UCSD Network Telescope, IPv4 /8 net: http://www.caida.org/projects/network_telescope/
数据挑战
- Bruteforce Database: https://github.com/duyetdev/bruteforce-database
- Challenges in Machine Learning: http://www.chalearn.org/
- CrowdANALYTIX dataX: http://data.crowdanalytix.com
- D4D Challenge of Orange: http://www.d4d.orange.com/en/home
- DrivenData Competitions for Social Good: http://www.drivendata.org/
- ICWSM Data Challenge (since 2009): https://www.icwsm.org/2018/datasets/datasets/#obtaining
- KDD Cup by Tencent 2012: http://www.kddcup2012.org/
- Kaggle Competition Data: https://www.kaggle.com/
- Localytics Data Visualization Challenge: https://github.com/localytics/data-viz-challenge
- Netflix Prize: http://netflixprize.com/leaderboard.html
- Space Apps Challenge: https://2015.spaceappschallenge.org/
- Telecom Italia Big Data Challenge: https://dandelion.eu/datamine/open-big-data/
- TravisTorrent Dataset - MSR2017 Mining Challenge: https://travistorrent.testroots.org/
- TunedIT - Data mining & machine learning data sets, algorithms, challenges: http://tunedit.org/challenges/
- Yelp Dataset Challenge: http://www.yelp.com/dataset_challenge
地球科学
- AQUASTAT - Global water resources and uses: http://www.fao.org/nr/water/aquastat/data/query/index.html?lang=en
- BODC - marine data of ~22K vars: https://www.bodc.ac.uk/data/
- EOSDIS - NASAs earth observing system data: http://sedac.ciesin.columbia.edu/data/sets/browse
- Earth Models: http://www.earthmodels.org/
- Integrated Marine Observing System (IMOS) - roughly 30TB of ocean measurements: https://imos.aodn.org.au/
- Marinexplore - Open Oceanographic Data: http://marinexplore.org/
- Alabama Real-Time Coastal Observing System: http://mymobilebay.com/
- National Estuarine Research Reserves System-Wide Monitoring Program -: http://nerrsdata.org/
- Smithsonian Institution Global Volcano and Eruption Database: http://volcano.si.edu/
- USGS Earthquake Archives: http://earthquake.usgs.gov/earthquakes/search/
经济学
- American Economic Association (AEA): https://www.aeaweb.org/resources/data
- EconData from UMD: http://inforumweb.umd.edu/econdata/econdata.html
- Economic Freedom of the World Data: http://www.freetheworld.com/datasets_efw.html
- Historical MacroEconomic Statistics: http://www.historicalstatistics.org/
- INFORUM - Interindustry Forecasting at the University of Maryland: http://inforumweb.umd.edu/
- International Economics Database: http://widukind.cepremap.org/
- International Trade Statistics: http://www.econistatistics.co.za/
- Internet Product Code Database: http://www.upcdatabase.com/
- Joint External Debt Data Hub: http://www.jedh.org/
- Jon Haveman International Trade Data Links: http://www.macalester.edu/research/economics/PAGE/HAVEMAN/Trade.Resources/TradeData.html
- OpenCorporates Database of Companies in the World: https://opencorporates.com/
- Our World in Data: http://ourworldindata.org/
- SciencesPo World Trade Gravity Datasets: http://econ.sciences-po.fr/thierry-mayer/data
- The Atlas of Economic Complexity: http://atlas.cid.harvard.edu/
- The Center for International Data: http://cid.econ.ucdavis.edu/
- The Observatory of Economic Complexity: http://atlas.media.mit.edu/en/
- UN Commodity Trade Statistics: http://comtrade.un.org/db/
- UN Human Development Reports: http://hdr.undp.org/en/
教育
- College Scorecard Data: https://collegescorecard.ed.gov/data/
- Student Data from Free Code Camp: https://github.com/freeCodeCamp/open-data
能源
- AMPds: http://ampds.org/
- BLUEd: http://nilm.cmubi.org/
- COMBED: http://combed.github.io/
- ECO: http://www.vs.inf.ethz.ch/res/show.html?what=eco-data
- EIA: http://www.eia.gov/electricity/data/eia923/
- Global Power Plant Database - The Global Power Plant Database is a: http://datasets.wri.org/dataset/globalpowerplantdatabase
- HES - Household Electricity Study, UK: http://randd.defra.gov.uk/Default.aspx?Menu=Menu&Module=More&Location=None&ProjectID=17359&FromSearch=Y&Publisher=1&SearchText=EV0702&SortString=ProjectCode&SortOrder=Asc&Paging=10#Description
- HFED: http://hfed.github.io/
- **PLAID - The Plug
搜集汇总
数据集介绍

构建方式
Awesome Public Datasets 是一个由社区驱动的开源项目,旨在收集和整理互联网上各种主题的公共数据集。该数据集的构建主要通过自动化脚本从多个来源抓取信息,并通过社区贡献进行维护和更新。
特点
该数据集的特点在于其覆盖领域广泛,包含农业、生物学、气候、复杂网络、计算机网络、数据挑战、地球科学、经济学、教育、能源、金融、GIS等多个领域的数据集。此外,大部分数据集都是免费的,并且提供了多种数据格式和访问方式。
使用方法
用户可以通过数据集的GitHub页面获取数据,页面中包含了数据集的详细描述和链接。用户应当遵守数据的使用条款,并根据需要选择合适的数据格式进行下载。部分数据集可能需要特定的软件或工具来处理。
背景与挑战
背景概述
Awesome Public Datasets是一个收集和整理高质量公共数据集的仓库,由sindresorhus维护。该数据集涵盖了多个领域,如农业、生物学、气候与天气、复杂网络、计算机网络、数据挑战、地球科学、经济学、教育、能源、金融、GIS等。这些数据集主要来源于博客、回答和用户响应,大部分免费,部分可能需要付费。该项目的目标是提供一个全面、高质量的公共数据集列表,以促进数据共享和科学研究。
当前挑战
尽管Awesome Public Datasets提供了一个广泛的数据集列表,但在构建过程中仍面临一些挑战。首先,数据集的质量和可靠性需要持续验证和更新。其次,数据集的整理和分类工作繁琐且容易出错。此外,部分数据集的获取可能存在法律和版权的限制。最后,随着数据集的增多,如何高效地管理和维护这个列表成为一个挑战。
常用场景
经典使用场景
Awesome Public Datasets 是一个集成了众多领域公共数据集的仓库,其经典使用场景在于为研究人员、数据科学家和开发者提供了一个全面的数据资源搜索和发现平台。用户可以在此找到从生物学、气候学、计算机科学到经济学、教育学和地理信息系统等多个学科领域的高质量数据集,以支持其研究和开发工作。
衍生相关工作
基于 Awesome Public Datasets,已经衍生出了一系列相关的工作,包括但不限于数据集的整理、清洗和集成工作,以及利用这些数据集进行的各种学术研究和应用开发。这些衍生工作进一步扩展了原始数据集的应用范围和影响力。
数据集最近研究
最新研究方向
Awesome Public Datasets数据集的近期研究方向主要集中在数据挖掘、数据分析和复杂网络等领域。学者们利用这些数据集进行前沿研究,如基于地理信息系统(GIS)的数据可视化分析,在金融领域中应用时间序列分析预测市场趋势,以及在能源领域中对智能电网数据进行分析,以优化能源分配和消费。这些研究不仅推动了数据科学领域的发展,也对相关行业产生了深远的影响。
以上内容由遇见数据集搜集并总结生成



