Awesome Public Datasets
收藏github2017-04-10 更新2024-05-31 收录
下载链接:
https://github.com/panisson/awesome-public-datasets
下载链接
链接失效反馈官方服务:
资源简介:
一个包含高质量公开数据集的精选列表,涵盖多个领域,如农业、生物学等。
A curated list of high-quality public datasets spanning various domains, such as agriculture, biology, and more.
创建时间:
2016-02-10
原始信息汇总
数据集概述
农业
- U.S. Department of Agricultures PLANTS Database
- 链接: http://www.plants.usda.gov/dl_all.html
生物学
- 1000 Genomes
- 链接: http://www.1000genomes.org/data
- American Gut (Microbiome Project)
- 链接: https://github.com/biocore/American-Gut
- Broad Cancer Cell Line Encyclopedia (CCLE)
- 链接: http://www.broadinstitute.org/ccle/home
- Cell Image Library
- 链接: http://www.cellimagelibrary.org
- Collaborative Research in Computational Neuroscience (CRCNS)
- 链接: http://crcns.org/data-sets
- Complete Genomics Public Data
- 链接: http://www.completegenomics.com/public-data/69-genomes/
- EBI ArrayExpress
- 链接: http://www.ebi.ac.uk/arrayexpress/
- EBI Protein Data Bank in Europe
- 链接: http://www.ebi.ac.uk/pdbe/emdb/index.html/
- ENCODE project
- 链接: https://www.encodeproject.org
- Ensembl Genomes
- 链接: http://ensemblgenomes.org/info/genomes
- Gene Expression Omnibus (GEO)
- 链接: http://www.ncbi.nlm.nih.gov/geo/
- Gene Ontology (GO)
- 链接: http://geneontology.org/page/download-annotations
- Global Biotic Interations (GloBI)
- 链接: https://github.com/jhpoelen/eol-globi-data/wiki#accessing-species-interaction-data
- Harvard Medical School (HMS) LINCS Project
- 链接: http://lincs.hms.harvard.edu
- Human Genome Diversity Project
- 链接: http://www.hagsc.org/hgdp/files.html
- Human Microbiome Project (HMP)
- 链接: http://www.hmpdacc.org/reference_genomes/reference_genomes.php
- ICOS PSP Benchmark
- 链接: http://ico2s.org/datasets/psp_benchmark.html
- International HapMap Project
- 链接: http://hapmap.ncbi.nlm.nih.gov/downloads/index.html.en
- Journal of Cell Biology DataViewer
- 链接: http://jcb-dataviewer.rupress.org
- MIT Cancer Genomics Data
- 链接: http://www.broadinstitute.org/cgi-bin/cancer/datasets.cgi
- NCBI Proteins
- 链接: http://www.ncbi.nlm.nih.gov/guide/proteins/#databases
- NCBI Taxonomy
- 链接: http://www.ncbi.nlm.nih.gov/taxonomy
- NeuroData
- 链接: http://neurodata.io
- NIH Microarray data
- 链接: http://bit.do/VVW6 或 ftp://ftp.ncbi.nih.gov/pub/geo/DATA/supplementary/series/GSE6532/
- OpenSNP genotypes data
- 链接: https://opensnp.org/
- Pathguid - Protein-Protein Interactions Catalog
- 链接: http://www.pathguide.org/
- Protein Data Bank
- 链接: http://www.rcsb.org/
- PubChem Project
- 链接: https://pubchem.ncbi.nlm.nih.gov/
- PubGene (now Coremine Medical)
- 链接: http://www.pubgene.org/
- Sanger Catalogue of Somatic Mutations in Cancer (COSMIC)
- 链接: http://cancer.sanger.ac.uk/cosmic
- Sanger Genomics of Drug Sensitivity in Cancer Project (GDSC)
- 链接: http://www.cancerrxgene.org/
- Sequence Read Archive(SRA)
- 链接: http://www.ncbi.nlm.nih.gov/Traces/sra/
- Stanford Microarray Data
- 链接: http://smd.stanford.edu/
- Stowers Institute Original Data Repository
- 链接: http://www.stowers.org/research/publications/odr
- Systems Science of Biological Dynamics (SSBD) Database
- 链接: http://ssbd.qbic.riken.jp
- Temple University Hospital EEG Database
- 链接: https://www.nedcdata.org/drupal/node/12
- The Cancer Genome Atlas (TCGA), available via Broad GDAC
- 链接: https://gdac.broadinstitute.org/
- The Catalogue of Life
- 链接: http://www.catalogueoflife.org/content/annual-checklist-archive
- The Personal Genome Project
- 链接: http://www.personalgenomes.org/ 或 https://my.pgp-hms.org/public_genetic_data
- UCSC Public Data
- 链接: http://hgdownload.soe.ucsc.edu/downloads.html
- Universal Protein Resource (UnitProt)
- 链接: http://www.uniprot.org/downloads
- UniGene
- 链接: http://www.ncbi.nlm.nih.gov/unigene
气候/天气
- Australian Weather
- 链接: http://www.bom.gov.au/climate/dwo/
- Brazilian Weather - Historical data (In Portuguese)
- 链接: http://sinda.crn2.inpe.br/PCD/SITE/novo/site/
- Canadian Meteorological Centre
- 链接: http://weather.gc.ca/grib/index_e.html
- Climate Data from UEA (updated monthly)
- 链接: https://crudata.uea.ac.uk/cru/data/temperature/ 和 ftp://ftp.cmdl.noaa.gov/
- European Climate Assessment & Dataset
- 链接: http://eca.knmi.nl/
- Global Climate Data Since 1929
- 链接: http://en.tutiempo.net/climate
- NASA Global Imagery Browse Services
- 链接: https://wiki.earthdata.nasa.gov/display/GIBS
- NOAA Bering Sea Climate
- 链接: http://www.beringclimate.noaa.gov/
- NOAA Climate Datasets
- 链接: http://www.ncdc.noaa.gov/data-access/quick-links
- NOAA Realtime Weather Models
- 链接: http://www.ncdc.noaa.gov/data-access/model-data/model-datasets/numerical-weather-prediction
- The World Bank Open Data Resources for Climate Change
- 链接: http://data.worldbank.org/developers/climate-data-api
- UEA Climatic Research Unit
- 链接: http://www.cru.uea.ac.uk/data
- WorldClim - Global Climate Data
- 链接: http://www.worldclim.org
- WU Historical Weather Worldwide
- 链接: https://www.wunderground.com/history/index.html
复杂网络
- CrossRef DOI URLs
- 链接: https://archive.org/details/doi-urls
- DBLP Citation dataset
- 链接: https://kdl.cs.umass.edu/display/public/DBLP
- NBER Patent Citations
- 链接: http://nber.org/patents/
- NIST complex networks data collection
- 链接: http://math.nist.gov/~RPozo/complex_datasets.html
- Protein-protein interaction network
- 链接: http://vlado.fmf.uni-lj.si/pub/networks/data/bio/Yeast/Yeast.htm
- PyPI and Maven Dependency Network
- 链接: https://ogirardot.wordpress.com/2013/01/31/sharing-pypimaven-dependency-data/
- Scopus Citation Database
- 链接: https://www.elsevier.com/solutions/scopus
- Small Network Data
- 链接: http://www-personal.umich.edu/~mejn/netdata/
- Stanford GraphBase (Steven Skiena)
- 链接: http://www3.cs.stonybrook.edu/~algorith/implement/graphbase/implement.shtml
- Stanford Large Network Dataset Collection
- 链接: http://snap.stanford.edu/data/
- Stanford Longitudinal Network Data Sources
- 链接: http://stanford.edu/group/sonia/dataSources/index.html
- The Koblenz Network Collection
- 链接: http://konect.uni-koblenz.de/
- The Laboratory for Web Algorithmics (UNIMI)
- 链接: http://law.di.unimi.it/datasets.php
- The Nexus Network Repository
- 链接: http://nexus.igraph.org/
- UCI Network Data Repository
- 链接: https://networkdata.ics.uci.edu/resources.php
- UFL sparse matrix collection
- 链接: http://www.cise.ufl.edu/research/sparse/matrices/
- WSU Graph Database
- 链接: http://www.eecs.wsu.edu/mgd/gdb.html
计算机网络
- 3.5B Web Pages from CommonCraw 2012
- 链接: http://www.bigdatanews.com/profiles/blogs/big-data-set-3-5-billion-web-pages-made-available-for-all-of-us
- 53.5B Web clicks of 100K users in Indiana Univ.
- 链接: http://cnets.indiana.edu/groups/nan/webtraffic/click-dataset/
- CAIDA Internet Datasets
- 链接: http://www.caida.org/data/overview/
- ClueWeb09 - 1B web pages
- 链接: http://lemurproject.org/clueweb09/
- ClueWeb12 - 733M web pages
- 链接: http://lemurproject.org/clueweb12/
- CommonCrawl Web Data over 7 years
- 链接: http://commoncrawl.org/the-data/get-started/
- CRAWDAD Wireless datasets from Dartmouth Univ.
- 链接: https://crawdad.cs.dartmouth.edu/
- Criteo click-through data
- 链接: http://labs.criteo.com/2015/03/criteo-releases-its-new-dataset/
- Open Mobile Data by MobiPerf
- 链接: https://console.developers.google.com/storage/openmobiledata_public/
- UCSD Network Telescope, IPv4 /8 net
- 链接: http://www.caida.org/projects/network_telescope/
上下文数据
- Context-aware data sets from five domains
- 链接: http://students.depaul.edu/~yzheng8/DataSets.html#Data 或 https://github.com/irecsys/CARSKit/tree/master/context-aware_data_sets
数据挑战
- Challenges in Machine Learning
- 链接: http://www.chalearn.org/
- CrowdANALYTIX dataX
- 链接: http://data.crowdanalytix.com
- D4D Challenge of Orange
- 链接: http://www.d4d.orange.com/en/home
- DrivenData Competitions for Social Good
- 链接: http://www.drivendata.org/
- ICWSM Data Challenge (since 2009)
- 链接: http://icwsm.cs.umbc.edu/
- Kaggle Competition Data
- 链接: https://www.kaggle.com/
- KDD Cup by Tencent 2012
- 链接: http://www.kddcup2012.org/
- Localytics Data Visualization Challenge
- 链接: https://github.com/localytics/data-viz-challenge
- Netflix Prize
- 链接: http://www.netflixprize.com/leaderboard
- Space Apps Challenge
- 链接: https://2015.spaceappschallenge.org/
- Telecom Italia Big Data Challenge
- 链接: https://dandelion.eu/datamine/open-big-data/
- Yelp Dataset Challenge
- 链接: http://www.yelp.com/dataset_challenge
经济学
- American Economic Ass (AEA)
- 链接: https://www.aeaweb.org/RFE/toc.php?show=complete
- EconData from UMD
- 链接: http://inforumweb.umd.edu/econdata/econdata.html
- Economic Freedom of the World Data
- 链接: http://www.freetheworld.com/datasets_efw.html
- Historical MacroEconomic Statistics
- 链接: http://www.historicalstatistics.org/
- International Trade Statistics
- 链接: http://www.econostatistics.co.za/
- Internet Product Code Database
- 链接: http://www.upcdatabase.com/
- Joint External Debt Data Hub
- 链接: http://www.jedh.org/
- Jon Haveman International Trade Data Links
- 链接: http://www.macalester.edu/research/economics/PAGE/HAVEMAN/Trade.Resources/TradeData.html
- OpenCorporates Database of Companies in the World
- 链接: https://opencorporates.com/
- Our World in Data
- 链接: http://ourworldindata.org/
- SciencesPo World Trade Gravity Datasets
- 链接: http://econ.sciences-po.fr/thierry-mayer/data
- The Atlas of Economic Complexity
- 链接: http://atlas.cid.harvard.edu/
- The Center for International Data
- 链接: http://cid.econ.ucdavis.edu/
- The Observatory of Economic Complexity
- 链接: http://atlas.media.mit.edu/en/
- UN Commodity Trade Statistics
- 链接: http://comtrade.un.org/db/
- UN Human Development Reports
- 链接: http://hdr.undp.org/en/
教育
- Student Data from Free Code Camp
- 链接: http://academictorrents.com/details/030b10dad0846b5aecc3905692890fb02404adbf
能源
- AMPds
- 链接: http://ampds.org/
- BLUEd
- 链接: http://nilm.cmubi.org/
- COMBED
- 链接: http://combed.github.io/
- Dataport
- 链接: https://dataport.pecanstreet.org/
- ECO
- 链接: http://www.vs.inf.ethz.ch/res/show.html?what=eco-data
- EIA
- 链接: http://www.eia.gov/electricity/data/eia923/
- HFED
- 链接: http://hfed.github.io/
- iAWE
- 链接: http://iawe.github.io/
- Plaid
- 链接: http://plaidplug.com/
- REDD
- 链接: http://redd.csail.mit.edu/
- UK-Dale
- 链接: http://www.doc.ic.ac
搜集汇总
数据集介绍

构建方式
Awesome Public Datasets 是一个收集和整理自博客、回答和用户响应的公共数据集列表。该数据集主要通过互联网资源进行整合,涵盖了多个领域的数据集,其中大部分是免费的,但也有一些是收费的。
特点
该数据集的特点在于其多样性、全面性和开放性。它包含了农业、生物学、气候/天气、复杂网络、计算机网络、上下文数据、数据挑战、经济学、教育、能源、金融、地质学、地理空间/GIS、政府、健康护理等多个领域的公共数据集。这些数据集来源广泛,既有官方机构发布的数据,也有非官方的收集整理。
使用方法
用户可以通过GitHub页面浏览和搜索所需的数据集。每个数据集都有详细的描述和链接,方便用户直接访问和下载。用户在使用前应仔细阅读数据集的说明,了解数据集的来源、格式和使用条款。
背景与挑战
背景概述
Awesome Public Datasets是一个收集和整理自博客、回答和用户响应的公开数据集列表。该数据集多数是免费的,但也有一些不是。此列表的创建旨在为研究者提供方便,帮助他们在各自领域找到所需的数据资源。该数据集的创建时间是未知的,主要研究人员或机构是sindresorhus,核心研究问题是提供公开可用的数据集列表,对相关领域的影响力在于方便了数据获取和共享。
当前挑战
该数据集在构建过程中所遇到的挑战主要包括数据的收集和整理,由于数据来源多样,格式和内容也各不相同,因此需要花费大量时间和精力进行筛选和整理。此外,数据集的维护和更新也是一个挑战,需要持续关注各数据源的变化,并相应地更新列表。在解决的领域问题方面,该数据集的挑战在于如何有效地帮助研究者找到适用于他们研究的数据集,以及如何确保数据集的质量和可用性。
常用场景
经典使用场景
Awesome Public Datasets集成了多个领域的大量公共数据集,其经典使用场景主要集中于数据科学家、研究人员和开发者,他们可以利用这个资源库来寻找和分析特定领域的数据,例如生物信息学、气候学、复杂网络等,以支持他们的研究和开发工作。
衍生相关工作
基于这个数据集,已经衍生出许多相关的经典工作,包括但不限于在生物信息学领域对基因组数据的研究,在气候学领域对气候变化趋势的分析,以及在复杂网络领域对社交网络结构的探索。
数据集最近研究
最新研究方向
该数据集涵盖了多个领域的公共数据集,最新研究方向主要集中在数据的整合、分析和挖掘上,以支持不同领域的研究和应用。例如,在生物医学领域,研究人员可能专注于利用基因组和微生物组数据集进行疾病关联分析;在气候和天气领域,则可能关注于气候变化对生态系统和人类社会的影响研究;在复杂网络和计算机网络领域,研究可能集中在网络结构特性分析和大数据流量模式挖掘。这些研究对于推动相关领域的发展具有重要的意义和影响。
以上内容由遇见数据集搜集并总结生成



