遇见数据集

SMILE: Systematic Microbiome Intelligence for Lost Ecosystems - Ancient Oral Microbiome Database v1.0

收藏
Zenodo2026-04-06 更新2026-05-26 收录
官方服务:

资源简介:

SMILE (Systematic Microbiome Intelligence for Lost Ecosystems) is a curated relational database of ancient oral microbiome data derived from published archaeogenomic literature. Version 1.0 integrates data from 45 peer-reviewed publications (2014–2024), spanning 1,414 samples, 16,196 microbiome records, and 2,150 distinct taxa. Temporal coverage extends from approximately 102,000 BP (Pesturina Cave, Serbia — Neanderthal era) to the medieval period, with near-global geographic coverage (-175 to 178° longitude, -34 to 79° latitude). The corpus includes records for Bacteria (16,120), Archaea (53), Fungi (16), and Viruses (7), with 1,401 human host samples and 13 non-human host samples (bear, gorilla, reindeer). Authentication metric records (n=212) and methodological metadata entries (n=42) are included to support reproducibility and downstream filtering. Data were extracted using an LLM-assisted pipeline (Anthropic API: Haiku for pre/post-processing stages, Sonnet for structured extraction) applied to primary literature, with multilingual extraction coverage including English, Japanese, French, and Russian sources. All extraction scripts are available at https://github.com/satoru-bio/smile-pipeline. The database is structured as a PostgreSQL 16 relational schema with PostGIS extension. This deposition includes a full SQL dump (restorable to any PostgreSQL 16+ instance), six CSV table exports for tool-agnostic access, a dataset summary JSON, and a README with restoration instructions. Limitation: 262 records identified as figure-only in source publications remain undigitised in v1.0 and are documented as a known gap. These are distributed across approximately 20 DOIs and will be addressed in a subsequent release. This dataset supports research in palaeomicrobiology, ancient DNA, evolutionary medicine, antimicrobial resistance baseline reconstruction, and biosecurity preparedness.

SMILE(Systematic Microbiome Intelligence for Lost Ecosystems,即失落生态系统系统性微生物组智能数据库)是一款经人工精选的关系型数据库,收录了源自已发表古基因组学文献的古代口腔微生物组数据。1.0版本整合了2014年至2024年间45篇同行评议论文的数据,涵盖1414份样本、16196条微生物组记录以及2150个独特分类单元。 该数据集的时间覆盖范围约从102000年前(塞尔维亚佩斯图里纳洞穴,尼安德特人时代)至中世纪,地理覆盖范围接近全球(经度区间为-175°至178°,纬度区间为-34°至79°)。数据集包含细菌(16120条)、古菌(53条)、真菌(16条)与病毒(7条)的记录,其中1401份为人类宿主样本,13份为非人类宿主样本(熊、大猩猩、驯鹿)。数据库同时收录212条验证指标记录与42条方法学元数据条目,以支撑研究的可重复性与下游筛选工作。 本数据集采用大语言模型(Large Language Model)辅助的流程从原始文献中提取数据:使用Anthropic API的Haiku模型完成预处理与后处理环节,Sonnet模型负责结构化数据提取;该流程支持多语言文献提取,涵盖英语、日语、法语及俄语数据源。所有提取脚本均可通过https://github.com/satoru-bio/smile-pipeline获取。 该数据库采用搭载PostGIS扩展的PostgreSQL 16关系型架构。本次发布包含完整的SQL转储文件(可恢复至任意PostgreSQL 16及以上版本的数据库实例)、6个CSV格式的表格导出文件以支持工具无关性访问、一份数据集摘要JSON文件,以及一份包含恢复说明的README文档。 局限性说明:在1.0版本中,源文献中仅以图表形式呈现的262条记录尚未完成数字化,这一已知数据缺口已被记录在案。这些未数字化记录分布于约20个数字对象唯一标识符(Digital Object Identifier,DOI)对应的文献中,将在后续版本中予以补充完善。 本数据集可支撑古微生物学、古代DNA、进化医学、抗生素耐药性基线重建以及生物安全防范等领域的研究工作。

提供机构:
Zenodo
创建时间:
2026-04-06
二维码
社区交流群
二维码
科研交流群
商业服务