遇见数据集

severo/embellishments

收藏
Hugging Face2022-10-25 更新2024-03-04 收录
官方服务:

资源简介:

--- annotations_creators: - no-annotation license: - cc0-1.0 size_categories: - n<1K source_datasets: - original pretty_name: Digitised Books - Images identified as Embellishments. c. 1510 - c. 1900. JPG --- # Dataset Card for severo/embellishments ## Dataset Description - **Homepage:** [Digitised Books - Images identified as Embellishments - Homepage](https://bl.iro.bl.uk/concern/datasets/59d1aa35-c2d7-46e5-9475-9d0cd8df721e) - **Point of Contact:** [Sylvain Lesage](mailto:sylvain.lesage@huggingface.co) ### Dataset Summary This small dataset contains the thumbnails of the first 100 entries of [Digitised Books - Images identified as Embellishments. c. 1510 - c. 1900. JPG](https://bl.iro.bl.uk/concern/datasets/59d1aa35-c2d7-46e5-9475-9d0cd8df721e). It has been uploaded to the Hub to reproduce the tutorial by Daniel van Strien: [Using 🤗 datasets for image search](https://danielvanstrien.xyz/metadata/deployment/huggingface/ethics/huggingface-datasets/faiss/2022/01/13/image_search.html). ## Dataset Structure ### Data Instances A typical row contains an image thumbnail, its filename, and the year of publication of the book it was extracted from. An example looks as follows: ``` { 'fname': '000811462_05_000205_1_The Pictorial History of England being a history of the people as well as a hi_1855.jpg', 'year': '1855', 'path': 'embellishments/1855/000811462_05_000205_1_The Pictorial History of England being a history of the people as well as a hi_1855.jpg', 'img': ... } ``` ### Data Fields - `fname`: the image filename. - `year`: a string with the year of publication of the book from which the image has been extracted - `path`: local path to the image - `img`: a thumbnail of the image with a max height and width of 224 pixels ### Data Splits The dataset only contains 100 rows, in a single 'train' split. ## Dataset Creation ### Curation Rationale This dataset was chosen by Daniel van Strien for his tutorial [Using 🤗 datasets for image search](https://danielvanstrien.xyz/metadata/deployment/huggingface/ethics/huggingface-datasets/faiss/2022/01/13/image_search.html), which includes the code in Python to do it. ### Source Data #### Initial Data Collection and Normalization As stated on the British Library webpage: > The images were algorithmically gathered from 49,455 digitised books, equating to 65,227 volumes (25+ million pages), published between c. 1510 - c. 1900. The books cover a wide range of subject areas including philosophy, history, poetry and literature. The images are in .JPEG format.d BCP-47 code is `en`. #### Who are the source data producers? British Library, British Library Labs, Adrian Edwards (Curator), Neil Fitzgerald (Contributor ORCID) ### Annotations The dataset does not contain any additional annotations. #### Annotation process [N/A] #### Who are the annotators? [N/A] ### Personal and Sensitive Information [N/A] ## Considerations for Using the Data ### Social Impact of Dataset [N/A] ### Discussion of Biases [N/A] ### Other Known Limitations This is a toy dataset that aims at: - validating the process described in the tutorial [Using 🤗 datasets for image search](https://danielvanstrien.xyz/metadata/deployment/huggingface/ethics/huggingface-datasets/faiss/2022/01/13/image_search.html) by Daniel van Strien, - showing the [dataset viewer](https://huggingface.co/datasets/severo/embellishments/viewer/severo--embellishments/train) on an image dataset. ## Additional Information ### Dataset Curators The dataset was created by Sylvain Lesage at Hugging Face, to replicate the tutorial [Using 🤗 datasets for image search](https://danielvanstrien.xyz/metadata/deployment/huggingface/ethics/huggingface-datasets/faiss/2022/01/13/image_search.html) by Daniel van Strien. ### Licensing Information CC0 1.0 Universal Public Domain

提供机构:
severo
原始信息汇总

数据集概述

数据集名称

  • 名称: Digitised Books - Images identified as Embellishments. c. 1510 - c. 1900. JPG
  • 别名: severo/embellishments

数据集描述

数据集摘要

  • 内容: 包含100个图像缩略图,这些图像来自1510年至1900年间出版的书籍,被标识为装饰性图像。
  • 用途: 用于复制Daniel van Strien的教程Using 🤗 datasets for image search

数据集结构

数据实例

  • 组成: 每个实例包含图像缩略图、文件名、以及图像来源书籍的出版年份。

  • 示例:

    { fname: 000811462_05_000205_1_The Pictorial History of England being a history of the people as well as a hi_1855.jpg, year: 1855, path: embellishments/1855/000811462_05_000205_1_The Pictorial History of England being a history of the people as well as a hi_1855.jpg, img: ... }

数据字段

  • fname: 图像文件名。
  • year: 字符串,表示图像来源书籍的出版年份。
  • path: 图像的本地路径。
  • img: 图像缩略图,最大高度和宽度为224像素。

数据分割

  • 分割方式: 单一的train分割,共100行。

数据集创建

源数据

  • 来源: 从49,455本数字化书籍中算法收集,涵盖1510年至1900年间出版的书籍。
  • 格式: JPEG格式。
  • 数据生产者: 英国图书馆、英国图书馆实验室、Adrian Edwards (策展人)、Neil Fitzgerald (贡献者ORCID)。

数据集创建者

  • 创建者: Sylvain Lesage at Hugging Face

许可证

  • 许可证: CC0 1.0 Universal Public Domain
搜集汇总
数据集介绍
构建方式
在数字人文研究的浪潮中,图像资源的系统化整理成为学界关注的焦点。该数据集源自大英图书馆大规模数字化馆藏项目,从49,455册出版于1510年至1900年间的古籍中,通过算法自动提取了装饰性图像。作为原始数据集的缩略版,本数据集精选了前100条记录,每份样本均包含图像缩略图、原始文件名及其所属书籍的出版年份。图像尺寸被统一规范为最大宽高不超过224像素,以兼顾存储效率与内容辨识度。数据集的构建初衷在于为图像检索技术提供轻量级验证样本,其内容涵盖哲学、历史、诗歌与文学等多学科领域,展现了早期印刷品中装饰艺术的演变脉络。
特点
本数据集虽规模精巧,却承载着丰富的历史文献价值。其核心特色在于图像来源的跨时代性,覆盖了从文艺复兴至工业革命末期的四个世纪,为研究书籍装帧艺术的风格演变提供了珍贵样本。每条数据实例包含文件名、年份与路径三项元数据,配合标准化缩略图,形成了结构清晰的多维信息单元。尤为独特的是,该数据集作为教学演示工具,完美适配图像检索教程的技术需求,通过FAISS等索引库可实现高效的特征比对。其CC0公共领域许可协议确保了学术应用的零门槛,而单训练集划分的设计则简化了实验流程,适合快速原型开发与算法验证。
使用方法
该数据集在HuggingFace平台上的使用极为便捷,开发者可通过datasets库直接加载。典型应用流程包括:首先利用load_dataset函数获取完整数据,随后提取img字段中的图像张量进行特征编码。配合torchvision等计算机视觉库,可轻松生成图像嵌入向量。结合FAISS索引库,用户能构建基于余弦相似度的图像搜索引擎,实现跨年代装饰图案的语义检索。教程中提供的Python代码完整覆盖了从数据加载到可视化检索的全链路,开发者仅需调整查询图像路径即可复现实验。对于批量处理场景,数据集支持分片加载与内存优化,其轻量特性特别适合在Jupyter Notebook等交互环境中进行教学演示。
背景与挑战
背景概述
在数字人文与文化遗产数字化领域,大规模古籍图像资源的自动化处理与检索正成为备受关注的研究方向。severo/embellishments数据集由Hugging Face的Sylvain Lesage于2022年创建,源自大英图书馆(British Library)馆藏的49,455册数字化古籍,涵盖约1510年至1900年间出版的哲学、历史、诗歌与文学等多领域文献。该数据集聚焦于从逾2500万页古籍页面中经由算法自动提取的装饰性图像(embellishments),旨在为图像检索与元数据管理提供验证性样本。尽管仅包含100张缩略图,但其作为Daniel van Strien教程中图像搜索流程复现的核心载体,展示了如何利用Hugging Face Datasets与FAISS实现高效图像检索,对推动数字图书馆中非文本内容的可发现性具有示范意义。
当前挑战
该数据集面临的主要挑战体现在两个层面。首先,在领域问题层面,古籍装饰性图像的自动识别与分类长期受困于风格多样性与年代跨度带来的视觉异质性——图像涵盖木刻版画、铜版画、手绘装饰等类型,且不同出版年份的印刷质量差异显著,传统图像特征提取方法难以鲁棒地应对此类变体。其次,在构建过程中,数据源自算法从海量页面中自动抽取,缺乏人工校验,导致部分图像可能包含文本误判或非装饰性内容;此外,数据集仅提供缩略图与出版年份等基础元数据,未标注图像的具体装饰类别或语义属性,限制了其在细粒度检索任务中的直接应用潜力。
常用场景
经典使用场景
该数据集源自大英图书馆数字化馆藏,精选了1510年至1900年间出版的书籍中算法识别的装饰性图像缩略图。其经典使用场景在于为图像检索与特征提取提供轻量级测试平台,特别适合验证基于深度学习模型(如卷积神经网络)的视觉相似性搜索流程。通过结合Hugging Face Datasets库与FAISS向量索引,研究者可快速搭建从图像加载、嵌入生成到最近邻检索的端到端流水线,体现了数据驱动下数字人文与计算机视觉的交叉应用。
解决学术问题
该数据集主要解决了数字人文领域中大规模历史图像语义检索的入门门槛问题。传统上,对古籍插图进行自动化分类与检索需要复杂的预处理与模型训练,而本数据集提供了标准化、低维度的图像缩略图,使研究者能聚焦于检索算法本身的性能评估。它推动了无标注图像在零样本学习场景下的应用探索,并为验证图像嵌入质量与跨模态对齐方法提供了可复现的基准,从而促进了文化遗产数字化中高效检索技术的学术发展。
衍生相关工作
该数据集衍生的经典工作包括Daniel van Strien撰写的《Using 🤗 datasets for image search》教程,该教程系统展示了如何利用Hugging Face生态实现图像检索,成为许多数字人文项目的技术起点。后续研究在此基础上扩展了多模态检索范式,例如结合文本元数据(如出版年份)进行条件过滤的混合查询。此外,该数据集被用于验证轻量级嵌入模型(如CLIP)在历史图像上的迁移效果,推动了面向特定领域(如古籍装饰图案)的微调策略研究,并催生了针对低资源图像集的特征蒸馏方法。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务