遇见数据集

fphool/usps

收藏
Hugging Face2026-04-20 更新2026-04-26 收录
官方服务:

资源简介:

--- dataset_info: features: - name: image dtype: image - name: label dtype: class_label: names: '0': '0' '1': '1' '2': '2' '3': '3' '4': '4' '5': '5' '6': '6' '7': '7' '8': '8' '9': '9' splits: - name: train num_bytes: 2194749.625 num_examples: 7291 - name: test num_bytes: 609594.125 num_examples: 2007 download_size: 2559509 dataset_size: 2804343.75 configs: - config_name: default data_files: - split: train path: data/train-* - split: test path: data/test-* license: unknown task_categories: - image-classification size_categories: - 1K<n<10K --- # Dataset Card for USPS USPS is a digit dataset automatically scanned from envelopes by the U.S. Postal Service containing a total of 9,298 16×16 pixel grayscale samples. ## Dataset Details The images are centered and normalized. They show a broad range of font styles. ### Dataset Sources - **Repository:** train set https://www.csie.ntu.edu.tw/~cjlin/libsvmtools/datasets/multiclass/usps.bz2, test set: https://www.csie.ntu.edu.tw/~cjlin/libsvmtools/datasets/multiclass/usps.t.bz2 - **Paper:** https://ieeexplore.ieee.org/abstract/document/291440 ## Uses In order to prepare the dataset for the FL settings, we recommend using [Flower Dataset](https://flower.ai/docs/datasets/) (flwr-datasets) for the dataset download and partitioning and [Flower](https://flower.ai/docs/framework/) (flwr) for conducting FL experiments. To partition the dataset, do the following. 1. Install the package. ```bash pip install flwr-datasets[vision] ``` 2. Use the HF Dataset under the hood in Flower Datasets. ```python from flwr_datasets import FederatedDataset from flwr_datasets.partitioner import IidPartitioner fds = FederatedDataset( dataset="flwrlabs/usps", partitioners={"train": IidPartitioner(num_partitions=10)} ) partition = fds.load_partition(partition_id=0) ``` ## Dataset Structure ### Data Instances The first instance of the train split is presented below: ``` { 'image': <PIL.PngImagePlugin.PngImageFile image mode=L size=16x16 at 0x133B4BA90>, 'label': 6 } ``` ### Data Split ``` DatasetDict({ train: Dataset({ features: ['image', 'label'], num_rows: 7291 }) test: Dataset({ features: ['image', 'label'], num_rows: 2007 }) }) ``` ## Citation When working with the USPS dataset, please cite the original paper. If you're using this dataset with Flower Datasets and Flower, cite Flower. **BibTeX:** Original paper: ``` @article{hull1994database, title={A database for handwritten text recognition research}, journal={IEEE Transactions on pattern analysis and machine intelligence}, volume={16}, number={5}, pages={550--554}, year={1994}, publisher={IEEE} } ```` Flower: ``` @article{DBLP:journals/corr/abs-2007-14390, author = {Daniel J. Beutel and Taner Topal and Akhil Mathur and Xinchi Qiu and Titouan Parcollet and Nicholas D. Lane}, title = {Flower: {A} Friendly Federated Learning Research Framework}, journal = {CoRR}, volume = {abs/2007.14390}, year = {2020}, url = {https://arxiv.org/abs/2007.14390}, eprinttype = {arXiv}, eprint = {2007.14390}, timestamp = {Mon, 03 Aug 2020 14:32:13 +0200}, biburl = {https://dblp.org/rec/journals/corr/abs-2007-14390.bib}, bibsource = {dblp computer science bibliography, https://dblp.org} } ``` ## Dataset Card Contact In case of any doubts about the dataset preprocessing and preparation, please contact [Flower Labs](https://flower.ai/).

dataset_info: 数据集信息: 特征: - 名称: image(图像) 数据类型: image(图像) - 名称: label(标签) 数据类型: 类别标签: 类别名称: '0': '0' '1': '1' '2': '2' '3': '3' '4': '4' '5': '5' '6': '6' '7': '7' '8': '8' '9': '9' 划分: - 名称: train(训练集) 字节数: 2194749.625 样本数: 7291 - 名称: test(测试集) 字节数: 609594.125 样本数: 2007 下载大小: 2559509 数据集总大小: 2804343.75 配置项: - 配置名称: default(默认配置) 数据文件: - 划分: train(训练集) 路径: data/train-* - 划分: test(测试集) 路径: data/test-* 许可证: 未知 任务类别: - 图像分类(image-classification) 样本规模类别: - 1K<n<10K(1千<样本数<1万) # USPS 数据集卡片 USPS 是由美国邮政总局(U.S. Postal Service)从信封自动扫描得到的手写数字数据集,总计包含9298张16×16像素的灰度图像样本。 ## 数据集详情 所有图像均经过居中对齐与归一化处理,涵盖了丰富多样的字体风格。 ### 数据集来源 - **代码仓库**: 训练集:https://www.csie.ntu.edu.tw/~cjlin/libsvmtools/datasets/multiclass/usps.bz2,测试集:https://www.csie.ntu.edu.tw/~cjlin/libsvmtools/datasets/multiclass/usps.t.bz2 - **学术论文**: https://ieeexplore.ieee.org/abstract/document/291440 ## 数据集用途 为将该数据集适配联邦学习(Federated Learning, FL)场景,我们推荐使用 [Flower 数据集](https://flower.ai/docs/datasets/)(flwr-datasets)完成数据集的下载与划分,并使用 [Flower 框架](https://flower.ai/docs/framework/)(flwr)开展联邦学习实验。 若需对数据集进行划分,请执行以下步骤: 1. 安装依赖包 bash pip install flwr-datasets[vision] 2. 在 Flower 数据集中底层使用 Hugging Face 数据集 python from flwr_datasets import FederatedDataset from flwr_datasets.partitioner import IidPartitioner fds = FederatedDataset( dataset="flwrlabs/usps", partitioners={"train": IidPartitioner(num_partitions=10)} ) partition = fds.load_partition(partition_id=0) ## 数据集结构 ### 数据实例 训练集划分的首个样本示例如下: { 'image': <PIL.PngImagePlugin.PngImageFile image mode=L size=16x16 at 0x133B4BA90>, 'label': 6 } ### 数据划分 DatasetDict({ train: Dataset({ features: ['image(图像)', 'label(标签)'], num_rows: 7291 }) test: Dataset({ features: ['image(图像)', 'label(标签)'], num_rows: 2007 }) }) ## 引用说明 在使用USPS数据集时,请引用其原始学术论文;若结合Flower数据集与Flower框架使用该数据集,请同时引用Flower相关文献。 **BibTeX 引用格式**: 原始论文: @article{hull1994database, title={A database for handwritten text recognition research}, journal={IEEE Transactions on pattern analysis and machine intelligence}, volume={16}, number={5}, pages={550--554}, year={1994}, publisher={IEEE} } Flower 框架: @article{DBLP:journals/corr/abs-2007-14390, author = {Daniel J. Beutel and Taner Topal and Akhil Mathur and Xinchi Qiu and Titouan Parcollet and Nicholas D. Lane}, title = {Flower: {A} Friendly Federated Learning Research Framework}, journal = {CoRR}, volume = {abs/2007.14390}, year = {2020}, url = {https://arxiv.org/abs/2007.14390}, eprinttype = {arXiv}, eprint = {2007.14390}, timestamp = {Mon, 03 Aug 2020 14:32:13 +0200}, biburl = {https://dblp.org/rec/journals/corr/abs-2007.14390.bib}, bibsource = {dblp computer science bibliography, https://dblp.org} } ## 数据集卡片联系人 若对数据集的预处理与制备流程存在任何疑问,请联系 [Flower 实验室](https://flower.ai/)。

提供机构:
fphool
二维码
社区交流群
二维码
科研交流群
商业服务