StephanAkkerman/fintwit-images
收藏资源简介:
--- language: - en license: mit task_categories: - image-classification - image-feature-extraction pretty_name: FinTwit Images dataset_info: features: - name: image dtype: image - name: label dtype: class_label: names: '0': charts '1': non-charts - name: id dtype: string splits: - name: train num_bytes: 320924227.116 num_examples: 4177 download_size: 444164916 dataset_size: 320924227.116 configs: - config_name: default data_files: - split: train path: data/train-* tags: - fintwit - twitter - charts - financial - financial charts - finance - stocks - crypto - image --- ## FinTwit Images This dataset is a collection of a sample of images from tweets that I scraped using my [Discord bot](https://github.com/StephanAkkerman/fintwit-bot) that keeps track of financial influencers on Twitter. The data consists of images that were part of tweets that did not mention a ticker. This dataset can be used for a wide variety of tasks, such as image classification or feature extraction. ### FinTwit Charts Collection This dataset is part of a larger collection of datasets, scraped from Twitter and labeled by a human (me). Below is the list of related datasets. - [Crypto Charts](huggingface.co/datasets/StephanAkkerman/crypto-charts): Images of financial charts of cryptocurrencies - [Stock Charts](https://huggingface.co/datasets/StephanAkkerman/stock-charts): Images of financial charts of stocks - [FinTwit Images](https://huggingface.co/datasets/StephanAkkerman/fintwit-images): Images that had no clear description, this contains a lot of non-chart images ## Dataset Structure Each images in the dataset is structured as follows: - **Image**: The image of the tweet, this can be of varying dimensions. - **Label**: A numerical label indicating the category of the image, with '1' for charts, and '0' for non-charts. ## Dataset Size The dataset comprises 4,579 images in total, categorized into: - 1,083 chart images - 3,496 non-chart images ## Usage I used this dataset for training my [chart-recognizer model](https://huggingface.co/StephanAkkerman/chart-recognizer) for classifying if an image is a chart or not. ## Acknowledgments We extend our heartfelt gratitude to all the authors of the original tweets. ## License This dataset is made available under the MIT license, adhering to the licensing terms of the original datasets.
语言: - 英语 许可证:MIT许可证 任务类别: - 图像分类 - 图像特征提取 友好名称:FinTwit Images 数据集信息: 特征: - 名称:图像 数据类型:图像 - 名称:标签 数据类型: 类别标签: 类别名称: '0': 非图表图像 '1': 图表图像 - 名称:ID 数据类型:字符串 划分集: - 名称:训练集 字节数:320924227.116 样本数量:4177 下载大小:444164916 数据集总大小:320924227.116 配置: - 配置名称:默认配置 数据文件: - 划分集:训练集 路径:data/train-* 标签集: - fintwit - Twitter - 图表 - 金融 - 金融图表 - 财经 - 股票 - 加密货币 - 图像 --- ## FinTwit 图像数据集 本数据集为笔者使用自研Discord机器人(用于追踪Twitter平台上的财经意见领袖)抓取的推文配图样本集合,该机器人的开源仓库地址为:https://github.com/StephanAkkerman/fintwit-bot。 本数据集包含的配图均来自未提及股票代码的推文。本数据集可适用于多种任务场景,例如图像分类与特征提取。 ### FinTwit 图表数据集集合 本数据集为更大规模数据集集合的一部分,该集合中的数据均从Twitter抓取并由笔者人工标注。以下为相关数据集列表: - [加密货币图表](https://huggingface.co/datasets/StephanAkkerman/crypto-charts): 加密货币金融图表配图 - [股票图表](https://huggingface.co/datasets/StephanAkkerman/stock-charts): 股票金融图表配图 - [FinTwit Images](https://huggingface.co/datasets/StephanAkkerman/fintwit-images): 无明确主题的配图集合,其中包含大量非图表类图像 ### 数据集结构 本数据集内的每张图像均包含以下信息: - **图像**:推文配图,尺寸不固定。 - **标签**:用于标识图像类别的数值标签,其中`1`代表图表图像,`0`代表非图表图像。 ### 数据集规模 本数据集总计包含4579张图像,分类如下: - 1083张图表图像 - 3496张非图表图像 ### 使用场景 笔者曾使用本数据集训练自研的[图表识别模型](https://huggingface.co/StephanAkkerman/chart-recognizer),用于实现图像是否为金融图表的分类任务。 ### 致谢 谨向所有原始推文的作者致以诚挚谢意。 ### 许可证 本数据集采用MIT许可证进行分发,同时遵循原始数据集的相关许可条款。
FinTwit Images 数据集概述
基本信息
- 语言: 英语
- 许可证: MIT
- 任务类别: 图像分类, 图像特征提取
- 美观名称: FinTwit Images
数据集特征
- 特征:
- image: 图像数据
- label: 标签数据,类别包括 charts 和 non-charts
- id: 字符串类型
数据集分割
- train:
- 字节数: 320924227.116
- 样本数: 4177
数据集大小
- 下载大小: 444164916
- 数据集大小: 320924227.116
配置
- default:
- 数据文件:
- split: train
- path: data/train-*
- 数据文件:
标签
- tags:
- fintwit
- charts
- financial
- financial charts
- finance
- stocks
- crypto
- image
数据集结构
- Image: 推文图像,尺寸不一
- Label: 数值标签,1 表示 charts,0 表示 non-charts
数据集大小
- 总图像数: 4579
- chart 图像数: 1083
- non-chart 图像数: 3496
使用
- 用于训练 chart-recognizer 模型,用于分类图像是否为 chart。
许可证
- 该数据集在 MIT 许可证下发布。




