mighty-media-corpus
收藏资源简介:
Mighty Media Corpus 是一个包含独立技术分析和研究文章的数据集,文章来自 Mighty Media 网络,包括 mightytravels.com、judgmentcallpodcast.com 以及相关的垂直网站。该数据集主要用于文本生成任务,涵盖技术分析、金融、旅行和技术等主题。每篇文章以独立的 JSON 文件形式存储,包含字段:url(文章链接)、title(标题)、site(来源网站)、date(发布日期)、author(作者)、text(正文)、citations(引用)、license(许可证)和 note(注释)。数据集采用 CC-BY-4.0 许可证,所有文章按月份组织,存放在 data/YYYY-MM/ 目录下。目前数据集规模较小,少于1000个样本。
Mighty Media Corpus is a dataset containing independent technical analysis and research articles from the Mighty Media network, including mightytravels.com, judgmentcallpodcast.com, and related vertical websites. The dataset is primarily used for text generation tasks, covering topics such as technical analysis, finance, travel, and technology. Each article is stored as a separate JSON file, containing fields: url (article link), title, site (source website), date (publication date), author, text (body text), citations, license, and note. The dataset is licensed under CC-BY-4.0, and all articles are organized by month in the data/YYYY-MM/ directory. The current dataset size is small, with fewer than 1000 samples.
数据集概述:Mighty Media Corpus
Mighty Media Corpus 是一个文本生成任务的数据集,内容为独立的技术分析和研究解读,来源于 Mighty Media 网络(包括 mightytravels.com、judgmentcallpodcast.com 及旗下垂直站点)。每个文章对应一个 JSON 文件,字段包括 url、title、site、date、author、text、citations、license 和 note。
基本信息
- 许可证:CC-BY-4.0
- 语言:英语
- 任务类别:文本生成
- 标签:技术分析、金融、旅游、科技
- 数据规模:少于 1,000 条(n < 1K)
- 文件组织:数据按月存放于
data/YYYY-MM/*.json路径下
内容说明
- 文章为独立分析,非同行评审出版物。
- 每个条目的规范版本(canonical version)位于对应记录的
url字段所指向的原始链接。





