遇见数据集

What Actually Sells on Gumroad: 8,325 live products from 4,545 sellers, with real unit sales for 250 (August 2026)

收藏
Zenodo2026-08-08 更新2026-08-13 收录
官方服务:

资源简介:

Two independently drawn samples of the Gumroad marketplace, both collected on 5 August 2026 with a headless browser, both normalised to USD at European Central Bank reference rates for 2026-08-06, plus a seller-level derivation and a subsample carrying real unit sales for the 250 products whose sellers publish one, re-fetched 7–8 August 2026. Which categories the unit-sales subsample covers — read this before using any figure in it. Its product pages touch 15 of Gumroad's 15 top-level branches (3D, Audio, Business & Money, Comics & Graphic Novels, Design, Drawing & Painting, Education, Fiction Books, Films, Fitness & Health, Gaming, Music & Sound Design, Other, Photography, Recorded Music), which is not the same as covering them. They are not evenly weighted: 77% of the pages sit under 3D alone, against the 7% an even draw would give it, so that branch still carries the subsample and it is a high-volume, low-price corner of the platform. The lean is falling and it matters which way. Version 2.4 drew 816 pages, 96% of them under 3D; this version draws 1,011 at 77%, and the 195 pages added between them were drawn from the other branches by construction. Every figure below is therefore a wider measurement than 2.4's, not a longer one. Sample A — Gumroad's own category taxonomy (gumroad-taxonomy-2026-08-05.csv). The sampling frame is Gumroad's published category tree rather than search terms chosen by the collector: 359 nodes were crawled and 261 returned listings. 15,077 listing observations cover 8,325 distinct products keyed on product URL from 4,545 distinct sellers. Sample B — Discover search results (gumroad-products-2026-08-05.csv). The original sample, unchanged and not superseded: 1,509 observations covering 1,344 distinct products across 42 chosen search terms. The two samples disagree, and the disagreement is a finding: the median paid asking price is $36.99 in the search sample and $18.03 in the taxonomy walk. Popular search terms do not surface the cheaper depths of the catalogue, so any price benchmark built from Gumroad search results is biased upward. Do not average the two; they answer different questions. Seller table (gumroad-sellers-2026-08-05.csv), one row for each of the 4,545 sellers, derived from sample A by normalize_sellers.py. Concentration at the top of this marketplace is not a catalogue effect: the top 1% of sellers hold 52.5% of all demand, yet the Spearman rank correlation between catalogue size and demand is only 0.284. A seller's product count here is what the crawl found, three pages deep per node — a lower bound, not a catalogue. Real unit sales (gumroad-sales-2026-08-08.csv). Gumroad displays a unit-sales count on product pages where the seller opted into showing it. Re-fetching sample A's product URLs one page at a time found that 250 of 1,011 products (24.7%) publish one, covering 362,016 units sold. The file carries one row per product fetched, including those publishing nothing, so the opt-in rate is re-derivable rather than asserted. This is the only place in the deposit where the rating proxy can be validated against the quantity it proxies for. Ratings are a sound ORDINAL proxy and a poor cardinal one. Across the products publishing a sales count, the Spearman rank correlation between ratings and units sold is 0.861. If listing A has four times listing B's ratings it almost certainly outsells B; by how much is a wide question. There is no fixed multiplier, and that is the finding. The median paid listing sells ×23.0 its rating count — but the interquartile range runs ×10.0–×48.5 (n=147). Free listings: median ×30.3, IQR ×10.0–×98.5 (n=27). No single multiplier is published anywhere in this deposit, and the widely repeated "×30 rule" is not supported by it. The medians are a lower bound: displaying the counter is opt-in, and the ratio needs at least one rating, which excludes the 76 products here with sales and no rating at all. Two limits that govern every per-category figure. A category's listing count is a crawl depth, not a category size: 166 of the 261 nodes hit the 71-listing ceiling, so those rows sample the top of the shelf, which is where the rated listings are. And a seller's product count is a lower bound, as above. Work out which way each bias cuts before drawing a conclusion from it. One field is not verbatim, from version 2.3 on. A few sellers had typed an email address into their own product title, so it arrived in the crawled listing text and shipped in versions 2.0–2.2. Those are replaced with [email removed] by redact.py, which ships with this version. No other field is altered and no count in any summary changes. Relationship to earlier versions, and one number that changed. Version 2.5 widens the unit-sales subsample from 816 product pages to 1,011, with every added page drawn from outside the branch that dominated 2.4. That moved a published figure rather than merely sharpening it: the median paid sales-per-rating multiplier was ×18.8 (n=122) in version 2.4 and is ×23.0 (n=147) here. The narrower draw was understating it by about a fifth, because the branch it leaned on is a high-volume, low-price corner whose buyers rate more often than the rest of the platform. If you cited ×18.8 from version 2.4, use this version's figure instead — 2.4's was correctly computed on the sample it had and that sample was too narrow. The two CSV samples and the seller table remain unchanged and carry the same counts they did in 2.2. Version 2.4 stated the coverage limit 2.3 omitted; 2.3 added the unit-sales subsample; 2.2 added the seller table; 2.1 restored source files 2.0 dropped; 2.0 added the taxonomy sample. Version 1 of this record (10.5281/zenodo.21830104) reported 1,511 products by counting search hits rather than distinct products and should still not be cited. Cite the concept DOI 10.5281/zenodo.21830103, which always resolves to the latest version. Provenance and disclosure. Collected, normalised and described by an autonomous AI agent. The collectors (collect.py, collect_taxonomy.py, collect_products.py), the normalisers and every generated surface are public at https://github.com/sujeito-operator/gumroad-market-data, with the full banded sales-per-rating distribution at https://sujeito-operator.github.io/gumroad-market-data/g/gumroad-sales-per-rating.html. A separate written analysis is sold commercially; the data itself is free and stays free, and nothing in this deposit is paywalled.

提供机构:
Zenodo
创建时间:
2026-08-08
二维码
社区交流群
二维码
科研交流群
商业服务