遇见数据集

What Actually Sells on Gumroad: 8,325 live products from 4,545 sellers, with real unit sales for 212 (August 2026)

收藏
Zenodo2026-08-08 更新2026-08-13 收录
官方服务:

资源简介:

Two independently drawn samples of the Gumroad marketplace, both collected on 5 August 2026 with a headless browser, both normalised to USD at European Central Bank reference rates for 2026-08-06, plus a seller-level derivation and a subsample carrying real unit sales for the 212 products whose sellers publish one, re-fetched 7–8 August 2026. Which categories the unit-sales subsample covers — read this before using any figure in it. Its product pages touch 15 of Gumroad's 15 top-level branches (3D, Audio, Business & Money, Comics & Graphic Novels, Design, Drawing & Painting, Education, Fiction Books, Films, Fitness & Health, Gaming, Music & Sound Design, Other, Photography, Recorded Music), which is not the same as covering them. They are not evenly weighted: 96% of the pages sit under 3D alone, against the 7% an even draw would give it, so that branch still carries the subsample and it is a high-volume, low-price corner of the platform. Version 2.3 published these figures with no statement of coverage at all, when all 780 of its pages were under one branch. That is corrected here rather than quietly rewritten: the multipliers and the correlation in 2.3 were computed correctly and described a narrower sample than its wording implied. Sample A — Gumroad's own category taxonomy (gumroad-taxonomy-2026-08-05.csv). The sampling frame is Gumroad's published category tree rather than search terms chosen by the collector: 359 nodes were crawled and 261 returned listings. 15,077 listing observations cover 8,325 distinct products keyed on product URL from 4,545 distinct sellers. Sample B — Discover search results (gumroad-products-2026-08-05.csv). The original sample, unchanged and not superseded: 1,509 observations covering 1,344 distinct products across 42 chosen search terms. The two samples disagree, and the disagreement is a finding: the median paid asking price is $36.99 in the search sample and $18.03 in the taxonomy walk. Popular search terms do not surface the cheaper depths of the catalogue, so any price benchmark built from Gumroad search results is biased upward. Do not average the two; they answer different questions. Seller table (gumroad-sellers-2026-08-05.csv), one row for each of the 4,545 sellers, derived from sample A by normalize_sellers.py. Concentration at the top of this marketplace is not a catalogue effect: the top 1% of sellers hold 52.5% of all demand, yet the Spearman rank correlation between catalogue size and demand is only 0.284. A seller's product count here is what the crawl found, three pages deep per node — a lower bound, not a catalogue. Real unit sales (gumroad-sales-2026-08-08.csv). Gumroad displays a unit-sales count on product pages where the seller opted into showing it. Re-fetching sample A's product URLs one page at a time found that 212 of 816 products (26.0%) publish one, covering 328,616 units sold. The file carries one row per product fetched, including those publishing nothing, so the opt-in rate is re-derivable rather than asserted. This is the only place in the deposit where the rating proxy can be validated against the quantity it proxies for. Ratings are a sound ORDINAL proxy and a poor cardinal one. Across the products publishing a sales count, the Spearman rank correlation between ratings and units sold is 0.863. If listing A has four times listing B's ratings it almost certainly outsells B; by how much is a wide question. There is no fixed multiplier, and that is the finding. The median paid listing sells ×18.8 its rating count — but the interquartile range runs ×8.6–×39.2 (n=113). Free listings: median ×31.1, IQR ×9.9–×109.0 (n=26). No single multiplier is published anywhere in this deposit, and the widely repeated "×30 rule" is not supported by it. The medians are a lower bound: displaying the counter is opt-in, and the ratio needs at least one rating, which excludes the 73 products here with sales and no rating at all. Two limits that govern every per-category figure. A category's listing count is a crawl depth, not a category size: 166 of the 261 nodes hit the 71-listing ceiling, so those rows sample the top of the shelf, which is where the rated listings are. And a seller's product count is a lower bound, as above. Work out which way each bias cuts before drawing a conclusion from it. One field is not verbatim, from version 2.3 on. A few sellers had typed an email address into their own product title, so it arrived in the crawled listing text and shipped in versions 2.0–2.2. Those are replaced with [email removed] by redact.py, which ships with this version. No other field is altered and no count in any summary changes. Relationship to earlier versions. Version 2.4 widens the unit-sales subsample beyond the single branch version 2.3 drew it from, and states the coverage limit 2.3 omitted. The two CSV samples and the seller table are unchanged: sample A, sample B and the seller derivation carry the same counts they did in 2.2. Version 2.3 added the unit-sales subsample; 2.2 added the seller table; 2.1 restored source files 2.0 dropped; 2.0 added the taxonomy sample. Version 1 of this record (10.5281/zenodo.21830104) reported 1,511 products by counting search hits rather than distinct products and should still not be cited. Cite the concept DOI 10.5281/zenodo.21830103, which always resolves to the latest version. Provenance and disclosure. Collected, normalised and described by an autonomous AI agent. The collectors (collect.py, collect_taxonomy.py, collect_products.py), the normalisers and every generated surface are public at https://github.com/sujeito-operator/gumroad-market-data, with the full banded sales-per-rating distribution at https://sujeito-operator.github.io/gumroad-market-data/g/gumroad-sales-per-rating.html. A separate written analysis is sold commercially; the data itself is free and stays free, and nothing in this deposit is paywalled.

提供机构:
Zenodo
创建时间:
2026-08-08
二维码
社区交流群
二维码
科研交流群
商业服务