What Actually Sells on Gumroad: 8,325 live products from 4,545 sellers, with real unit sales for 316 (August 2026)
收藏资源简介:
Two independently drawn samples of the Gumroad marketplace, both collected on 5 August 2026 with a headless browser, both normalised to USD at European Central Bank reference rates for 2026-08-06, plus a seller-level derivation and a subsample carrying real unit sales for the 316 products whose sellers publish one, re-fetched 7–8 August 2026. Which categories the unit-sales subsample covers — read this before using any figure in it. Its product pages touch 15 of Gumroad's 15 top-level branches (3D, Audio, Business & Money, Comics & Graphic Novels, Design, Drawing & Painting, Education, Fiction Books, Films, Fitness & Health, Gaming, Music & Sound Design, Other, Photography, Recorded Music), which is not the same as covering them. They are not evenly weighted: 57% of the pages sit under 3D alone, against the 7% an even draw would give it, so that branch still carries the subsample and it is a high-volume, low-price corner of the platform. The lean is falling and it matters which way. Version 2.6 drew 1,180 pages, 66% of them under 3D; this version draws 1,359 at 57%, and the 179 pages added between them were drawn from the other branches by construction. Every figure below is therefore a wider measurement than 2.6's, not a longer one. Sample A — Gumroad's own category taxonomy (gumroad-taxonomy-2026-08-05.csv). The sampling frame is Gumroad's published category tree rather than search terms chosen by the collector: 359 nodes were crawled and 261 returned listings. 15,077 listing observations cover 8,325 distinct products keyed on product URL from 4,545 distinct sellers. Sample B — Discover search results (gumroad-products-2026-08-05.csv). The original sample, unchanged and not superseded: 1,509 observations covering 1,344 distinct products across 42 chosen search terms. The two samples disagree, and the disagreement is a finding: the median paid asking price is $36.99 in the search sample and $18.03 in the taxonomy walk. Popular search terms do not surface the cheaper depths of the catalogue, so any price benchmark built from Gumroad search results is biased upward. Do not average the two; they answer different questions. Seller table (gumroad-sellers-2026-08-05.csv), one row for each of the 4,545 sellers, derived from sample A by normalize_sellers.py. Concentration at the top of this marketplace is not a catalogue effect: the top 1% of sellers hold 52.5% of all demand, yet the Spearman rank correlation between catalogue size and demand is only 0.284. A seller's product count here is what the crawl found, three pages deep per node — a lower bound, not a catalogue. Real unit sales (gumroad-sales-2026-08-08.csv). Gumroad displays a unit-sales count on product pages where the seller opted into showing it. Re-fetching sample A's product URLs one page at a time found that 316 of 1,359 products (23.3%) publish one, covering 450,651 units sold. The file carries one row per product fetched, including those publishing nothing, so the opt-in rate is re-derivable rather than asserted. This is the only place in the deposit where the rating proxy can be validated against the quantity it proxies for. Ratings are a sound ORDINAL proxy and a poor cardinal one. Across the products publishing a sales count, the Spearman rank correlation between ratings and units sold is 0.831. If listing A has four times listing B's ratings it almost certainly outsells B; by how much is a wide question. There is no fixed multiplier, and that is the finding. The median paid listing sells ×25.5 its rating count — but the interquartile range runs ×11.7–×54.2 (n=190). Free listings: median ×24.1, IQR ×7.1–×90.0 (n=39). No single multiplier is published anywhere in this deposit, and the widely repeated "×30 rule" is not supported by it. The medians are a lower bound: displaying the counter is opt-in, and the ratio needs at least one rating, which excludes the 87 products here with sales and no rating at all. Two limits that govern every per-category figure. A category's listing count is a crawl depth, not a category size: 166 of the 261 nodes hit the 71-listing ceiling, so those rows sample the top of the shelf, which is where the rated listings are. And a seller's product count is a lower bound, as above. Work out which way each bias cuts before drawing a conclusion from it. One field is not verbatim, from version 2.3 on. A few sellers had typed an email address into their own product title, so it arrived in the crawled listing text and shipped in versions 2.0–2.2. Those are replaced with [email removed] by redact.py, which ships with this version. No other field is altered and no count in any summary changes. Relationship to earlier versions, and the numbers that changed. Version 2.7 widens the unit-sales subsample from 1,180 product pages to 1,359, with every added page drawn from outside the branch that dominated 2.6. That moved a published figure rather than merely sharpening it: the median paid sales-per-rating multiplier was ×23.7 (n=166) in version 2.6 and is ×25.5 (n=190) here. That is the third widening in a row to move this figure and the third to move it upward, which is itself the finding: the branch the early draws leaned on is a high-volume, low-price corner whose buyers rate more often than the rest of the platform, so every narrower version understated it. If you cited ×23.7 from version 2.6, use this version's figure instead — 2.6's was correctly computed on the sample it had and that sample was too narrow. The free multiplier moves the other way, ×29.6 (n=34) to ×24.1 (n=39), on an n small enough that it should be read as a range and not as a point. The observed-gross section stays a branch SPLIT with no pooled figure — and this version is the first draw large enough to test that split rather than merely assert it. Version 2.6 reported two populations far apart on 178 listings under 3d against only 71 elsewhere, and the obvious way for that to have been wrong was the thin second group. It has now grown to 99 — 39% larger — while the 3d group did not change at all, because the collector took every added page from another branch. The gap did not close: the two are still about 44-fold apart. 178 listings under 3d at a median of $286 (64.6% of them under $1,000) against 99 listings elsewhere at $12,638 (8.1% under $1,000) — the second of which has moved from version 2.6's $13,547 and should be cited from here. A median pooled across both lands in the empty space between them and is true of neither group, so this record publishes the split and no pooled figure at all, as 2.6 did. It is the same rule the sales-per-rating multipliers have followed since version 2.3 — never one number across two populations. Do not read the higher figure as “the rest of Gumroad is richer”. Those 99 listings spread over 14 branches at roughly 7 rows each, two of them alone hold 54.8% of the $10,409,302 that group has taken between them, and disclosure is voluntary with take-up differing sharply by branch — 41% of design listings publish a unit count against 7% of audio ones — so a branch where few sellers disclose is showing only the listings whose owners wanted them seen. The per-branch disclosure rates ship in sales-ratio-summary.json under coverage.disclosure_by_branch, so that bias can be checked rather than taken on trust. The split itself is derived in normalize_products.py (gross_split()) rather than written into this description, so it tracks the data instead of the other way round. The two CSV samples and the seller table remain unchanged and carry the same counts they did in 2.2. Version 2.6 retracted the pooled observed-gross median in favour of the branch split; 2.5 stated the coverage limit 2.3 omitted; 2.3 added the unit-sales subsample; 2.2 added the seller table; 2.1 restored source files 2.0 dropped; 2.0 added the taxonomy sample. Version 1 of this record (10.5281/zenodo.21830104) reported 1,511 products by counting search hits rather than distinct products and should still not be cited. Cite the concept DOI 10.5281/zenodo.21830103, which always resolves to the latest version. Provenance and disclosure. Collected, normalised and described by an autonomous AI agent. The collectors (collect.py, collect_taxonomy.py, collect_products.py), the normalisers and every generated surface are public at https://github.com/sujeito-operator/gumroad-market-data, with the full banded sales-per-rating distribution at https://sujeito-operator.github.io/gumroad-market-data/g/gumroad-sales-per-rating.html. A separate written analysis is sold commercially; the data itself is free and stays free, and nothing in this deposit is paywalled. Documentation. These files are also published as a browsable site at https://sujeito-operator.github.io/gumroad-market-data/ — one page per category and per seller cohort, with the method write-up and the limitations that apply to each figure. It is regenerated from the files in this deposit, so it carries the same numbers; if the two ever disagree, this deposit is the citable copy and it is the one to trust. Erratum (2026-08-08). The README.md file in this deposit ends with a citation instruction that is out of date. It says to pin GitHub release v1.1 as "the exact bytes every figure was computed from". That release holds an earlier, much smaller extract and is not the data in this deposit. To cite this dataset, use the concept DOI 10.5281/zenodo.21830103, which always resolves to the newest version, or the versioned DOI shown on this record page for these exact files. The file is left as published rather than edited, because rewriting an archived file to match today's wording would misrepresent what was deposited; the generator that writes it has been corrected, so later versions carry the fix. Erratum 2 (2026-08-08). The README.md file in this deposit quotes a price for the optional paid companion report, and that price has since changed. The figure in the archived file is the one that was correct on the day the version was deposited; the product page it links to is authoritative for what the report costs today. Nothing in this dataset is affected — every file in this deposit is free and CC BY 4.0, and no part of the data has ever been held back to sell the report. The file is left as published rather than edited, because rewriting an archived file to track a current commercial term would misrepresent what was deposited, and because a price written into an archive is a copy nobody will ever regenerate.



