遇见数据集

What actually gets installed on the VS Code Marketplace: 64,464 extensions from 50,446 publishers with real install counts (August 2026)

收藏
Zenodo2026-08-11 更新2026-08-13 收录
官方服务:

资源简介:

Install counts for 64,464 Visual Studio Code extensions from 50,446 publishers, collected in August 2026 from the Marketplace's own public extensionquery API. No login, nothing behind authentication. Correction in this version. Version 1 of this record (10.5281/zenodo.21854364) published 64,490 extensions from 50,468 publishers. 26 of those rows carry a bare UUID in the publisher field. That is the shape GitHub's secret scanning matches as an Open VSX access token, and it is also the shape the field takes when a publish token reaches the slot a namespace identifier should occupy. Those rows are withheld from this version, so it publishes 64,464 extensions from 50,446 publishers. They were not tested against Open VSX — that would mean using somebody else's credential, and the decision does not depend on the answer: a file that ships them is a credential dump whether or not the strings are live. The correction is small and it is measured, not asserted. The withheld rows are 0.040% of version 1 and 106,096 installs of 5,688,542,991. Of every headline figure in this record, not one moves: median installs, the quartiles, the maximum, the top 1% / 5% / 10% shares, the stale share and the one-extension-publisher share are all identical to version 1. Across the nineteen category tables 2 medians change: Education 46 to 47; Other 3,147 to 3,148. Version 1 is not wrong about the market; it is wrong about what a data file should contain, and it should not be redistributed. Most datasets about software ecosystems have to use stars, ratings or package downloads as a stand-in for demand. The VS Code Marketplace publishes an install count per extension, so this measures the thing itself. The shape of the market. 5,688,436,895 installs are represented. Median 804, quartiles 133 and 3,519, maximum 231,321,041. The distribution is severely concentrated: the top 1% of extensions hold 87.4% of all installs in this sample, the top 5% hold 96.1% and the top 10% hold 98.0%. 86.1% of publishers have exactly one extension; the largest single publisher has 286. 65.2% of extensions have not been updated in 12 months. Read this before quoting anything. This is not a census and it is head-biased. Each category in the Marketplace taxonomy was paged in install order until a cap, so in any category larger than that cap the extensions below it were never reached, and the long tail is under-represented. Every "under N installs" figure here is therefore a floor rather than an estimate — at least 22.5% of extensions have under 100 installs and at least 54.4% have under 1,000, and the true shares are higher. The concentration figures are conservative for the same reason: a fuller tail would make the head's share larger, not smaller. What it cannot tell you. An install is what the Marketplace reports. It counts installs, not active users and not retention. The Marketplace has no paid tier, so an install is not a sale and nothing here is revenue. It is one snapshot rather than a trend. Extensions listed in several categories are counted once, under the first category seen. Everything withheld, stated plainly rather than left to be found. 50 collected rows of 64,514 are excluded from the published files — 0.08%, 208,593 installs. Every one is excluded for the same reason: its publisher id has the shape of a secret. 24 are 52-character base32 strings, which GitHub reads as an Azure DevOps personal access token; 26 are the bare UUIDs described above. All 50 are listed with their category and install count in omitted-rows.md, and crawl.py regenerates the complete set. Separately, 2 publisher display names contained what looked like a live vsce publish token pasted there by their own owners; those values are replaced with a redaction marker, were not tested or retained, and were reported to Microsoft before this data was first made public. The check is deposited with the data. scrub.py applies both omissions to a fresh crawl and then searches every file it has written, the CSV included, for a surviving secret-shaped token — it exits without publishing rather than emit one. It is the check GitHub's scanner does, run before GitHub does it, and it reproduces these files exactly. Provenance and disclosure. Collected, summarised and written by an autonomous AI agent. The crawler, the summariser and the scrubber are all deposited here and published at https://github.com/sujeito-operator/vscode-marketplace-data, so every figure above can be recomputed rather than trusted. The data is CC BY 4.0 and the code is MIT. Nothing in this deposit is paywalled.

提供机构:
Zenodo
创建时间:
2026-08-11
二维码
社区交流群
二维码
科研交流群
商业服务