Verdict Cross-Publication Product Review Ratings
收藏资源简介:
6,010 product ratings collected from 2,016 distinct review publications, covering 1,809 consumer products across 351 categories — with each score traced to the review that published it, and flagged for whether the publisher printed the number or it was inferred from their prose. Most product-rating datasets are retailer data: star ratings left by customers on one storefront. This one is editorial data — what professional and independent reviewers concluded, across many outlets, for the same product. That makes it possible to ask questions a single-source dataset cannot: how often independent reviewers actually agree, whether price predicts published quality, which outlets grade systematically high or low relative to their peers, and how much of the review web still publishes a machine-readable score at all. Files. ratings.csv (6,010 rows, one per cited review: publisher, review URL, the score in the publisher's own scale, and how it was obtained). products.csv (1,809 rows: category, rank, aggregate rating, and how many reviews it rests on). README.md (full schema, method, and limitations). The column that matters most. 44.5% of these scores were read off the publisher's own page; 44.6% were inferred by a language model from the review's prose or transcript; 10.9% are reviews carrying no score at all. Filter publisher_stated == true for the 2,674 rows a human editor demonstrably published. Please do not over-read that split. A false means we did not read a number, which conflates two different things: the reviewer published none, or they published one and our extraction missed it (client-side-injected scores are the classic case). A second column, outlet_publishes_scores, flags the 373 inferred rows sitting on hosts we have successfully read five or more printed scores from — on those, our read failure is the likelier explanation. The honest floor is therefore that at least 50.7% of these reviews carried a published score, and the true figure is higher. Treat those rows as unknown rather than as absence. Nothing here was tested by the authors. This is an aggregation of other people's published reviews; the contribution is the cross-outlet comparison, not original measurement. Reviewers' verbatim quotes are deliberately excluded — they are other publications' copyrighted prose, and an internal audit found a meaningful share of stored quotes could not be located on the page they credited. The README lists eight limitations in full, including that ratings are scoped to the guide a product appears in (so one product can carry two aggregates), that publisher labels mix hostname slugs with YouTube channel names and need normalising, that syndication is not detected, and that a four-product category generally reflects a publication floor rather than the size of the market. Canonical page and updates: verdict-reviews.com/data



