遇见数据集

Landing Page Hero Copy: 239 Real Pages, Deterministically Scored

收藏
Zenodo2026-07-25 更新2026-08-02 收录
官方服务:

资源简介:

What changed in v1.1 No score, flag or summary figure differs from v1.0. The 239 rows are identical. This version adds a method column naming the scoring rule that produced every row (static-fetch-regex-v1), plus a README section explaining why that matters: the maintained tool has since changed its rules, and re-scoring this corpus with the current rule would move 42 of the 239 rows. A dataset is scored by the rule of its date; re-scoring an archived row with a later rule produces a different dataset, not a correction. v1.0 remains citable at DOI 10.5281/zenodo.21543620. A snapshot of the hero section (headline, sub-headline, primary call-to-action) of 239 real landing pages — YC-backed companies, developer tools, SaaS products, AI products — collected on 2026-07-24, each scored 0–100 by a deterministic rule-based grader (no LLM, no model, no network call: the same input always yields the same score). Headline finding: 195 of the 239 pages (82%) contain no digit anywhere in the hero. Distribution: median 79, mean 80.1, 19 pages at 100/100, 31 pages below 70. Files landing_page_heroes.csv — 239 rows: domain, score, flags, headline, subheadline, cta landing_page_heroes.jsonl — the same records, one JSON object per line README.md — full method, rubric, exclusions and limits How the hero was extracted A single static HTTP fetch per domain, with JavaScript disabled, then regex extraction: first <h1> (fallback og:title, then <title>, repeated sentences collapsed); first following <h2>/<p> of 20–400 characters, cookie boilerplate excluded; first following <a>/<button> label of 2–40 characters, skip-links/consent/login excluded. A page yielding under 200 characters of text without JavaScript was rejected rather than guessed at. How the score is computed Five additive dimensions, 100 points: Anti-hype (25) = 25 − (7×hype words + 5×"!" + 4×emoji + 4×ALL-CAPS words); Specificity (25) = 25 if any digit is present, else 8 (+9 if a soft quantifier such as %, x, hours, days, minutes, no, zero appears); Clarity (25) = 25 − 6×filler words; Headline shape (13) = 13 for 3–10 words, 9 for 11–12, 5 above 12, 7 for 1–2, 0 if empty; CTA (12) = 12, or 4 if the label is in a fixed generic-CTA list, or 0 if absent. Word lists and reference implementation: github.com/parweb/landing-copy-grader (MIT). Flag frequencies in this corpus nonum 195 (82%) · filler 82 (34%) · weakcta 35 (15%) · caps 33 (14%) · hype 16 (7%) · shorthl 13 (5%) · longhl 9 (4%) · emoji 7 (3%) · excl 7 (3%). What this dataset is NOT The score is not an AI-detection verdict. There is no ground truth about authorship in this dataset and none is claimed. Nothing here can tell you whether a human or a model wrote a page. The score has no validation against human judgement — no annotated reference set, no inter-rater agreement, no measured precision or recall. It is a transparent heuristic, not a calibrated classifier. Good copy can score badly. stripe.com scores 61 here: its headline runs 173 characters and its CTA is "Get started", so it trips longhl, weakcta, filler and nonum at once. A low score means the rubric fired, not the page is bad. The word lists are opinionated and English-only. "Innovative", "platform" and "get started" are penalised by choice; the lists are published so that disagreement can be precise. The hero is not the page, and the snapshot describes 2026-07-24 only. Collection and exclusions (303 attempted → 239 retained) The 64 exclusions: 23 domains blocked by the network filter of the collecting machine (a corporate DNS filter blocking AI-related domains: cursor, huggingface, elevenlabs, mistral…) — excluded and not retried; 22 HTTP errors, mostly 403/503 bot protection (coinbase, openai, perplexity, reddit, medium…); 8 pages under 200 characters without JavaScript (spotify, twitch, duolingo, snowflake…); 7 pages served in French by geolocation (salesforce, paypal, klarna…), excluded because the word lists are English; 4 documented manual exclusions (airbnb.com screen-reader-only h1, dev.to feed, checkout.com fragmented headline, substack.com feed content). These exclusions are not random. Pages behind heavy bot protection or heavy client-side rendering are systematically absent, and because a network filter removed 23 AI-company domains, this is not a representative sample of AI products. It is a convenience sample of pages a plain HTTP client could read on one day from one machine. The authors' own page is in the corpus — meridian-demo-flax.vercel.app, 77/100, flagged filler and nonum — left in rather than removed, so the rubric can be checked against the people who wrote it. Rights in the quoted text The scores, flags, schema and grader code are the authors' (MIT); the deposit is released CC-BY-4.0. The headline/subheadline/cta strings are short verbatim excerpts of third-party public web pages: copyright in that text remains with the respective site owners, and each row names its source domain so any excerpt can be traced. Owners who want a row removed may contact the authors. Reproducing the figures Every aggregate above is recomputed from the rows in this deposit: 195/239 = 81.59%, which rounds to 82%. An earlier public write-up by the authors said "194 / 81%"; that was wrong and was corrected publicly. If a number elsewhere disagrees with these files, the files are the source of truth. An interactive view of all 239 rows: 1h-money-store.vercel.app/leaderboard.

提供机构:
Zenodo
创建时间:
2026-07-25
二维码
社区交流群
二维码
科研交流群
商业服务