遇见数据集

landing-page-hero-copy-239

收藏
OpenML2026-07-25 更新2026-08-16 收录
官方服务:

资源简介:

Hero sections (headline, sub-headline, primary call-to-action) of 239 real landing pages - YC-backed companies, developer tools, SaaS products, AI products - collected on 2026-07-24 by a single static HTTP fetch per domain with JavaScript disabled, then regex extraction. Each row carries a 0-100 score from a deterministic rule-based grader: no LLM, no model, no network call, so the same input always yields the same score. Headline finding: 195 of the 239 pages (82%) contain no digit anywhere in the hero. Median 79, mean 80.1, 19 pages at 100/100, 31 below 70. Rubric (five additive dimensions, 100 points). Anti-hype 25 = 25 - (7 x hype words + 5 x "!" + 4 x emoji + 4 x ALL-CAPS words). Specificity 25 = 25 if any digit is present, else 8, +9 if a soft quantifier appears (%, x, hours, days, minutes, no, zero). Clarity 25 = 25 - 6 x filler words. Headline shape 13 = 13 for 3-10 words, 9 for 11-12, 5 above 12, 7 for 1-2, 0 if empty. CTA 12 = 12, or 4 for a generic label from a fixed list, or 0 if absent. Word lists and reference implementation: https://github.com/parweb/landing-copy-grader (MIT). The "flags" column is a pipe-separated list of the tells that fired: nonum (no digit anywhere) 195 rows, filler 82, weakcta 35, caps 33, hype 16, shorthl 13, longhl 9, emoji 7, excl 7. WHAT THIS IS NOT. The score is not an AI-detection verdict: this dataset contains no ground truth about authorship and none is claimed. Nothing here can tell you whether a human or a model wrote a page. The score has never been validated against human judgement - no annotated reference set, no inter-rater agreement, no measured precision or recall. It is a transparent heuristic, not a calibrated classifier. Good copy can score badly: stripe.com scores 61 here because its headline runs 173 characters and its CTA is "Get started", tripping longhl, weakcta, filler and nonum at once. The word lists are opinionated and English-only. The hero is not the page - anything below the fold or rendered client-side is invisible to this method. It is a snapshot of 2026-07-24 and nothing else. COLLECTION AND EXCLUSIONS: 303 URLs attempted, 239 retained. The 64 exclusions: 23 domains blocked by the network filter of the collecting machine (a corporate DNS filter blocking AI-related domains - cursor, huggingface, elevenlabs, mistral), excluded and not retried; 22 HTTP errors, mostly 403/503 bot protection (coinbase, openai, perplexity, reddit, medium); 8 pages returning under 200 characters without JavaScript (spotify, twitch, duolingo, snowflake); 7 pages served in French by geolocation (salesforce, paypal, klarna), excluded because the word lists are English; 4 documented manual exclusions (airbnb.com screen-reader-only h1, dev.to feed, checkout.com fragmented headline, substack.com feed content). These exclusions are not random: pages behind heavy bot protection or heavy client-side rendering are systematically absent, and because a network filter removed 23 AI-company domains this is not a representative sample of AI products. It is a convenience sample of what a plain HTTP client could read on one day from one machine. The authors' own page is in the corpus - meridian-demo-flax.vercel.app, 77/100, flagged filler and nonum - left in rather than removed. RIGHTS: scores, flags, schema and grader code are the authors' (MIT); the compilation is CC-BY-4.0. The headline / subheadline / cta strings are short verbatim excerpts of third-party public web pages; copyright in that text remains with the respective site owners, and every row names its source domain so any excerpt can be traced. Owners who want a row removed may contact the authors. Archival copy of record, with the full README: https://doi.org/10.5281/zenodo.21543620

创建时间:
2026-07-25
二维码
社区交流群
二维码
科研交流群
商业服务