FlyRank/internship-starter
收藏资源简介:
--- license: other language: - en tags: - seo - search-console - content-performance - tabular - education - flyrank-internship pretty_name: FlyRank Internship — Starter (Content Refresh, Anonymized) size_categories: - 10K<n<100K --- # FlyRank Internship — Starter Dataset (Anonymized) The public, safe starting point for the FlyRank **Applied Search Intelligence** ML internship. **30,000** anonymized content-performance rows across **32** pseudonymized clients (53 columns). **Public-safe:** hashed `content_id` / `client_id` + numeric/categorical metrics only — **no** titles, URLs, keywords, domains, or client names. ## What it's for Week 1–2 quick wins and the ready-now capstone lanes (ranking-signal analysis, lifecycle / opportunity scoring, content-archetype clustering). ## Verified reference results (this 30k slice) - Rule baseline **Precision@50 = 0.26** → Random Forest **Precision@50 = 0.74** - `search_volume` vs `impressions_90d` correlation ≈ **0.0012** (essentially zero — a real myth-buster) - Weighted CTR by position: `top_3` **0.49%** → `page_1` 0.35% → `deep` **0.04%** - Length is *not* the differentiator: growing vs declining word count ≈ 2,850 vs 2,910 ## Safety rules Anonymized, but still treat row-level outputs as not-for-careless-publishing. Do **not** use product flags (`health_score`, `needs_ctr_fix`, `is_quick_win`, …) as model features — they leak the decline label. Keep all public outputs anonymized/aggregate.
FlyRank Internship — Starter Dataset (Anonymized) is the public, safe starting point for the FlyRank Applied Search Intelligence ML internship. It contains 30,000 anonymized content-performance rows across 32 pseudonymized clients (53 columns). The data is public-safe with hashed content_id/client_id and numeric/categorical metrics only — no titles, URLs, keywords, domains, or client names. It is intended for Week 1–2 quick wins and the ready-now capstone lanes (ranking-signal analysis, lifecycle/opportunity scoring, content-archetype clustering). Verified reference results include: rule baseline Precision@50 = 0.26 → Random Forest Precision@50 = 0.74; search_volume vs impressions_90d correlation ≈ 0.0012 (essentially zero); weighted CTR by position: top_3 0.49% → page_1 0.35% → deep 0.04%; length is not the differentiator: growing vs declining word count ≈ 2,850 vs 2,910. Safety rules: anonymized but treat row-level outputs carefully; do not use product flags (health_score, needs_ctr_fix, is_quick_win, …) as model features to avoid label leakage; keep all public outputs anonymized/aggregate.





