From Pilots to Practice: A Governance-Technology Integration Framework for Circular Water in the Baltic Sea Region
收藏资源简介:
Dataset overview This dataset accompanies a study is based on a three-tier text corpus (Tier 1–3) totaling 372,856 tokens, spanning 2000–2025 (inclusive). Documents were selected for direct relevance to nutrient recycling, wastewater management, and circular-economy objectives in the Baltic Sea Region (BSR). Tokens denote total running words and types denote unique wordforms. The dataset is designed to enable cross-scale comparison of governance–technology integration mechanisms across policy framing, practitioner discourse, and implementation documentation. Tiered corpus structure Tier 1. Supranational policy corpus (enabling governance) 12 EU/HELCOM/EUSBSR policy documents totaling 96,841 tokens and 8,270 types (cleaned). These include EU directives, HELCOM strategies, and the EUSBSR Action Plan, forming a “vertical governance chain” from supranational frameworks to macro-regional coordination and national implementation contexts. Tier 2. Transnational forum discourse corpus (practice translation and learning) 7 Europe Forum Turku transcript files (2024–2025) totaling 135,151 tokens (2024: 38,467; 2025: 96,684) and comprising 526 coded statements of practitioner and policymaker discourse. The manuscript reports the thematic distribution of these statements: 44.5% contain water-related terms (often geographic references), 9.9% contain multiple water/circular-economy keywords, and 1.7% are explicitly coded “Sea Recycling”; dominant topics include EU funding mechanisms (36.1%), innovation policy (17.5%), and regional security (12.7%). Tier 3. Implementation corpus (project-level experimentation) 5 Interreg BSR projects (≈€20.5 million) totaling 140,864 tokens and ≈9,525 types, drawn from project proposals, deliverables, implementation reports, stakeholder-engagement records, and scaling strategies. The five projects are: ReNutriWater, WaterMan, NURSECOAST-II, City Blues, and CiNURGI. Preprocessing and derived “analysis-ready” datasets Across tiers, texts were cleaned and preprocessed consistently (lowercasing; stoplists; modals/negations retained; numbers ignored unless analyzing dates). Analytical results are reported using per-million-tokens (pmw) normalization and fixed collocation windows (main L5–R5, with sensitivity checks elsewhere). In addition to (or instead of) redistributing full source texts, the dataset is naturally structured into derived outputs that reproduce the paper’s reported evidence, including: AntConc exports (wordlists, keywords, collocates with LL and MI/“Effect”, n-grams, KWIC samples). Cross-tier mechanism concordance outputs (mechanism seed list and a mechanism × tier coverage matrix computed as files hit / total files). InfraNodus “features-only” outputs per tier/year slice (community summaries; modularity; top betweenness/degree tables; network metrics). Corpus statistics tables (tokens/types per file; tier labels), plus preprocessing assets (core stoplist; transcript add-on stoplist; reporting conventions).



