Dataset A Survey Charting Trust Models and Systems in Self-Sovereign Identity
收藏资源简介:
PRISMA research dataset This dataset accompanies our literature review of Self-Sovereign Identity Trust Models and Trust Management Systems research. This dataset contains: The list of papers that can be retrieved on the OpenAlex and CrossRef platforms for the period of published works considered. The dataset of papers processed in PRISMA phases up to the final dataset The files are organized as a literature-selection flow: identified records, deduplicated records, title/abstract screening outcomes, final corpus, final-corpus recovery report, flow counts, and per-paper final-assessment exclusion reasons. Files 1. `01_identified_records_curated.csv` - All records identified by the search protocol before deduplication. 2. `02_deduplicated_records_curated.csv` - Unique records retained after DOI and title-based deduplication. 3. `03_title_abstract_screened_in_curated.csv` - Records retained after title/abstract screening. 4. `04_title_abstract_excluded_curated.csv` - Records excluded after title/abstract screening, with exclusion reasons. 5. `05_final_corpus_curated.csv` - Final curated survey corpus with enriched bibliographic metadata. 6. `06_final_corpus_recovery_report_curated.csv` - Matching report comparing the final corpus against the search dataset. 7. `07_snowballing_citation_evidence.csv` - Citation-link audit artifact. In this curated flow, citation searching contributes 0 records to the PRISMA other-methods arm. 8. `08_literature_selection_flow_counts.csv` - Summary counts for the PRISMA 2020 literature-selection flow, reported as two identification arms (databases/registers and other methods) that merge at the final included count (`metric`, `count`). 9. `09_final_assessment_exclusion_reasons.csv` - Per-paper status for records that were assessed for eligibility but not included in the final corpus, plus any report that was sought but could not be retrieved (the `exclusion_stage` column distinguishes `final_assessment` from `reports_not_retrieved`). `curated_final_corpus.csv` is kept as a convenience copy of `05_final_corpus_curated.csv`. The PRISMA 2020 flow diagram is provided as `prisma_flow_diagram.pdf` (rendered figure), `prisma_flow_diagram.md` (Mermaid source), and `prisma_flow_diagram.tex` / `prisma_flow_diagram_figure_only.tex` (LaTeX source using `prisma-flow-diagram.sty`). Selection Flow The corpus was selected manually (publisher-curated libraries -- ScienceDirect/ Elsevier, IEEE Xplore, ACM Digital Library, SpringerLink -- plus Google Scholar and snowball sampling, followed by manual full-text assessment). The counts below describe the automated reproducibility pipeline that reconstructs this selection over public metadata APIs (OpenAlex, Crossref, DataCite), reported as two PRISMA 2020 identification arms: - Databases and registers: the database/registers search identified 433 records; after deduplication 73 records were screened at title/abstract level (0 excluded). Of the 73 reports sought for retrieval, 1 could not be retrieved because its full text was unavailable, so 72 were assessed for eligibility, 51 were excluded with reasons (see `09_final_assessment_exclusion_reasons.csv`), and 21 were included. The query `SSI trust VC` recovers both Bistarelli versions in Crossref: the ACM SAC poster (10.1145/3672608.3707964) and the PerCom Workshops paper ref73 (10.1109/percomworkshops65533.2025.00036). The PerCom Workshops paper is the included final-corpus record; the ACM SAC poster is retained in `01_identified_records_curated.csv` as a DOI-confirmed duplicate/ alternate database hit and removed before deduplication/screening. - Other methods: 2 further studies were identified as hand-selected primary sources -- the two Sovrin governance/technical documents -- and assessed for eligibility, with none excluded. Citation searching is retained only as an audit artifact and contributes 0 records to the PRISMA other-methods arm. Records identified through other methods do not pass through title/abstract screening, so the two arms are kept separate and merge only at the final count: 21 + 2 = 23 included studies. Dataset The final corpus includes academic articles, conference papers, book chapters, and primary ecosystem documents. The first four CSV files use a shared schema: - `dataset_record_id`: stable identifier for the row in this dataset. - `title`: work title. - `authors`: semicolon-separated author list. - `year`: publication year where available. - `publication_date`: publication date where available. - `venue_or_container`: journal, conference, proceedings, book, or document container. - `publisher`: publisher or issuing organization. - `doi`: DOI where available. - `url`: DOI URL or public record URL. - `abstract`: abstract text where available. - `search_phrase`: search phrase that identified the record. - `matched_identity_terms`: identity-related terms found in title or abstract. - `matched_trust_terms`: trust-related terms found in title or abstract. - `duplicate_status`: `unique` or `duplicate`. - `duplicate_of`: canonical dataset record if duplicate. - `screening_decision`: selection decision for the stage. - `exclusion_reason`: exclusion reason, where applicable. - `curation_notes`: additional notes. Final Corpus Columns The final corpus file uses an enriched bibliography schema: - `corpus_id`: stable final-corpus identifier. - `source_ref`: reference number in the survey manuscript (pre-publication draft numbering; the published bibliography numbers are offset by +1, e.g. draft `[73]` corresponds to published `[74]`). - `title`: work title. - `authors`: semicolon-separated author list. - `year`: publication year. - `document_type`: publication or document category. - `venue_or_container`: journal, conference, proceedings, book, or document container. - `editors`: semicolon-separated editor list, where applicable. - `publisher`: publisher or issuing organization. - `doi`: DOI where available. - `url`: DOI URL or public document URL. - `pages`: page range where available. - `notes`: corpus curation notes. Final-Assessment Exclusion Reason Columns The final-assessment exclusion file uses this schema: - `record_id`: stable record identifier from the staged search dataset. - `title`: work title. - `doi`: DOI where available. - `year`: publication year where available. - `venue`: journal, conference, proceedings, book, or document container. - `exclusion_stage`: assessment stage at which the record was excluded. - `exclusion_reason`: machine-readable exclusion reason. - `exclusion_reason_label`: human-readable exclusion reason. - `superseded_by_corpus_id`: final-corpus identifier for the included full version, where applicable. - `superseded_by_title`: title of the included full version, where applicable. - `notes`: curation notes supporting the exclusion decision.



