Algae Senate — Data S2: Assembled Senate graphs and transformations
收藏资源简介:
Supporting Data S2 for "Searched-text length confounds substring screening in algal and cyanobacterial literature-graph extraction." This collection contains the assembled JSON graph outputs for all five extraction tracks: Senate-A, Senate-B, Senate-C, Senate-Cc, and Senate-Cq. Each track provides an entity catalog, relationship records, and entity provenance, together with grounded, high-confidence, normalized, and scored graph views; name canonicalization and abbreviation or embedding merge records; relation mappings; grounding reports; graph-quality summaries; and transformation reports. The Senate-Cc track is the fixed downstream substrate used for candidate enumeration, and its source-text-screened state contains 256,677 entities and 2,208,495 relationships, each linked to a source-document identifier. Contents are enumerated in the accompanying 101-file manifest and checksum record. These are model-derived graph records and deterministic transformations; inclusion in a graph does not establish biological validity.Companion Code S1: doi:10.5281/zenodo.22975324.



