POISK: Patent-extracted Organization of Immunoglobulin Sequence Knowledge
收藏资源简介:
Overview This dataset comprises a curated collection of antibody sequences and associated functional metadata extractedfrom the global patent literature. By converting unstructured legal disclosures into a structured bioinformaticformat, this resource enables the large-scale analysis of therapeutic antibody space, epitope landscape mapping,and the benchmarking of protein-protein interaction (PPI) prediction models. The dataset includes variable region sequences (VH and VL), target protein specifications, experimentalbinding affinities, and epitope residue definitions where available. Dataset Structure and Documentation The data is provided in a flattened CSV format, where each row represents an antibody-target pair or aspecific binding record. For a full breakdown of the data schema, please refer to the accompanying file:dataset_column_description.csv. Key Data Categories: Identifiers: Patent IDs, internal antibody nomenclature, and unique record identifiers. Sequences: Amino acid sequences for Heavy (VH) and Light (VL) variable regions, as well as the antigensequence, including associated SEQ ID NOs from the original patent filings. Target Metadata: Target protein names, gene symbols, UniProt IDs, and source organisms. Biophysical Properties: Experimental binding affinity, assay methods (e.g., SPR, BLI), and specific bindingconditions. Epitope Information: Descriptions of binding sites, including linear ranges, structural contact residues,and competition groups. Methodology: LLM-Based Extraction The extraction of data from high-volume, unstructured patent text was performed using an automated pipelinepowered by Large Language Models (LLMs). This approach allows for the capture of nuanced data points,such as assay conditions and specific epitope residues that are typically inaccessible to standard rule-basedparsers. Quality Control and Algorithmic Audit Users should be aware that LLM-based extraction is subject to generative inconsistencies, commonly referredto as hallucinations. Discrepancies may include the misassociation of affinity values with specific antibodyvariants or errors in residue indexing. To ensure data integrity, we have implemented a multi-stage audit layer. Every record is accompanied byprovenance and confidence metadata. Users are strongly advised to consult the following columns if a datapoint appears inconsistent: audit.verdict: A categorical reliability score (e.g., CONTRADICTED, PLAUSIBLE, CONFIRMED). audit.flags: Field names with potential mismatches. audit.reasoning: A natural language explanation detailing the logic used by the model to reach thefinal extracted value. Usage and Limitations This dataset is intended for research and benchmarking purposes. Given the legal and technical complexityof patent filings, users are encouraged to verify critical sequences or affinity data against the original patentdocuments (accessible via the patent_id column) before initiating downstream experimental validation.For further details regarding the extraction pipeline, performance benchmarks, and validation metrics,please refer to the main manuscript: [Link to Manuscript/DOI] (will be added later). We also provide metrics from the six folding models used for the benchmarking including ipTM and ipSAE (poisk_all_models_folding_metrics.zip).



