CFDE-REVEAL: LLM prompts and implementation provenance for the hypothesis-generation workflow
收藏资源简介:
This archive documents the LLM prompts used in the CFDE-REVEAL workflow described in The Common Fund Data Ecosystem paper. REVEAL interprets biomedical research questions into structured search terms and generates mechanistic hypotheses from retrieved gene-set–trait association evidence. The archive includes the complete query-interpretation and hypothesis-generation prompts, an optional guided-query construction prompt, an exploratory-mode override, and an additional legacy extraction prompt retained for completeness. It also contains the original frontend source documenting dynamic user-message assembly, a description of prompt roles and diagnostic behavior, source provenance, and SHA-256 checksums. The materials were extracted from the dk-reveal-multi-direction-test branch of the Broad Institute’s dig-dug-portal repository at commit 9e583fc92ef2453c8f61c3a60d3a077a7ec8ca12: https://github.com/broadinstitute/dig-dug-portal/commit/9e583fc92ef2453c8f61c3a60d3a077a7ec8ca12 This deposit preserves the prompt implementation for transparency and citation. It is not a complete executable deployment or an evaluation dataset and does not establish the accuracy of generated hypotheses or the reliability of refusal and warning instructions. The exact deployed commit and example-run mode are not independently established by the source extraction.



