Test size and power using detection in subsets.
收藏资源简介:
1Number cases and controls reduced to 100, so sequencing exhausts cases.For each line, except the last, 500 cases and 500 controls are generated in 5,000 simulated samples to estimate test size or power for a nominal 0.05-level test comparing the collective frequency of rare alleles. In each scenario, the baseline disease rate is 1%, so relative risk (RR) of 2.5 implies a penetrance of 2.5%. Rare is the number of unknown rare alleles in the population, all assumed to have the same frequency and penetrance. Freq is the total frequency of all rare alleles (e.g. 20 rare alleles with a combined frequency of 0.2 imply a frequency of 0.01 each). We make the simplifying assumption that rare alleles are mutually exclusive. Seq is the total number sequenced, either concentrated in cases or equally divided (balanced) among cases and controls. All four p-value columns are from Fisher's exact text. The first three count the number of cases and controls with any of the rare alleles detected among the indiduals that are sequenced. In the Naive and Corrected columns, all sequences are from controls, but the number of detected distinct rare alleles is subtracted from the case count in the ‘Corrected’ column. Balanced indicates that the individuals sequenced for allele detection were equally divided between cases and controls. Complete denotes the test based on sequencing all cases and all controls — a much larger sequencing effort. The parenthetic numbers indicate 25th and 75th percentiles of the number of rare alleles detected in the cases-only and balanced detection strategies.



