遇见数据集

Data from: The reproducibility of research and the misinterpretation of P values

收藏
DataONE2017-11-02 更新2024-06-26 收录
数据链接:
官方服务:

资源简介:

We wish to answer this question If you observe a “significant” P value after doing a single unbiased experiment, what is the probability that your result is a false positive?. The weak evidence provided by P values between 0.01 and 0.05 is explored by exact calculations of false positive risks. When you observe P = 0.05, the odds in favour of there being a real effect (given by the likelihood ratio) are about 3:1. This is far weaker evidence than the odds of 19 to 1 that might, wrongly, be inferred from the P value. And if you want to limit the false positive risk to 5 %, you would have to assume that you were 87% sure that there was a real effect before the experiment was done. If you observe P = 0.001 in a well-powered experiment, it gives a likelihood ratio of almost 100:1 odds on there being a real effect. That would usually be regarded as conclusive, But the false positive risk would still be 8% if the prior probability of a real effect were only 0.1. And, in this case, if you wanted to achieve a false positive risk of 5% you would need to observe P = 0.00045. It is recommended that the terms “significant” and “non-significant” should never be used. Rather, P values should be supplemented by specifying the prior probability that would be needed to produce a specified (e.g. 5%) false positive risk. It may also be helpful to specify the minimum false positive risk associated with the observed P value. Despite decades of warnings, many areas of science still insist on labelling a result of P < 0.05 as “statistically significant. This practice must contribute to the lack of reproducibility in some areas of science. This is before you get to the many other well-known problems, like multiple comparisons, lack of randomisation and P-hacking. Precise inductive inference is impossible and replication is the only way to be sure, Science is endangered by statistical misunderstanding, and by senior people who impose perverse incentives on scientists.

我们旨在解答如下问题:若在单次无偏实验中得到具有「显著性」的P值,那么该结果为假阳性的概率是多少?针对P值介于0.01与0.05之间的弱证据效力,可通过精确计算假阳性风险展开探讨。当观测到P=0.05时,支持真实效应存在的优势比(由似然比给出)约为3:1。这一证据效力远弱于从该P值错误推断出的19:1的优势比。若希望将假阳性风险限制在5%以内,则需在实验开展前便认定真实效应存在的先验置信度为87%。若在效能充足的实验中观测到P=0.001,其对应的真实效应存在的似然比优势比接近100:1,通常可被视为决定性证据。但倘若真实效应存在的先验概率仅为0.1,此时假阳性风险仍高达8%。在此情形下,若要将假阳性风险降至5%,则需观测到P=0.00045。建议永远不再使用「显著」与「不显著」这类表述,取而代之的是,应通过补充说明为达到特定(例如5%)假阳性风险所需的先验概率,来完善P值的解读。此外,指明与观测到的P值相关的最小假阳性风险也会有所助益。尽管已有数十年的警示,诸多科学领域仍坚持将P<0.05的结果标注为「具有统计显著性」。这一做法无疑加剧了部分科学领域的结果不可重复性问题。上述问题尚未涉及诸多广为人知的其他缺陷,例如多重比较、随机化缺失以及P值操纵等问题。精准的归纳推理并不现实,重复实验才是确证结果的唯一途径。统计学认知误区,以及那些向科研人员施加不当激励的资深人士,正将科学置于险境之中。

创建时间:
2017-11-02
二维码
社区交流群
二维码
科研交流群
商业服务