The paper “The critical use of statistical inference in occupational epidemiology: essay”1 is a welcome contribution to the discussion and interpretation of the results of epidemiological studies. The focus of the paper is on inferential statistics, but it touches upon several other points that are relevant to all biomedical fields, not only occupational epidemiology. I work in a hospital, and I have been collaborating with many health professionals, including physicians (occupational and not), nurses, and biologists: over the years, I have learned that they have unfounded misconceptions (I often call them myths), the most dangerous being the centrality of “statistical significance”. Based on this experience, I here offer my personal perspective on a few key epidemiological issues that I find quite relevant and often misinterpreted.
Representativeness
Many non-epidemiologists are confused about representativeness. When coming into my office, some apologize in advance by saying: “My study is selected, so it is not representative”. I try to reassure them by saying that no epidemiological study is representative of some target populations, that all studies are selected in a way or the other, and that sometimes selection is beneficial. Of course, it is important to make a clear distinction according to the objective of the study.
When the objective is just one variable (one occurrence measure, e.g., incidence or prevalence of a given disease, prevalence of smokers), representativeness is necessary2. One must study either the entire population (e.g., through a cancer registry) or conduct a survey of a representative sample of the population, usually randomly selected from an appropriate data source. By selecting an incorrect sample, one can easily obtain invalid estimates (too high or too low). The analogy with electoral polls is evident. Of course, selection or information biases may affect the study (e.g. not all people accept participating and participants may not be accurate).
On the contrary, when the interest is the association between two variables (often with the aim to investigate if a causal relationship exists between exposure and outcome), representativeness is not important2,3. In this case, what matters is to have a study sample that gives the correct exposure-disease association, that is, the right absolute (difference) or relative (ratio) measure of the association. In epidemiological jargon, what matters is the validity of the study, not whether the study sample was randomly selected to be representative of some target population.
It is easy to appreciate the irrelevance of representativeness: clinical studies on drug or vaccine efficacy are performed among non-representative groups of patients selected (non-randomly) in one or several health facilities; occupational studies are performed on (non-random) samples of workers. One eminent historical example is the study of lung cancer and smoking among British male doctors, which is clearly a case of non-representativeness. Indeed, careful selection (“intentional non-representativeness”) is sometimes instrumental in reducing bias2; for example, in order to avoid very strong confounding by active smoking, the effect of passive smoking has been validly studied by deliberately restricting enrollment to never-smokers.
The only situation in which representativeness is important occurs in case-controls studies: the controls should be representative of the study base (the population-time) from which the cases originated. Exceptions do exist, for example, when a disease registry routinely collects all cases in an area, but resources to collect control subjects are limited4.
Validity
In epidemiology, validity refers to the capacity to obtain estimates of exposure or disease (i.e., occurrence measures: incidence or prevalence) or of exposure-disease association (i.e., measures of associations: broadly speaking differences or ratios of risk and rates), which are, on average, close to the true value. Since the true value is generally unknown, the assessment of validity is indirect, consisting of evaluating the absence of important systematic errors or biases (confounding, selection, and information bias). Note that systematic errors cannot be reduced by increasing the study size.
The Authors of the essay1 correctly remarked that sufficient validity (perfection does not exist) is a fundamental prerequisite of a good epidemiological study. Indeed, in recent years a large body of techniques, called “quantitative bias analyses” have been developed to quantify confounding, selection, and information biases5. Unfortunately, these tools are still not widely used in biomedical research.
I noted that much confusion about confounding exists outside the epidemiological field. First, a common myth is that any “third variable” (beyond exposure and outcome) is a confounder to be adjusted for with multivariable analyses; not recognizing that these third variables may have different roles in the causal pathways, acting as mediators, effect modifiers, synergistic factors, or colliders, each requiring specific treatment in the design or analysis phases6-9. Second, many still think incorrectly that potential confounders are variables that were “statistically significant” in the univariate analysis, when in fact one should use other non-statistical tools to identify confounders, for example directed acyclic graphs (DAGs)6-9.
Precision
In epidemiology, precision refers to the capacity to obtain estimates of measures of occurrence or of exposure-disease association that have little variability (in hypothetical repetitions of the study). Formally, Precision = 1/Variance(Estimate), where variance is the squared standard error (SE) of the estimate. For relative measures (“relative risks”, RR), it is better to consider the SE of ln(RR). Precision reflects the amount of information (or, conversely the amount of uncertainty, the random error) included in the study. Unlike validity, precision can be increased by increasing the study size.
The statistical precision of a study is easily gauged from the width of the confidence interval (CI), which depends on SE. However, the authors of the essay1 correctly pointed out that too often CI is misinterpreted by labeling a result as “statistically significant” (or not) based on the fact that CI does not (or does) include the null value, thus degrading CI to hypothesis testing.
Statistical significance
We can say that “the main goal of a statistical analysis should be the production of the most accurate (valid and precise) effect estimates obtainable from the data”10. Unfortunately, while validity issues often have a larger impact on study accuracy, much of the emphasis of statistical teaching is still on inferential statistics to address random errors. There are two broad classes of statistical methods used in this scope: hypothesis testing and CI.
Hypothesis testing produces P-values that are largely misinterpreted. The authors of the essay1 remind us that inferential statistics are fully appropriate only when randomization is used. But most research is observational: in these situations, one may take statistics as an aid to appreciate the uncertainty in the results. Unfortunately, this is often not the case. Many books (e.g., Rothman et al., 2008)11 and papers (e.g., Sterne and Davey Smith, 2001)12 have discussed the many problems associated with P-values. I think that the single most pernicious misuse is the widespread tendency to dichotomize the P-value in “P < 0.05” (called statistically significant and thus “positive”) and “P > 0.05” (qualified as statistically not significant and thus “negative”). This behavior has been labeled dichotomania13.
The confidence interval is a much better way to express and communicate scientific results. By focusing on precision rather than statistical significance, three numbers (the effect estimate and the lower and upper confidence limits) convey both health significance and statistical uncertainty14. Unfortunately, as noted above, too often CIs are degraded to hypothesis testing, falling back in the dichotomania trap.
As correctly noted by the authors of the essay1, knowledge in science advances through replication of results under different conditions, with the aid of several other disciplines outside epidemiology (e.g. biology, toxicology). In this regard, note that the famous “Hill’s criteria” to gauge causality do not include statistical significance 11. In most situations, a single study cannot lead to a conclusive evaluation by itself. Systematic reviews, in particular quantitative systematic reviews or meta-analyses, are important tools in determining causal relationships. It would be easy to appreciate the irrelevance of statistical significance by taking a few seconds to reflect on the fact that a meta-analysis makes use of confidence intervals, not considering at all the P-values of each study.
Science is more complex than calculating a single number and comparing it to a completely arbitrary threshold7. A strong cultural change (which includes abandoning the focus on hypothesis testing) is needed in the way statistics is taught, used, and interpreted13,15. While awaiting that time, readers could follow three simple recommendations when submitting manuscripts16: 1) Present effect estimates along with their confidence intervals instead of P-values, avoiding to qualify the result as statistically significant or not based on the fact that the interval includes (or not) the null value. 2) Do not write in the methods section of the paper sentences like “We considered statistically significant a P-value < 0.05.” 3) If you report P-values, avoid labeling them as statistically significant or not; instead, evaluate them in a non-rigid, qualitative way and consider the health relevance of the findings (e.g., the magnitude of relative risk or risk difference). Irrespective of the P-values, your study may be included in a meta-analysis in the future.
I like to finish by quoting a sharp sentence from a recent Editorial in an important statistical journal17: “statistically significant – don’t say it and don’t use it”.
References
-
1 Fernandes RCP, Lima VMC, Carvalho FM. O uso crítico da inferência estatística na epidemiologia ocupacional: ensaio. Res Bras Saude Ocup 2024;49:e13. Disponível em: https://doi.org/10.1590/2317-6369/35622pt2024v49e13
» https://doi.org/10.1590/2317-6369/35622pt2024v49e13 -
2 Richiardi L, Pizzi C, Pearce N. Commentary: Representativeness is usually not necessary and often should be avoided. Int J Epidemiol. 2013 Aug;42(4):1018-22. https://doi.org/10.1093/ije/dyt103
» https://doi.org/10.1093/ije/dyt103 -
3 Rothman KJ, Gallacher JE, Hatch EE. Why representativeness should be avoided. Int J Epidemiol. 2013 Aug;42(4):1012-4. https://doi.org/10.1093/ije/dys223
» https://doi.org/10.1093/ije/dys223 -
4 Consonni D, Calvi C, De Matteis S, Mirabelli D, Landi MT, Caporaso NE, et al. Peritoneal mesothelioma and asbestos exposure: a population-based case-control study in Lombardy, Italy. Occup Environ Med. 2019 Aug;76(8):545-53. https://doi.org/10.1136/oemed-2019-105826
» https://doi.org/10.1136/oemed-2019-105826 -
5 Lash TL, Fox MP, MacLehose RF, Maldonado G, McCandless LC, Greenland S. Good practices for quantitative bias analysis. Int J Epidemiol. 2014 Dec;43(6):1969-85. https://doi.org/10.1093/ije/dyu149
» https://doi.org/10.1093/ije/dyu149 -
6 Corraini P, Olsen M, Pedersen L, Dekkers OM, Vandenbroucke JP. Effect modification, interaction and mediation: an overview of theoretical insights for clinical investigators. Clin Epidemiol. 2017 Jun;9:331-8. https://doi.org/10.2147/CLEP.S129728
» https://doi.org/10.2147/CLEP.S129728 -
7 Pearce N, Lawlor DA. Causal inference-so much more than statistics. Int J Epidemiol. 2016 Dec;45(6):1895-903. https://doi.org/10.1093/ije/dyw328
» https://doi.org/10.1093/ije/dyw328 -
8 Digitale JC, Martin JN, Glymour MM. Tutorial on directed acyclic graphs. J Clin Epidemiol. 2022 Feb;142:264-7. https://doi.org/10.1016/j.jclinepi.2021.08.001
» https://doi.org/10.1016/j.jclinepi.2021.08.001 -
9 Lipsky AM, Greenland S. Causal directed acyclic graphs. JAMA. 2022 Feb;327(11):1083-4. https://doi.org/10.1001/jama.2022.1816
» https://doi.org/10.1001/jama.2022.1816 -
10 Greenland S, Daniel R, Pearce N. Outcome modelling strategies in epidemiology: traditional methods and basic alternatives. Int J Epidemiol. 2016 Apr;45(2):565-75. https://doi.org/10.1093/ije/dyw040
» https://doi.org/10.1093/ije/dyw040 - 11 Rothman KJ, Greenland S, Lash TL. Modern epidemiology 3rd rd. Philadelphia: Lippincott Williams & Wilkins; 2008.
-
12 Sterne JA, Davey Smith G. Sifting the evidence: what's wrong with significance tests? BMJ. 2017 Jan;322(7280):226-31. https://doi.org/10.1136/bmj.322.7280.226
» https://doi.org/10.1136/bmj.322.7280.226 -
13 Greenland S. Invited commentary: the need for cognitive science in methodology. Am J Epidemiol. 2017;186(6):639-45. https://doi.org/10.1093/aje/kwx259
» https://doi.org/10.1093/aje/kwx259 -
14 Poole C. Low P-values or narrow confidence intervals: which are more durable? Epidemiology. 2001 May;12(3):291-4. https://doi.org/10.1097/00001648-200105000-00005
» https://doi.org/10.1097/00001648-200105000-00005 -
15 Lash TL. The harm done to reproducibility by the culture of null hypothesis significance testing. Am J Epidemiol. 2017 Sep;186(6):627-35. https://doi.org/10.1093/aje/kwx261
» https://doi.org/10.1093/aje/kwx261 -
16 Consonni D, Bertazzi PA. Health significance and statistical uncertainty: the value of P-value. Med Lav. 2017 oct;108(5):327-31. https://doi.org/10.23749/mdl.v108i5.6603
» https://doi.org/10.23749/mdl.v108i5.6603 -
17 Wasserstein RL, Schirm AL, Lazar NA. Moving to a world beyond “p < 0.05” Am Stat. 2019;73(sup1):1-19. https://doi.org/10.1080/00031305.2019.1583913
» https://doi.org/10.1080/00031305.2019.1583913
-
Data availability:
The entire data set supporting the results of this study has been published in the article itself.
-
Presentation at a scientific event:
The author informs that this Discussion is original and was not presented at a scientific event.
Edited by
-
Editor-in-Chief:
Eduardo Algranti
The entire data set supporting the results of this study has been published in the article itself.
