A p value measures how incompatible the observed data are with a specified statistical model, usually one that assumes no effect or no difference. It does not tell researchers the probability that the null hypothesis is true, whether a result occurred by chance, or whether an effect is clinically important.
What Is a P Value?
A p value is calculated after defining a statistical hypothesis, selecting an appropriate test, and analysing the observed data. In conventional null hypothesis significance testing, the calculation begins with a null hypothesis such as “there is no difference between the treatment groups.”
The p value is the probability of obtaining results at least as extreme as those observed, assuming the null hypothesis and the other assumptions of the statistical model are correct.
This definition has two essential conditions:
- The calculation assumes the null hypothesis for the test.
- The calculation also depends on assumptions about sampling, measurement, study design, probability distributions, model specification, and statistical analysis.
A small p value suggests that the data are relatively incompatible with the tested model. It does not, by itself, identify why that incompatibility occurred.
How to Interpret a P Value
Example: P = 0.03
Suppose a clinical study compares two treatments and reports p = 0.03 for the difference in mean blood pressure.
A reasonable interpretation is:
If there were truly no difference between the treatments, and if the statistical model and its assumptions were correct, results at least as extreme as those observed would occur in approximately 3% of comparable repetitions.
It would be incorrect to say that there is a 3% probability that the null hypothesis is true. A conventional p value calculates a probability of data under a hypothesis; it does not calculate the probability of the hypothesis given the data.
What Does P < 0.05 Mean?
The threshold of 0.05 is commonly used to label findings as “statistically significant.” When researchers select an alpha level of 0.05 before analysing the data, a p value below that threshold leads them to reject the null hypothesis under the chosen decision rule.
However, 0.05 is a convention rather than a universal boundary separating true findings from false ones. Results with p = 0.049 and p = 0.051 provide nearly the same statistical information, even though rigid threshold-based language may place them in different categories.
The p value should generally be reported as an exact value when practical, such as p = 0.032, rather than only as p < 0.05. Very small values may be reported using an appropriate bound, such as p < 0.001.
What a P Value Does Not Tell You
It Does Not Give the Probability That the Null Hypothesis Is True
A statement such as “p = 0.04 means there is a 4% chance of no effect” reverses the conditional probability. Estimating the probability of a hypothesis requires additional assumptions and a different inferential framework, such as a Bayesian analysis.
It Does Not Give the Probability That the Result Was Due to Chance
Chance is not a single competing explanation that can be assigned the p value. Random sampling variation is incorporated into the statistical model, but bias, confounding, measurement error, selective reporting, missing data, and model misspecification may also influence the result.
It Does Not Measure Effect Size
A p value does not indicate whether a difference is large or small. The same effect can produce different p values depending on sample size, outcome variability, study design, and analytical precision.
Researchers should report an effect measure appropriate to the research question, such as:
- Mean difference
- Risk difference
- Risk ratio
- Odds ratio
- Hazard ratio
- Correlation coefficient
- Standardised mean difference
It Does Not Establish Clinical Importance
A statistically detectable difference may be too small to matter to patients, clinicians, or health systems. Conversely, a clinically important effect may produce a p value above 0.05 when the study is small or imprecise.
Clinical interpretation should consider the estimated magnitude of benefit or harm, confidence interval, minimum clinically important difference, adverse effects, costs, feasibility, and patient preferences.
It Does Not Prove Causation
A small p value cannot correct weaknesses in study design. An association may still be explained by confounding, selection bias, information bias, reverse causation, analytical flexibility, or other systematic errors.
It Does Not Measure Replicability
A single p value does not state the probability that another study will obtain the same finding. Replication depends on the true effect, sample size, population, measurement quality, study procedures, analytical decisions, and random variation.
Common Misuses of P Values
Treating Statistical Significance as a Yes-or-No Truth Test
Dividing findings into “significant” and “not significant” can hide important differences in effect magnitude and precision. Evidence is continuous, while the 0.05 threshold creates an artificial category boundary.
Concluding That No Significant Difference Means No Difference
A p value above 0.05 does not demonstrate equivalence or prove that the null hypothesis is correct. It may reflect limited statistical power, high variability, measurement error, or a wide confidence interval that remains compatible with meaningful benefit and harm.
Equivalence and non-inferiority questions require appropriately designed studies, predefined margins, and suitable statistical analyses.
Comparing Two Effects by Comparing Their P Values
One subgroup may have p < 0.05 while another has p > 0.05, but that does not prove that the effects differ between the subgroups. The appropriate approach is to directly estimate and test the interaction or difference between effects.
Using P Values Without Checking Assumptions
Every statistical test relies on assumptions. Depending on the method, these may concern independence, distributional form, proportional hazards, linearity, variance, sampling, missingness, or correct model specification. A precisely calculated p value from an unsuitable model may still be misleading.
Testing Many Hypotheses Without Adjustment
When researchers perform many tests, the probability of obtaining at least one small p value increases. For example, if 20 independent null hypotheses are tested at an alpha level of 0.05, the expected number of false-positive results is one, although the realised number may differ.
Possible responses include:
- Predefining primary and secondary outcomes
- Limiting unnecessary analyses
- Using family-wise error rate procedures when appropriate
- Controlling the false discovery rate in suitable exploratory settings
- Clearly identifying analyses as confirmatory or exploratory
P-Hacking and Selective Reporting
P-hacking refers to analytical practices that increase the likelihood of obtaining a desirable p value. Examples include repeatedly changing exclusion criteria, outcome definitions, covariates, subgroups, transformations, or stopping rules after examining the results.
Selective publication of analyses with small p values can distort the scientific literature. Preregistration, registered reports, protocol transparency, data sharing where ethical, and complete outcome reporting can reduce these risks.
Why Sample Size Affects P Values
A statistical test evaluates an estimated effect relative to its uncertainty. Larger studies often produce smaller standard errors, making even modest differences capable of generating small p values.
Small studies may produce large estimated effects with wide uncertainty and p values above 0.05. Large studies may produce very small p values for effects that have little practical importance.
This is why the p value should never be interpreted without the effect estimate and its confidence interval.
P Values and Confidence Intervals
A confidence interval presents a range of effect values that are reasonably compatible with the data and the statistical model. It therefore provides more information about magnitude and precision than a p value alone.
For a conventional two-sided test, a 95% confidence interval that excludes the null value generally corresponds to p < 0.05. The null value depends on the measure:
| Effect measure | Typical null value |
|---|---|
| Mean difference | 0 |
| Risk difference | 0 |
| Risk ratio | 1 |
| Odds ratio | 1 |
| Hazard ratio | 1 |
| Correlation coefficient | 0 |
Confidence intervals should also be interpreted cautiously. They depend on the same data quality and modelling assumptions as the corresponding hypothesis test, and they do not automatically account for bias or multiple analyses.
Type I and Type II Errors
Type I Error
A Type I error occurs when a testing procedure rejects a true null hypothesis. The predefined alpha level controls the long-run Type I error rate under the assumptions of the procedure. Alpha is not the probability that a specific statistically significant finding is false.
Type II Error
A Type II error occurs when a testing procedure fails to reject the null hypothesis even though a relevant effect exists. Statistical power is the probability that the procedure will reject the null hypothesis under a specified alternative effect and set of assumptions.
Power depends on factors including sample size, outcome variability, effect magnitude, study design, significance level, and statistical method.
Better Ways to Report Statistical Results
A complete statistical report should answer more than whether a threshold was crossed. Researchers should usually report:
- The research question and prespecified hypothesis
- The estimated effect and its units
- An appropriate confidence interval
- The exact p value when useful
- The statistical method and important assumptions
- Absolute effects as well as relative effects when clinically relevant
- Whether the analysis was primary, secondary, exploratory, or post hoc
- Any adjustment for multiplicity
- Missing-data methods and sensitivity analyses
- Clinical or practical interpretation
Weak Reporting Example
“Treatment A was significantly better than treatment B, p < 0.05.”
More Informative Reporting Example
“The mean outcome was 4.2 points lower with Treatment A than with Treatment B (95% confidence interval 1.1 to 7.3 points lower; p = 0.008). The prespecified clinically important difference was 5 points, so the estimate suggests possible benefit, although the interval includes effects below and above that threshold.”
The second version communicates direction, magnitude, precision, statistical evidence, and clinical context.
Practical Checklist for Interpreting a P Value
- Identify the exact null hypothesis being tested.
- Confirm that the test matches the study design and variable types.
- Check whether important model assumptions were assessed.
- Review the effect estimate rather than the p value alone.
- Examine the confidence interval for precision and clinically relevant values.
- Consider sample size and statistical power.
- Check how many outcomes, subgroups, and models were examined.
- Determine whether the analysis was prespecified or exploratory.
- Assess risks of bias, confounding, missing data, and measurement error.
- Interpret the result alongside prior evidence and biological or clinical plausibility.
Should Researchers Stop Using P Values?
P values can be useful when they answer a clearly defined question within a well-designed analysis. The main problem is not the calculation itself but the practice of treating one threshold as a substitute for scientific reasoning.
Some researchers advocate abandoning declarations of statistical significance, while others support retaining p values with more careful interpretation. A practical approach is to avoid binary conclusions, report effect sizes and uncertainty, describe analytical decisions transparently, and interpret findings within the full clinical and scientific context.
Conclusion
P values describe the compatibility of observed data with a specified statistical model; they do not provide the probability that a hypothesis is true. Meaningful interpretation requires effect sizes, confidence intervals, study quality, model assumptions, multiplicity, prior evidence, and clinical relevance. Researchers should use p values as one component of an analysis rather than as a final verdict on whether a finding is real or important.
Medical disclaimer: This article is intended for education about medical research methods. It does not provide individual medical advice or replace consultation with qualified statistical, research, or healthcare professionals.
Key takeaways
- A p value is calculated from the probability of the observed or more extreme data under a specified null model; it is not the probability that the null hypothesis is true.
- The threshold of 0.05 is a convention and should not be treated as a sharp boundary between true and false findings.
- P values do not measure effect size, clinical importance, causality, study quality, or the probability of replication.
- Effect estimates, confidence intervals, study design, assumptions, bias, and multiplicity should be evaluated alongside p values.
- A result above 0.05 does not prove that no effect exists, and a result below 0.05 does not prove that an important effect exists.
- Transparent prespecification and complete reporting help reduce p-hacking, selective analysis, and misleading conclusions.
Frequently asked questions
What does a p value mean in medical research?
Does p less than 0.05 prove that a result is real?
Does a high p value prove that there is no difference?
Is statistical significance the same as clinical significance?
Why should p values be reported with confidence intervals?
What is p-hacking?
References
- Wasserstein RL, Lazar NA. The ASA Statement on p-Values: Context, Process, and Purpose. The American Statistician. 2016;70(2):129-133. https://doi.org/10.1080/00031305.2016.1154108
- American Statistical Association. ASA Statement on Statistical Significance and P-Values. 2016. https://www.amstat.org/asa/files/pdfs/p-valuestatement.pdf
- Greenland S, Senn SJ, Rothman KJ, Carlin JB, Poole C, Goodman SN, Altman DG. Statistical tests, P values, confidence intervals, and power: a guide to misinterpretations. European Journal of Epidemiology. 2016;31(4):337-350. https://doi.org/10.1007/s10654-016-0149-3
- Wasserstein RL, Schirm AL, Lazar NA. Moving to a World Beyond p < 0.05. The American Statistician. 2019;73(sup1):1-19. https://doi.org/10.1080/00031305.2019.1583913
- Amrhein V, Greenland S, McShane B. Scientists rise up against statistical significance. Nature. 2019;567:305-307. https://doi.org/10.1038/d41586-019-00857-9
- Rothman KJ. Disengaging from statistical significance. European Journal of Epidemiology. 2016;31(5):443-444. https://doi.org/10.1007/s10654-016-0158-2
- Aguinis H, Vassar M, Wayant C. On reporting and interpreting statistical significance and p values in medical research. BMJ Evidence-Based Medicine. 2021;26(2):39-42. https://doi.org/10.1136/bmjebm-2019-111264