Skip to content

How to Calculate Sample Size for a Clinical Study

Table of contents

To calculate sample size for a clinical study, first define the primary outcome and study design, then specify the clinically meaningful effect, expected variability or event rate, significance level, statistical power, allocation ratio, and anticipated loss to follow-up. These inputs must be entered into a formula or validated statistical program appropriate for the planned analysis.

Sample size calculation is not a single universal formula. A trial comparing two means requires a different approach from a prevalence study, diagnostic-accuracy study, survival analysis, non-inferiority trial, or cluster-randomised trial. The calculation should therefore follow the primary objective and primary statistical analysis specified in the protocol.

Why sample size matters in clinical research

A sample that is too small may have insufficient statistical power to detect a clinically important effect. This can produce an inconclusive result even when a meaningful difference exists. An unnecessarily large study may expose more participants than required, consume avoidable resources, and detect differences that are statistically significant but clinically trivial.

An appropriate sample size supports scientific validity, ethical recruitment, operational feasibility, and transparent interpretation. However, the result is only as credible as the assumptions entered into the calculation.

Information needed before calculating sample size

1. Primary research question

State the main question using a structured framework such as population, intervention or exposure, comparator, and outcome. The sample size should normally be based on one clearly designated primary outcome rather than whichever secondary outcome produces the largest or most convenient number.

2. Study design

Identify whether the study is a cross-sectional survey, cohort study, case-control study, parallel randomised trial, crossover trial, diagnostic study, survival study, cluster-randomised trial, superiority trial, equivalence trial, or non-inferiority trial. Each design has different statistical requirements.

3. Outcome type

Determine how the primary outcome will be measured:

  • Continuous: blood pressure, weight, laboratory concentration, or symptom score.
  • Binary: death, response, infection, readmission, or treatment success.
  • Time-to-event: time to death, relapse, or disease progression.
  • Count: number of hospital admissions or exacerbations.
  • Diagnostic: sensitivity, specificity, or area under the receiver operating characteristic curve.

4. Clinically meaningful effect

The target effect is the smallest difference that would be important enough to influence clinical practice, policy, or future research. For a continuous outcome, this may be a mean difference. For a binary outcome, it may be an absolute risk reduction, risk ratio, or odds ratio.

The target difference should not be chosen only because it produces an affordable sample. It should be justified using clinical expertise, previous trials, systematic reviews, observational data, pilot work, patient perspectives, or a recognised minimal clinically important difference.

5. Significance level

The significance level, denoted by alpha, is the probability of a type I error under the statistical model: rejecting the null hypothesis when it is true. A two-sided alpha of 0.05 is common, although the correct value depends on the design, number of primary comparisons, regulatory context, and analysis plan.

6. Statistical power

Power is the probability of detecting the specified effect when that effect truly exists under the assumptions of the calculation. It equals 1 minus beta, where beta is the probability of a type II error. Values of 80% or 90% are frequently used. Increasing power increases the required sample size.

7. Variability or expected event rate

Calculations for continuous outcomes require an estimate of the standard deviation. Calculations for binary outcomes require expected proportions or event risks. Survival analyses require assumptions about the event rate, follow-up duration, recruitment period, censoring, and treatment effect.

These values should come from the most applicable evidence available. An unrealistically small standard deviation or exaggerated treatment effect can substantially underestimate the required sample.

8. Allocation ratio

A two-group clinical trial often uses equal allocation, such as 1:1, because this usually provides the greatest power for a fixed total sample when group costs and variances are similar. Unequal allocation may be reasonable when one treatment is more expensive, safety data are especially valuable, or recruitment constraints differ between groups.

9. Expected dropout or missing data

The number that must be recruited is usually larger than the number required for the final analysis. Investigators should account for withdrawals, loss to follow-up, non-adherence, unusable measurements, and other causes of missing outcome data.

Core sample size formulas

The following formulas illustrate common situations. They assume relatively simple designs and should not replace design-specific statistical advice for complex clinical studies.

Estimating a single proportion

For a cross-sectional study estimating prevalence or another proportion:

n = Z² × p(1 − p) ÷ d²

Where:

  • n is the required sample size.
  • Z is the standard normal value for the confidence level, commonly 1.96 for a 95% confidence interval.
  • p is the expected proportion.
  • d is the desired absolute margin of error.

Example: estimating prevalence

Suppose researchers expect a prevalence of 30% and want a 95% confidence interval with a margin of error of 5 percentage points:

n = 1.96² × 0.30 × 0.70 ÷ 0.05²

n = 322.69

The result is rounded upward to 323 participants before adjusting for non-response or design effects.

Estimating a single mean

When estimating a population mean with a specified precision:

n = Z² × SD² ÷ d²

Here, SD is the expected standard deviation and d is the desired half-width of the confidence interval.

Comparing two independent means

For two equally sized groups with a continuous outcome, a commonly used approximation is:

n per group = 2 × (Z1−α/2 + Z1−β)² × SD² ÷ Δ²

Where:

  • SD is the assumed common standard deviation.
  • Δ is the difference the study is designed to detect.
  • Z1−α/2 reflects the two-sided significance level.
  • Z1−β reflects the desired power.

Example: comparing two means

Assume a trial is designed to detect a mean difference of 5 units, with an expected standard deviation of 10 units, a two-sided alpha of 0.05, and 80% power. The approximate normal values are 1.96 and 0.84:

n per group = 2 × (1.96 + 0.84)² × 10² ÷ 5²

n per group = 62.72

Rounding upward gives 63 participants per group, or 126 in total, before accounting for dropout.

Comparing two independent proportions

For equal group sizes, one commonly used form is:

n per group = [Z1−α/2√(2p̄(1−p̄)) + Z1−β√(p₁(1−p₁) + p₂(1−p₂))]² ÷ (p₁−p₂)²

Where p₁ and p₂ are the expected proportions in the two groups and is their average.

The absolute difference between the expected event rates is critical. Detecting a small absolute risk reduction generally requires a much larger sample than detecting a large reduction.

Paired or before-and-after studies

Paired designs analyse the within-person difference rather than treating observations as independent. Their sample size depends on the standard deviation of the paired differences or, equivalently, the correlation between repeated measurements. Using an independent-groups formula for paired data may give an inappropriate result.

Survival and time-to-event studies

Many survival studies are driven primarily by the required number of outcome events rather than only the number of recruited participants. The calculation commonly incorporates the hazard ratio, significance level, power, allocation ratio, anticipated event incidence, recruitment period, follow-up duration, and censoring.

Once the required number of events has been calculated, investigators estimate how many participants must be enrolled to observe that number of events within the study period.

How to adjust for dropout

Use the following adjustment when a proportion of recruited participants is expected to be unavailable for the primary analysis:

Adjusted sample = required analysable sample ÷ (1 − dropout proportion)

For example, if 126 analysable participants are needed and 15% dropout is expected:

Adjusted sample = 126 ÷ 0.85 = 148.24

The study should recruit at least 149 participants. Simply increasing 126 by 15% would produce 145, which would leave fewer than 126 participants if exactly 15% were lost.

Finite population correction

If sampling without replacement from a small, clearly defined population, an initial sample estimate may be reduced using:

nadjusted = n₀ ÷ [1 + ((n₀ − 1) ÷ N)]

Here, n₀ is the initial sample estimate and N is the total eligible population. This correction is mainly relevant when the proposed sample represents a substantial fraction of the source population. It is generally unnecessary when the population is large or conceptually open-ended.

Important adjustments for complex clinical studies

Cluster-randomised trials

Participants within the same clinic, hospital, school, or community may have correlated outcomes. The individually randomised sample is therefore multiplied by a design effect:

Design effect = 1 + (m − 1) × ICC

Where m is the average cluster size and ICC is the intracluster correlation coefficient. Calculations should also consider the number of clusters, variation in cluster size, cluster-level dropout, and the planned analytical model.

Non-inferiority and equivalence studies

These designs require a pre-specified clinical margin and assumptions that differ from ordinary superiority testing. The margin must be clinically justified and supported by appropriate historical evidence. The calculation should match the planned confidence-interval or hypothesis-testing approach.

Multiple primary outcomes or comparisons

Testing several primary hypotheses can increase the overall type I error rate. Multiplicity adjustments may reduce the alpha available for each comparison and increase the required sample. The protocol should define whether outcomes are co-primary, hierarchical, or independently confirmatory.

Covariate-adjusted analyses

Adjustment for strongly prognostic baseline variables may improve precision, but the expected gain should be justified conservatively. The sample size method must be compatible with the planned regression model rather than relying on an unsupported reduction.

Repeated measurements

Longitudinal studies require assumptions about within-participant correlation, the number and timing of observations, covariance structure, missing measurements, and the intended mixed-effects or repeated-measures analysis. Multiplying the number of participants by the number of visits does not create an equivalent number of independent observations.

Diagnostic-accuracy studies

Sample size may be calculated separately for sensitivity and specificity. The number of participants with and without the target condition depends on the desired confidence-interval precision, expected accuracy, and disease prevalence. A low prevalence may require a large screened population to obtain enough participants with the condition.

Multivariable prediction models

Simple rules such as a fixed number of events per predictor can be inadequate. Sample size planning should consider outcome frequency, number of candidate predictor parameters, anticipated model performance, overfitting, shrinkage, and the intended validation strategy.

Step-by-step sample size calculation workflow

  1. Write the primary objective. Define exactly what parameter, difference, association, or treatment effect the study will estimate or test.
  2. Specify the primary outcome. State its measurement scale and assessment time.
  3. Choose the study design and primary analysis. The sample size method should match the intended statistical test or model.
  4. Define the target effect. Justify the minimum clinically important difference, expected risk difference, hazard ratio, accuracy parameter, or other effect.
  5. Estimate nuisance parameters. Obtain plausible values for standard deviation, event risk, prevalence, correlation, intracluster correlation, or censoring.
  6. Select alpha and power. State whether testing is one-sided or two-sided and address multiplicity where relevant.
  7. Set the allocation ratio. Include any stratification, clustering, or unequal randomisation assumptions.
  8. Calculate the analysable sample. Use a validated formula, statistical program, simulation, or specialist software.
  9. Inflate for missing data. Account for dropout, unusable samples, loss of clusters, and expected non-response.
  10. Perform sensitivity analyses. Repeat the calculation using several plausible effect sizes, standard deviations, event rates, or dropout assumptions.
  11. Check feasibility. Compare the recruitment target with the eligible population, recruitment rate, study duration, budget, and site capacity.
  12. Document every assumption. Record the formula or software, version, statistical test, inputs, rationale, and final rounding decision.

Sample size calculation versus power calculation

An a priori sample size calculation asks how many participants are needed to achieve a chosen power for a specified effect. A power calculation may also examine the power achievable with a fixed feasible sample.

Post hoc power calculated using the observed effect after a study has finished is generally not a useful substitute for confidence intervals and direct interpretation of the observed estimate. The achieved result and its uncertainty should be reported instead.

Common mistakes to avoid

  • Using a prevalence formula for an analytical study comparing groups.
  • Basing the calculation on a secondary outcome while describing another outcome as primary.
  • Using an unrealistically large expected treatment effect.
  • Taking a standard deviation from a population that differs substantially from the planned study population.
  • Confusing relative risk reduction with absolute risk reduction.
  • Using a one-sided test without a strong scientific justification.
  • Failing to account for clustering, repeated measurements, unequal allocation, or multiplicity.
  • Adding the dropout percentage directly instead of dividing by one minus the dropout proportion.
  • Choosing a convenient sample first and reverse-engineering assumptions to justify it.
  • Reporting only the final number without documenting the inputs and calculation method.

How to report a sample size calculation

A transparent protocol or manuscript should report:

  • The primary outcome and primary comparison.
  • The expected outcome values in each group.
  • The target effect and its clinical justification.
  • The assumed standard deviation, event rate, correlation, or hazard ratio.
  • The significance level and whether the test is one-sided or two-sided.
  • The target power.
  • The allocation ratio.
  • Adjustments for multiplicity, clustering, or repeated measurements.
  • The expected dropout or non-evaluable proportion.
  • The formula, software, package, version, or simulation method used.
  • The final number per group and total recruitment target.

Example reporting statement

The study was designed to detect a 5-unit difference in the primary continuous outcome between two equally allocated groups. Assuming a common standard deviation of 10 units, a two-sided significance level of 0.05, and 80% power, 63 analysable participants were required per group. Allowing for 15% loss to follow-up, the recruitment target was increased to 75 participants per group, giving a total target of 150.

Which software can calculate sample size?

Common options include R, Stata, SAS, PASS, nQuery, G*Power, Epi Info, OpenEpi, and validated design-specific online calculators. Software should be selected according to the study design and primary analysis. Researchers should preserve the input settings and, where possible, independently verify important calculations.

Simulation may be preferable when the design includes complex longitudinal models, adaptive features, multiple levels of clustering, competing risks, unusual outcome distributions, or several interacting assumptions.

When to involve a statistician

Statistical input is particularly valuable before data collection for non-inferiority or equivalence studies, cluster trials, survival analyses, adaptive trials, diagnostic studies, repeated-measures designs, prediction modelling, multiple primary outcomes, rare events, and regulatory clinical trials.

A statistician should help align the clinical question, estimand, design, primary analysis, missing-data strategy, and sample size assumptions. Consulting only after recruitment has started may make important design problems difficult or impossible to correct.

Key conclusion

Reliable sample size calculation begins with a clearly defined primary question, not with a formula. Investigators must select assumptions that are clinically defensible, statistically compatible with the planned analysis, and realistic for the study population. The final target should then be adjusted for expected losses and tested under alternative assumptions before recruitment begins.

Medical disclaimer: This article provides general educational information about clinical research methods. It does not replace consultation with a qualified biostatistician, research ethics committee, regulatory specialist, or clinical research professional for a specific protocol.

Key takeaways

  • Base the calculation on the primary objective, primary outcome, study design, and planned statistical analysis.
  • Specify and justify the target effect, alpha, power, variability or event rate, and allocation ratio.
  • Use a design-specific method because no single sample size formula applies to every clinical study.
  • Adjust the analysable sample for dropout by dividing by one minus the expected loss proportion.
  • Perform sensitivity analyses because uncertain effect sizes, standard deviations, and event rates can materially change the result.
  • Document every assumption, calculation method, software setting, and adjustment in the protocol.

Frequently asked questions

What factors determine sample size in a clinical study?
The main factors are the study design, primary outcome, clinically meaningful effect, significance level, statistical power, outcome variability or event rate, allocation ratio, and expected dropout. Complex studies may also require adjustments for clustering, repeated measurements, censoring, or multiple comparisons.
Is 30 participants enough for a clinical study?
There is no universal rule that 30 participants are sufficient. The required number depends on the study objective, expected effect, variability, analysis method, power, and desired precision. A sample of 30 may be suitable for some feasibility work but inadequate for many confirmatory clinical studies.
Why are 80% and 90% power commonly used?
These values limit the assumed type II error rate to 20% or 10%, respectively, for the effect specified in the calculation. Ninety percent power provides a lower risk of missing the target effect but requires a larger sample than 80% power.
How should sample size be adjusted for dropout?
Divide the required analysable sample by one minus the expected dropout proportion. For example, 200 required participants with an expected 20% dropout gives 200 divided by 0.80, producing a recruitment target of 250.
Can sample size be calculated without previous studies?
Yes, but uncertainty is greater. Researchers may use a justified clinically important difference, pilot data, plausible parameter ranges, expert consensus, or precision-based methods. Sensitivity analyses should show how the required sample changes under alternative assumptions.
Should every clinical study include a formal sample size calculation?
Most confirmatory and analytical studies should provide a justified sample size or precision assessment. Pilot and feasibility studies may instead base their size on feasibility objectives, parameter estimation, or operational constraints, but the rationale should still be stated clearly.

References

  1. Serdar CC, Cihan M, Yücel D, Serdar MA. Sample size, power and effect size revisited: simplified and practical approaches in pre-clinical, clinical and laboratory studies. Biochemia Medica. 2021;31(1):010502. https://pmc.ncbi.nlm.nih.gov/articles/PMC7745163/
  2. Wang X, Ji X. Sample Size Estimation in Clinical Research: From Randomized Controlled Trials to Observational Studies. Chest. 2020;158(1 Suppl):S12-S20. https://pubmed.ncbi.nlm.nih.gov/32658647/
  3. Charan J, Biswas T. How to Calculate Sample Size for Different Study Designs in Medical Research? Indian Journal of Psychological Medicine. 2013;35(2):121-126. https://pmc.ncbi.nlm.nih.gov/articles/PMC3775042/
  4. Suresh KP, Chandrashekara S. Sample size estimation and power analysis for clinical research studies. Journal of Human Reproductive Sciences. 2012;5(1):7-13. https://pmc.ncbi.nlm.nih.gov/articles/PMC3409926/
  5. Althubaiti A. Sample size determination: A practical guide for health researchers. Journal of General and Family Medicine. 2023;24(2):72-78. https://pmc.ncbi.nlm.nih.gov/articles/PMC10000262/
  6. Cook JA, Julious SA, Sones W, et al. DELTA2 guidance on choosing the target difference and undertaking and reporting the sample size calculation for a randomised controlled trial. BMJ. 2018;363:k3750. https://www.bmj.com/content/363/bmj.k3750
  7. International Council for Harmonisation of Technical Requirements for Pharmaceuticals for Human Use. ICH E9: Statistical Principles for Clinical Trials. European Medicines Agency. 1998. https://www.ema.europa.eu/en/ich-e9-statistical-principles-clinical-trials-scientific-guideline
  8. Hopewell S, Chan AW, Collins GS, et al. CONSORT 2025 explanation and elaboration: updated guideline for reporting randomised trials. BMJ. 2025;389:e081124. https://www.bmj.com/content/389/bmj-2024-081124

Advertisement

References

  1. [1] Replace with the first reference. Include title, authors, journal/source, year, and DOI or URL.
  2. [2] Replace with the second reference.
  3. [3] Replace with the third reference.

[aw_author_credentials]

Previous post
Next post