Skip to content
MDThesis

Free tool · no sign-up

Sample size calculation for a medical postgraduate thesis


Five designs, the formula written out, your values substituted into it, and a methods sentence you can paste into your synopsis. It runs in your browser, it asks for nothing, and it shows its working — because the number is only worth as much as your answer to the question that follows it.

  • Five designs
  • Formula and working shown
  • Attrition applied
  • Methods sentence you can paste
  • Nothing stored

The calculator

The instrument

Calculate it, and see how it was calculated


It opens on a worked example so you can see what a complete answer looks like before you type anything. Change any value and the figures, the working and the sentence are written again.

Step one

Which of these is your study?

Nearly every postgraduate thesis is one of these five. If two of them seem to fit, the one that matches your primary objective is the one to calculate on, and the other belongs in your limitations.

Step two

Your assumed values


From a published study in a comparable population, which you cite. Where nothing is published, 50% is the most conservative assumption and gives the largest n.

How wide a confidence interval you are willing to report, in percentage points. 5% means your estimate is reported as p ± 5.

Optional. Supply it only where your sampling frame is genuinely finite and known — one hospital's registered patients, one year's admissions. Leave it blank otherwise.


0.05 is the convention. State it in your synopsis rather than leaving it implied.

What proportion you expect to lose to follow-up, withdrawal or unusable records. Justify it from your department's own experience.

Nothing you type here leaves your browser

Step threeRecalculated as you type

The answer, and the working behind it


Required sample size

359

323 participants in total, 359 once 10% attrition is allowed for.

From the formula
323
After the attrition allowance
359

The formula

n = Zα² × p(1 − p) ÷ d²

Your values, substituted

n = (1.9600)² × 0.3 × (1 − 0.3) ÷ (0.05)²

n = 3.8416 × 0.21 ÷ 0.0025

n = 322.6944 → 323 (rounded up)

n = 323 ÷ (1 − 0.1) = 358.89 → 359 (rounded up)

Every figure is rounded up, and the rounded figure is what goes into the next step — so each number above can be re-derived from the one before it, which is exactly how it will be checked.

What this number assumes

Expected prevalence (p)
30%
From published work in a comparable population, which you cite
Absolute precision (d)
± 5 percentage points
Your choice, and you must be able to say why it is tight enough
Significance level (two-sided α)
0.05 — 95% confidence, Zα = 1.9600
A convention you state explicitly, not a value you derive
Allowance for attrition
10%
Your estimate, justified from your department's experience
For your synopsisYour citation, not ours

The methods sentence, ready to paste


This is the paragraph a protocol committee looks for: the formula, every assumed value, and both figures — before and after attrition. It deliberately leaves <your reference> in place. Replace it with the study your assumed value came from, in your university's reference style. We cannot supply that citation, and a sentence submitted with the placeholder still in it will be the first thing an examiner notices.

Your edits stay until you change one of the values above, at which point the sentence is written again from the new figures.

304 characters

Before the viva

What your examiner will ask about this number

The arithmetic is the easy half. These are the questions that actually get asked about a sample size, and what a committee is listening for in the answer.

Where did that assumed value come from?
Name the study, the population it was done in and why it is comparable to yours. A value taken from a textbook, a pilot of six patients or another specialty is the commonest place this question ends badly. If you used your own pilot, say how many patients it had.
Why that level of precision, power and alpha?
Say them out loud rather than leaving them implied: two-sided alpha, the confidence level, and the power. Then say what the precision means in practice — a precision of five percentage points on a prevalence of thirty means you will report somewhere between twenty-five and thirty-five.
What happens if you cannot recruit that many?
Answer with your department's own numbers: how many eligible patients came through last year, what proportion would have consented, and therefore how long recruitment takes. If the number is out of reach, change the design, the duration or the precision — and say which — rather than quietly reporting a smaller n than you calculated.
Is your precision absolute or relative?
This calculator uses absolute precision, in percentage points. Relative precision — ten per cent of the prevalence itself — gives a very different n, so state which one you used.
Why not assume fifty per cent?
p = 50% gives the largest possible n and needs no citation at all. It is the honest choice when nothing comparable is published, and a sound answer if you can say that is why you chose it.
This is a planning aid, not an authority. A calculator applies a formula to the values you give it. It cannot tell you whether those values are the right ones, and that is the part a viva is about. Every assumed figure must come from published work you cite or from your own department's records, and your guide and your institution's statistician have the final word on the design and the number. Nothing here has been approved by anyone, and no calculation is a promise that your protocol will be accepted.

The designs

Which formula, and when

The five designs, and what each one asks you to justify


The hard part of a sample size calculation is not the arithmetic. It is choosing the design that matches your primary objective and defending the values you put into it.

Your primary objective decides the formula. Write the objective as a single sentence first — what you are measuring, in whom, and against what — and the design usually falls out of it. Where your thesis has several objectives, the sample size is calculated on the primary one, and you say so; the secondary objectives are then reported as what the same sample can support, which is an honest position and a defensible one.

  1. Design 01

    One proportion or prevalence

    You are estimating a single figure — how common something is — to a stated margin of error. There is no comparison group.

    n = Zα² × p(1 − p) ÷ d²

    What you must supply
    An expected proportion from published work you cite, and an absolute precision — how wide a confidence interval you are prepared to report. Where the sampling frame is genuinely finite and known, a finite population correction reduces the figure.
    Where it goes wrong
    Absolute precision, in percentage points, is not the same as relative precision, a proportion of the estimate itself. They give very different numbers, so state which you used. An expected proportion of 50 per cent gives the largest possible sample and needs no citation at all.
  2. Design 02

    Two independent means

    You are comparing an average between two groups — mean HbA1c, mean duration of stay, a mean score — and the outcome is a measurement.

    n per group = 2 × (Zα + Zβ)² × σ² ÷ Δ²

    What you must supply
    A common standard deviation for your outcome, from a study using the same instrument and the same units, and the smallest difference that would change what a clinician does.
    Where it goes wrong
    Δ must be clinically meaningful, not merely detectable. Taking the largest difference in the literature shrinks the sample and is the first thing a statistician on a protocol committee notices. The formula assumes equal groups; unequal allocation needs a different one and a stated ratio.
  3. Design 03

    Two independent proportions

    You are comparing how often something happens in two groups — a complication rate, a cure rate, a positivity rate — and the outcome is yes or no.

    n per group = (Zα + Zβ)² × [p₁(1 − p₁) + p₂(1 − p₂)] ÷ (p₁ − p₂)²

    What you must supply
    The proportion you expect in each group, usually the control group from published work and the intervention group from the difference you consider worth finding.
    Where it goes wrong
    The closer the two proportions, the larger the sample, and it rises steeply. No continuity correction is applied here; where numbers are small some statisticians prefer the corrected formula, which gives a slightly larger figure. Say which you used.
  4. Design 04

    Diagnostic accuracy

    You are testing a new or index test against a reference standard and will report sensitivity and specificity.

    n = Zα² × Sn(1 − Sn) ÷ d², then total N = n ÷ prevalence

    What you must supply
    The sensitivity and specificity you expect, the precision you want on each, and the prevalence of the condition among the patients you will actually recruit.
    Where it goes wrong
    Two calculations, not one: the sensitivity arm is divided by the prevalence and the specificity arm by one minus the prevalence. Report the larger, show both, and take the prevalence from your own department's records where you can — it is what converts a number of cases into a number of patients to screen.
  5. Design 05

    Correlation between two measurements

    You are asking whether two continuous measurements move together, and will report a correlation coefficient.

    C = 0.5 × ln[(1 + r) ÷ (1 − r)]; n = [(Zα + Zβ) ÷ C]² + 3

    What you must supply
    The correlation coefficient you expect to find, from a study reporting the same pair of measurements.
    Where it goes wrong
    The assumed r does almost all the work: a small reduction in it raises the sample steeply. The formula is for Pearson's r on a reasonably linear relationship. If you plan Spearman's rho, say so, because the estimate is then approximate.

The standard normal values the formulas use

Z for alpha is two-sided; Z for beta is one-sided. These are the figures the calculator substitutes, quoted to four decimal places, and they are the ones to write in your synopsis beside the formula.

Z for the significance level — two-sided
αConfidence
0.1090%1.6449
0.0595%1.9600
0.0199%2.5758
Z for the power — one-sided
Power (1 − β)
80%0.8416
85%1.0364
90%1.2816
95%1.6449

Standard normal deviates · two-sided for α, one-sided for β

Attrition, and why it is applied last

Every design here applies the same final step: n divided by one minus your expected drop-out rate, rounded up. It comes last because the formulas give you the number of participants you need to analyse, and attrition is about how many you have to recruit to be left with that many. Applying it the other way round — adding ten per cent to the recruitment target rather than dividing by 0.9 — quietly under-recruits, and the difference grows as the drop-out rate rises.

Each figure is rounded up before it is used in the next step, so every number the calculator prints can be re-derived from the number above it. That matters more than carrying decimals: your working has to reproduce your answer in front of someone checking it with a calculator of their own.

The fuller treatment of each design, with worked examples in prose, is in the guide: sample size calculation for an MD or MS thesis, explained without fear.


The defence

Six things to write down

What a protocol committee asks about the number


A sample size is accepted or queried on the basis of its assumptions, not its arithmetic. These are the six things worth writing into the synopsis beside the figure.

Where the assumed value came from
Name the study, the population it was done in, and why it is comparable to yours. A value from a different specialty, a different country's population or a pilot of six patients is the commonest place this question ends badly. If the value is from your own records, say how many records and over what period.
Why that precision or that difference
Say what it means in practice. A precision of five percentage points on a prevalence of thirty means you will report somewhere between twenty-five and thirty-five — is that useful enough to justify the work? A difference worth detecting has to be a difference that would change management.
Alpha, power and sidedness, stated out loud
Two-sided alpha of 0.05 and 80 per cent power are the usual conventions, and the word to include is usually: they are conventions you are adopting, not facts you derived. Writing them down is what makes the calculation reproducible by whoever reads it.
The attrition allowance, and its basis
State the percentage and where it came from. Your department's own experience of losing patients to follow-up is a better justification than a round number everyone uses.
Feasibility, in your own department
How many eligible patients came through last year, what proportion would consent, and therefore how many months of recruitment the figure implies. A sample size that cannot be recruited in the time your programme allows is a problem to solve now, not after data collection has started.
The software or source you used
Most formats ask you to name the formula or the tool. You are welcome to say the figure was calculated with the formula above; whether your department wants a named statistical package instead is a question for your guide.

Where this sits in your timetable

The sample size is settled before data collection, not after, because the ethics committee approves a protocol that contains it and because approval has to precede the first patient.[5] For DNB and DrNB trainees that protocol, with ethics committee approval, is uploaded within 180 days of joining, with the thesis itself due at 26 months.[2] For MD and MS candidates under the National Medical Commission's 2023 regulations the thesis remains compulsory and is assessed within the practical and viva examination, which is where these questions get asked out loud.[1]

If the calculation says your department cannot supply the patients in the time you have, that is useful information and it has arrived early. Changing the design, the duration or the precision now is ordinary research practice. Discovering it in month twenty is not.

  • NMC PGMER-2023
  • NBEMS 180 days / 26 months
  • UGC 2018 · under 10%
  • ICMR 2017 · ethics

Questions

Common questions

Asked often, answered plainly


Which sample size formula should I use for my MD or MS thesis?
It follows from your primary objective. If you are estimating one figure, such as how common something is, use the single proportion formula with an absolute precision. If you are comparing an average between two groups, use the two independent means formula. If you are comparing a rate or a proportion between two groups, use the two proportions formula. If you are evaluating a test against a reference standard, use the diagnostic accuracy formula and divide by the prevalence in your setting. If you are asking whether two measurements move together, use the correlation formula. Where two designs seem to fit, calculate on your primary objective and say so.
Where do I get the prevalence or standard deviation to put into the formula?
From published work in a comparable population, which you cite by name in your synopsis. Your own department's records are acceptable and often better for a prevalence, provided you say how many records you looked at and over what period. A pilot of a handful of patients is not a sound source for a standard deviation. Where nothing comparable is published, 50 per cent is the most conservative assumption for a prevalence and gives the largest sample, and saying that is why you chose it is a sound answer in a viva.
How much should I add for drop-outs?
Only what you can justify. Ten per cent is commonly used for a prospective study with short follow-up, and more is reasonable where follow-up is long or the population is mobile. Justify it from your own department's experience rather than convention, and state the figure in your synopsis so the final number is explained. A retrospective record review usually needs no attrition allowance at all, but does need an allowance for incomplete records if that is a real risk.
What if my department cannot supply that many patients?
Change the design, the duration or the precision, and say which. Counting last year's eligible patients, estimating the proportion who would consent and working out how many months recruitment takes is the calculation a committee actually wants to see beside the formula. Quietly reporting a smaller number than you calculated, without saying why, is the version that causes trouble at submission.
Is this calculator enough on its own for my synopsis?
It gives you the formula, the arithmetic and a methods sentence, which is what most synopsis formats ask for. It cannot tell you whether your assumed values are the right ones, and it is not a substitute for your guide or your institution's statistician, both of whom have the final word. Where your university's format or your ethics committee asks for something specific, theirs is the requirement that governs.
Do you store what I type into the calculator?
No. The arithmetic runs in your browser. Nothing is sent to us, nothing is stored, and there is no account, no email field and no partial answer held back.

What this tool is not

It is not an approval, and it is not a statistician. It applies a published formula to the values you type and shows you the arithmetic. It cannot judge whether your assumed prevalence is the right one for your population, whether your outcome is measured well enough for the standard deviation you have used, or whether your design answers your question at all. Those are the judgements your guide and your institution's statistician are for, and their decision governs over anything on this page.

We also make no claim about whether your committee will accept a figure calculated here. Where your university publishes its own format, or your ethics committee asks for a particular method, theirs is the requirement.

The same commitments, in summary

Protection of your work

  • Row-level security

    Every table enforces row-level access. You read your own record, and nothing else.

  • View-only streaming

    Drafts are streamed to you through an authenticated route, not handed over as a file.

  • Watermarked to you

    Every page you read carries your own name and email across it.

  • Download gated

    The final file unlocks when the fee is settled in full, and not before.

  • Mumbai region · DPDP 2023

    Your record and your documents are held in the Mumbai region, so India's Digital Personal Data Protection Act 2023 applies to them.

  • Anonymised data only

    We accept no patient identifiers. An NDA is available on request.

How this works

Instruments we work to

  • NMC PGMER-2023

    The thesis obligations set out in the postgraduate medical education regulations.

  • NBEMS

    DNB and DrNB protocol and thesis timelines, and the page limit, as published.

  • UGC 2018 · <10%

    The academic integrity convention we work to on every draft.

  • ICMJE · Vancouver

    Authorship criteria and reference style, applied as published.

  • No affiliation

    We work to these published instruments. We are affiliated to none of the bodies that issue them.

How this works

Authorship and the uniqueness check

  • Sole author

    Mentoring, editing, statistics and compliance. You remain the sole author of your thesis.

  • Not ghostwriting

    We will not write your thesis for you, and we will not be named in it.

  • MDSoftune

    Word-level uniqueness checking, built with REDENN Informatics Inc., Canada.

  • Every version

    Each draft is checked word by word before your university sees it.

How this works

MDThesis is an independent academic mentorship practice. It is not affiliated with, endorsed by, or acting for the NMC, NBEMS, UGC or any university.

Free feasibility call

Got your number. Want a senior doctor to check the design behind it?


A free feasibility call: a senior doctor in your specialty looks at your objective, your design and what your department can actually recruit, and tells you plainly what to do next.

Request for a feasibility call

No obligation


No spam. A senior mentor replies personally. Your details stay private.


Reply within one working day · No obligation · Your details are not shared

Prefer to write to us first? Contact the practice. We mentor and edit; you remain the sole author of your thesis.

Document: Sample size calculator · Revision 1 · Last reviewed

Issued by MDThesis, a brand of REDENN Informatics Private Limited