'Simple random sampling' is the phrase that turns up in sampling sections where no sampling frame ever existed. It is easy for an examiner to find, because the rest of the methods usually contradicts it, and it is entirely avoidable: an accurate description of what you actually did is almost always acceptable, and the inaccurate one is not.
Probability sampling
In probability sampling every member of the population has a known, non-zero chance of being selected. That is what lets a sample stand for a population, and it is what the standard sample size formulas assume you are doing.
Simple random sampling needs a sampling frame, a numbered list of the whole eligible population, from which you draw at random using a random number generator or table. The frame is the requirement. If you cannot produce that list before you start sampling, you are not doing simple random sampling, and in a hospital-based prospective study where patients will arrive over the next eighteen months no such list can exist.
Systematic sampling works from the same frame: a random starting point, then every kth record, where k is the frame size divided by the sample size you need. This is honest and practical for record-based studies, where the frame genuinely exists, because every admission in the last three years is a list you can number. Watch for periodicity: if the list repeats a pattern, one unit's cases always falling on the same day, a fixed interval can lock onto it.
Stratified random sampling divides the population into strata that matter, sex, age band, disease severity, urban and rural, then samples at random within each. It needs a frame and it needs the strata known in advance. The payoff is precision and a guarantee that small groups appear at all. The cost is that the analysis must respect the strata, and that proportionate and disproportionate allocation are different decisions which have to be stated.
Cluster sampling samples groups rather than individuals, wards, villages, schools, anganwadi centres, and studies everyone or a random sample within each selected cluster. It is the right design for community-based work where no individual frame exists. It also carries a statistical cost, which is below.
Non-probability sampling
Consecutive sampling enrols every eligible patient who presents, in order, until the sample size is reached or the study period closes. It is what a hospital-based prospective study can actually do, and it is a recognised and defensible method. Say it. 'All consecutive patients meeting the inclusion criteria and attending the department between these two dates were enrolled' is a complete and accurate sampling statement, and it is far stronger than a random-sampling claim that collapses under one question.
Convenience sampling takes whoever is available and willing, without enrolling consecutively: whoever is in the ward when you do your round, whoever agrees. There is no shame in the phrase and a great deal of harm in disguising it. Write it, then do the two things that make it defensible. State specifically how those participants came to be the available ones, and discuss in your limitations what that selection is likely to have done to your results. An examiner accepts a stated limitation and pursues a hidden one.
Purposive sampling selects participants deliberately because of a characteristic you need. It is the correct and intended method for qualitative work and for some diagnostic studies, where information-rich cases matter more than a representative cross-section. It is a design choice and should be justified as one, not apologised for.
Quota sampling sits next to convenience sampling: you fill predetermined numbers in each category by whatever means come to hand, with no randomisation inside the categories. In a table it looks like stratified sampling and it is not, and calling it stratified is the misdescription most likely to be caught.
Snowball sampling, where participants recruit other participants, belongs to hard-to-reach populations and carries its own biases, which you state rather than hope nobody raises.
What usually gets misdescribed
- 'Simple random sampling' written where no sampling frame existed: consecutive sampling is the honest replacement
- 'Random' used to mean arbitrary or haphazard, when it has a technical meaning, a defined chance of selection
- 'Stratified' written where quotas were filled without randomisation inside each stratum
- 'Randomly allocated' confused with 'randomly sampled': allocation to groups after enrolment is a different act from selection out of a population, and a randomised trial commonly samples consecutively and allocates at random
- A sampling method named with no study period, setting or eligibility criteria alongside it, so a reader cannot tell who was available to be sampled in the first place
All five are repaired by describing what happened. The objection is to the claim, not to the method.
How sampling interacts with the sample size calculation
The standard formulas, for estimating a prevalence and for comparing two means or two proportions, assume simple random sampling. Any other design modifies the number they give you.
Cluster sampling inflates it. Because people within a cluster resemble one another, each extra person from the same cluster adds less information than an independent person would, so the required sample size is multiplied by a design effect to compensate. The design effect depends on how alike observations within a cluster are and on how many you take per cluster, and it has to be stated and justified rather than assumed. A community study using a cluster design and an unadjusted sample size is under-powered by construction.
Stratified sampling can reduce it, where the strata really are more homogeneous than the population as a whole, and the calculation is then done stratum by stratum.
Consecutive sampling leaves the formula alone but changes what the number means. You calculated n for a representative sample; what you will have is every eligible patient from one department over a fixed period. Report the number, and acknowledge in your limitations that the population sampled is your hospital's.
Run the arithmetic in the other direction too, before the protocol goes in. If your department sees a certain number of eligible patients a month and your calculated n is well beyond what the data-collection window can deliver, the design has to change: a longer period, wider inclusion criteria, a second site, or a different question. Loosening the precision or the power until the number fits is the one move that cannot be defended, and it is the first thing a statistician on the panel looks for.
Our free sample size calculator at /tools/sample-size covers the designs a thesis usually needs and shows the formula and the assumed values next to the answer, so the working can go into the synopsis instead of a bare number.
Writing the section
- Name the design: prospective observational, cross-sectional, record-based retrospective, randomised
- Name the setting and the exact dates of the study period
- Give the inclusion and exclusion criteria that define eligibility
- Name the sampling technique accurately, in one sentence that matches what you did
- For a probability method, state what the sampling frame was and how randomisation was performed
- For consecutive or convenience sampling, say so, and carry it through to your limitations
- Keep the sample size calculation in its own paragraph, with the formula, the assumed values, the source of those values and the final number
A reader should be able to say, from that paragraph alone, exactly who could have entered your study and how they were chosen. When they can, the section is finished.