Skip to content
MDThesis

Statistics · Data analysis · Methodology

Choosing the right statistical test for your thesis data


A decision path for postgraduate data: how many groups, paired or independent, what kind of outcome, and whether it is normally distributed — plus what a p value means.

  • Entry no. 08 of 18
  • Published
  • 6 min read

The guide

Almost every statistical test a postgraduate thesis needs can be reached by answering four questions in order. Answer them about your primary objective before you look at any menu of tests, because the test follows from the design, and the design was fixed when you wrote your protocol.

The four questions that decide it

  • What kind of outcome variable is it? Categorical (present or absent, grade I to IV) or numerical (a measurement)?
  • How many groups are being compared? One, two, or more than two?
  • Are the groups independent, or are the observations paired — the same patients measured twice, or matched pairs?
  • If the outcome is numerical, is it approximately normally distributed within each group?

Write the four answers down for your primary objective. In most theses they select exactly one test, which means your statistical analysis section can name it instead of listing everything you might conceivably use.

Categorical outcome: chi-square and Fisher's exact

Comparing proportions between independent groups, a complication rate in group A against group B, is a chi-square test on a contingency table. The test relies on expected counts being large enough. The usual convention is to move to Fisher's exact test when an expected count in any cell falls below five, which happens constantly in thesis-sized tables. Fisher's exact is not the weaker option; it is the correct one for small counts, and every statistical package will compute it.

Paired categorical data, the same patient before and after or two tests run on the same patient, is McNemar's test and not chi-square. This is an easy error to make because the table looks identical, and an easy one for a statistician to spot.

Report counts and percentages alongside the p value. A percentage with no denominator is not a result.

Numerical outcome, two groups

Two independent groups with a normally distributed outcome: the unpaired, or independent samples, t test, and where the two groups' variances are clearly unequal the Welch version of it, which most packages offer and which costs nothing to prefer. The same patients measured twice, or matched pairs: the paired t test. Where the distribution is not normal, or the sample is too small to tell, the non-parametric equivalents are Mann-Whitney U for independent groups and Wilcoxon signed-rank for paired data.

Non-parametric does not mean second best. It means you have not assumed a distribution you cannot demonstrate. With twenty patients an arm and a skewed outcome, and length of stay, duration of ventilation, cost and many biochemical markers are skewed, Mann-Whitney is the honest test, and an examiner who knows statistics will prefer it to a t test applied hopefully. Report median and interquartile range with it, not mean and standard deviation.

More than two groups

Three or more independent groups with a normal outcome: one-way ANOVA, followed, only if the ANOVA is significant, by a post hoc test to find which pairs differ. Tukey, Bonferroni and Scheffé are the ones software usually offers. For non-normal data: Kruskal-Wallis, with Dunn's test or Mann-Whitney plus a correction for the pairwise comparisons.

Do not run three separate t tests across three groups. Each comparison carries its own chance of a false positive and running several multiplies it. That is precisely what the post hoc correction exists to handle, and skipping it is visible in the results table.

Repeated measurements on the same patients at three or more time points need repeated-measures ANOVA, or the Friedman test if the outcome is not normal, rather than a series of paired tests.

Normality: decide it, do not guess it

Decide before analysis and on evidence. Shapiro-Wilk is the usual formal test at thesis sample sizes, with Kolmogorov-Smirnov also offered by most software, and both should be read next to a histogram and a Q-Q plot rather than instead of them. A significant Shapiro-Wilk result means the data depart from normality; a non-significant one in a small sample means only that you could not detect a departure.

Then state it. 'Normality was assessed using the Shapiro-Wilk test; non-normally distributed variables were compared using the Mann-Whitney U test' is one sentence in your methods and it closes the question permanently.

Correlation: Pearson or Spearman

Pearson's correlation coefficient measures how closely two numerical variables follow a straight line, and it assumes both are roughly normal and the relationship is linear. Spearman's rank correlation works on ranks instead, and is the choice for ordinal data, skewed data, or a relationship that rises steadily without being straight.

Two cautions that come up in vivas. Correlation is not agreement: if you are comparing two methods of measuring the same quantity, the Bland-Altman approach answers the question you actually have and a high correlation coefficient does not. And a correlation coefficient establishes neither direction nor cause, however large it is.

Logistic regression, simple and multivariable

When the outcome is binary, logistic regression estimates the effect of a predictor on the odds of that outcome and reports it as an odds ratio with a confidence interval. With one predictor it is simple logistic regression. Add further predictors and it becomes multivariable logistic regression, which estimates the effect of each one while holding the others constant, and that is usually the version a hospital study needs, because a difference between two groups of patients is rarely attributable to a single variable. Both are well within reach of a postgraduate thesis.

Two practical limits. You need enough outcome events to support the number of predictors you include; a widely quoted rule of thumb is of the order of ten events per predictor, and it is a rule of thumb rather than a regulation. And the variables you adjust for should come from clinical reasoning written down in advance, not from feeding everything in and keeping whatever emerged significant. Report the odds ratio, its confidence interval and the reference category, because an odds ratio with no stated reference category cannot be read.

What a p value is, and what it is not

A p value is the probability of obtaining a result at least as extreme as the one you observed, if the null hypothesis were true. That is the whole of it. It is not the probability that the null hypothesis is true, not the probability that your finding is real, and not a measure of how large or how clinically important an effect is.

It follows that 0.04 and 0.06 are not different in kind, and that 0.05 is a convention rather than a boundary in nature. A very large study can return a tiny p value for a difference too small to change any management decision. A small study can miss a real and important difference and return a non-significant p. 'Not significant' means you did not demonstrate a difference, not that there is none.

Why confidence intervals belong in the results

A confidence interval reports the effect and its precision together: how big the difference was, and how much uncertainty surrounds it. That is the clinically useful statement, and it is the one a p value cannot make. A mean difference of four units with a 95% confidence interval from one to seven tells a reader something. A p value of 0.01 tells them only that something happened.

Give the estimate, its interval and the p value for every primary and secondary outcome, and give exact p values rather than 'p < 0.05' wherever your software prints them. Round sensibly: two decimals for most measurements, three for a p value, and never more digits than your instrument could measure.

Before you run anything

  • Name your tests in the protocol, not after you have seen the data
  • One primary objective, one outcome, one primary test; everything else is secondary and says so
  • Check your master chart codes and units before analysis rather than during it
  • Keep the output file your software produced, not only the numbers you copied out of it
  • If the design needs survival analysis, a multivariable model, or a sample size you cannot justify, involve a statistician at protocol stage

MDThesis puts a biostatistician on the analysis and on the analysis plan inside the protocol, which is the cheaper of the two places to involve one.


Keep reading

Elsewhere on the register

Keep reading


Ordered by how close the subject is to this one.

Free feasibility call

Tell us where your thesis stands. A senior doctor will tell you what to do next.


A senior doctor replies within one working day. No obligation. MD, MS, DNB, DrNB, DM, MCh, MDS and international programmes.

Request for a feasibility call

No obligation


No spam. A senior mentor replies personally. Your details stay private.


Reply within one working day · No obligation · Your details are not shared

Prefer to write to us first? Contact the practice. We mentor and edit; you remain the sole author of your thesis.

Document: Guide no. 08 — Choosing the right statistical test for your thesis data · Revision 1 · Last reviewed

Issued by MDThesis, a brand of REDENN Informatics Private Limited