How are sensitivity and specificity calculated from a 2×2 table?
Every screening-test question starts from the same 2×2 table: test result (positive or negative) against true disease status (from the gold standard). Label the cells a, b, c, d and every formula follows.
| Disease present | Disease absent | Total | |
|---|---|---|---|
| Test positive | a — true positive (TP) | b — false positive (FP) | a + b |
| Test negative | c — false negative (FN) | d — true negative (TN) | c + d |
| Total | a + c (all diseased) | b + d (all healthy) | N |
| Measure | Formula | Meaning |
|---|---|---|
| Sensitivity | a / (a + c) | Proportion of diseased people who test positive |
| Specificity | d / (b + d) | Proportion of healthy people who test negative |
| Positive predictive value (PPV) | a / (a + b) | Of all positives, how many truly have the disease |
| Negative predictive value (NPV) | d / (c + d) | Of all negatives, how many are truly disease-free |
| Positive likelihood ratio | Sensitivity / (1 − specificity) | True-positive rate divided by false-positive rate |
| Negative likelihood ratio | (1 − sensitivity) / specificity | False-negative rate divided by true-negative rate |
How does prevalence change PPV and NPV — a worked example
Suppose 1000 people are screened, 100 have the disease (prevalence 10%). The test picks up 90 of the 100 diseased and wrongly flags 45 of the 900 healthy.
| Diseased | Healthy | Total | |
|---|---|---|---|
| Test + | 90 (TP) | 45 (FP) | 135 |
| Test − | 10 (FN) | 855 (TN) | 865 |
| Total | 100 | 900 | 1000 |
- Sensitivity = 90/100 = 90%; specificity = 855/900 = 95%.
- PPV = 90/135 = 66.7% — one in three positives is a false alarm.
- NPV = 855/865 = 98.8%.
- LR+ = 0.90/0.05 = 18; LR− = 0.10/0.95 ≈ 0.11.

What are mean, median and mode — and when is each used?
| Measure | Definition | Key property |
|---|---|---|
| Mean | Sum of observations ÷ number of observations | Uses every value, but is pulled by extreme values (outliers) |
| Median | Middle observation when data are arranged in order | Not affected by outliers; preferred for skewed data |
| Mode | Value that occurs most often (maximum frequency) | Simple to read off a frequency table; equals mean and median in a symmetric distribution |
If the mean, median and mode coincide, the distribution is symmetric (skewness = 0). When a few very high or very low values pull the mean away, the median gives a fairer 'typical' value — which is why non-parametric tests, which work on medians or ranks, are used for skewed data.

What are standard deviation, standard error and coefficient of variation?
Measures of dispersion describe how spread out the data are: range, interquartile range, variance, standard deviation (SD), standard error (SE) and coefficient of variation (CV).
| Measure | What it shows | Formula / note |
|---|---|---|
| Standard deviation (SD) | How far values spread around the mean | Square root of the variance |
| Standard error (SE) | How far a sample mean is likely to be from the population mean (SD of many sample means) | SE = SD ÷ √n — falls as sample size rises |
| Coefficient of variation (CV) | SD relative to the mean; compares variability of different measurements | CV = 100 × SD ÷ mean |
| Interquartile range (IQR) | Spread of the middle 50% | Q3 (75th percentile) − Q1 (25th percentile) |
What does the normal curve show — the 68–95–99.7 rule?
The normal (Gaussian) distribution is a bell-shaped, symmetric curve fully described by its mean and SD; mean, median and mode are equal at its centre.
| Range | Observations included |
|---|---|
| Mean ± 1 SD | 68.2% |
| Mean ± 2 SD | 95.4% |
| Mean ± 3 SD | 99.7% |
| Mean ± 1.96 SD | 95% (the Z value used for 5% significance) |
| Mean ± 2.58 SD | 99% (Z value for 1% significance) |

- Testing normality: Shapiro–Wilk test (better for small samples, n < 50) and Kolmogorov–Smirnov test (n ≥ 50); skewness and kurtosis between −1 and +1 suggest approximate normality.
- Quick check: if the SD is less than half the mean (CV < 50%), data are often taken as roughly normal.
- Why it matters: normally distributed data are compared with parametric tests; otherwise non-parametric tests are used.
Which test of significance should be used — t-test, chi-square or ANOVA?
Choosing a test depends on three things: the aim of the study, the type and distribution of the data, and whether observations are paired (same subjects measured twice) or unpaired (different subjects in each group). Tests that compare means are parametric; tests that compare medians, ranks or proportions are non-parametric.
| Question | Normal data (parametric) | Non-normal / ordinal (non-parametric) |
|---|---|---|
| Compare means of 2 independent groups | Unpaired (independent samples) t-test | Mann–Whitney U test |
| Compare 2 paired measurements (before–after) | Paired t-test | Wilcoxon test |
| Compare means of 3 or more groups | One-way ANOVA (F test) | Kruskal–Wallis H test |
| Repeated measures, 3 or more times | Repeated-measures ANOVA | Friedman test |
| Correlation between 2 variables | Pearson correlation | Spearman rank correlation |
| Situation | Test |
|---|---|
| Association between 2 categorical variables, independent groups | Pearson chi-square test or Fisher exact test |
| Change in proportions, 2 paired groups | McNemar test |
| Change in proportions, 3 or more paired groups | Cochran Q test |
What does a p value mean, and how does it relate to confidence intervals?
The null hypothesis states that there is no difference between groups. It is assumed true until the data give enough evidence to reject it. Before the study the researcher fixes the significance level, alpha — usually 0.05, a 5% chance of being wrong when claiming a difference.
- A p value is the probability, under the statistical model, of getting a result equal to or more extreme than the one observed.
- If p < alpha (for example p = 0.02 with alpha 0.05), the null hypothesis is rejected and the result is called statistically significant.
- The p value is not the probability that the null hypothesis is true.
- Confidence intervals are reported with or instead of p values; their width depends on the standard error and sample size — a smaller sample gives a wider, less precise interval.
- Statistical vs clinical significance: a tiny difference can be statistically significant in a huge study yet not matter to the patient; clinical significance asks whether the size of the effect is important.
What are type I and type II errors, power and sample size?
| Null hypothesis actually true | Null hypothesis actually false | |
|---|---|---|
| Null rejected (difference claimed) | Type I error (alpha) — false positive | Correct decision (probability = power, 1 − β) |
| Null not rejected (no difference claimed) | Correct decision | Type II error (beta) — false negative |
- Type I error: rejecting the null hypothesis and claiming a difference that does not exist. Its probability is set in advance as alpha (the significance level).
- Type II error: declaring no difference when one really exists; its probability is beta.
- Power = 1 − β: the probability of correctly rejecting a false null hypothesis. 80% is the usual minimum target (Z = 0.84 at 80% power, 1.28 at 90%).
- Power depends on the significance level, sample size and effect size; a small study with low power is prone to type II error.
- Sample size must be larger when the SD is larger, when the outcome is categorical rather than continuous, and when cluster or other complex sampling adds a design effect.