High-Yield Biostatistics — Central Tendency, SD, Normal Curve, 2×2 Table, Tests of Significance, P Value and Errors

Written & medically reviewed by the Kinase Medical Team · Last reviewed

Quick Answer

Sensitivity is true positives among all diseased and specificity is true negatives among all healthy; predictive values change with prevalence. In a normal curve, mean ± 1, 2 and 3 SD cover about 68%, 95% and 99.7%. Compare two means with a t-test, three or more with ANOVA, proportions with chi-square; a type I error is a false positive.

How are sensitivity and specificity calculated from a 2×2 table?

Every screening-test question starts from the same 2×2 table: test result (positive or negative) against true disease status (from the gold standard). Label the cells a, b, c, d and every formula follows.

The 2×2 table
Disease presentDisease absentTotal
Test positivea — true positive (TP)b — false positive (FP)a + b
Test negativec — false negative (FN)d — true negative (TN)c + d
Totala + c (all diseased)b + d (all healthy)N
Test performance formulas
MeasureFormulaMeaning
Sensitivitya / (a + c)Proportion of diseased people who test positive
Specificityd / (b + d)Proportion of healthy people who test negative
Positive predictive value (PPV)a / (a + b)Of all positives, how many truly have the disease
Negative predictive value (NPV)d / (c + d)Of all negatives, how many are truly disease-free
Positive likelihood ratioSensitivity / (1 − specificity)True-positive rate divided by false-positive rate
Negative likelihood ratio(1 − sensitivity) / specificityFalse-negative rate divided by true-negative rate
Sensitivity and Specificity Explained Clearly (Biostatistics)Dr Roger Seheult works through sensitivity, specificity and the 2×2 table with clinical examples.Video: MedCram - Medical Lectures Explained CLEARLY · 12:15 · Watch on YouTube · Loads from YouTube (privacy-enhanced mode) only when you press play.

How does prevalence change PPV and NPV — a worked example

Suppose 1000 people are screened, 100 have the disease (prevalence 10%). The test picks up 90 of the 100 diseased and wrongly flags 45 of the 900 healthy.

Worked example (illustrative numbers)
DiseasedHealthyTotal
Test +90 (TP)45 (FP)135
Test −10 (FN)855 (TN)865
Total1009001000
  • Sensitivity = 90/100 = 90%; specificity = 855/900 = 95%.
  • PPV = 90/135 = 66.7% — one in three positives is a false alarm.
  • NPV = 855/865 = 98.8%.
  • LR+ = 0.90/0.05 = 18; LR− = 0.10/0.95 ≈ 0.11.
Diagram with relevant elements (sick people) on the left and others on the right; a circle of selected elements contains true positives on the left and false positives on the right, with false negatives and true negatives outside; small pie-style fractions below show sensitivity and specificity.
Sensitivity is the share of all truly sick people the test selects; specificity is the share of all healthy people it correctly leaves out.Image: Walber, FeanDoe, CC BY-SA 4.0

What are mean, median and mode — and when is each used?

Measures of central tendency
MeasureDefinitionKey property
MeanSum of observations ÷ number of observationsUses every value, but is pulled by extreme values (outliers)
MedianMiddle observation when data are arranged in orderNot affected by outliers; preferred for skewed data
ModeValue that occurs most often (maximum frequency)Simple to read off a frequency table; equals mean and median in a symmetric distribution

If the mean, median and mode coincide, the distribution is symmetric (skewness = 0). When a few very high or very low values pull the mean away, the median gives a fairer 'typical' value — which is why non-parametric tests, which work on medians or ranks, are used for skewed data.

Two curves side by side: on the left a negatively skewed curve with its long tail to the left, on the right a positively skewed curve with its long tail to the right, each drawn over a dashed symmetric curve.
Skewed distributions have a long tail on one side; in such data the median describes the centre better than the mean.Image: Rodolfo Hermans (Godot) at en.wikipedia, CC BY-SA 3.0

What are standard deviation, standard error and coefficient of variation?

Measures of dispersion describe how spread out the data are: range, interquartile range, variance, standard deviation (SD), standard error (SE) and coefficient of variation (CV).

Measures of dispersion
MeasureWhat it showsFormula / note
Standard deviation (SD)How far values spread around the meanSquare root of the variance
Standard error (SE)How far a sample mean is likely to be from the population mean (SD of many sample means)SE = SD ÷ √n — falls as sample size rises
Coefficient of variation (CV)SD relative to the mean; compares variability of different measurementsCV = 100 × SD ÷ mean
Interquartile range (IQR)Spread of the middle 50%Q3 (75th percentile) − Q1 (25th percentile)

What does the normal curve show — the 68–95–99.7 rule?

The normal (Gaussian) distribution is a bell-shaped, symmetric curve fully described by its mean and SD; mean, median and mode are equal at its centre.

Area under the normal curve
RangeObservations included
Mean ± 1 SD68.2%
Mean ± 2 SD95.4%
Mean ± 3 SD99.7%
Mean ± 1.96 SD95% (the Z value used for 5% significance)
Mean ± 2.58 SD99% (Z value for 1% significance)
Bell-shaped normal distribution curve divided into bands one standard deviation wide, labelled 34.1% on each side of the mean, then 13.6%, 2.1% and 0.1% moving outwards to plus and minus three sigma.
About 68% of values lie within 1 SD of the mean, about 95% within 2 SD and about 99.7% within 3 SD.Image: M. W. Toews, CC BY 2.5
  • Testing normality: Shapiro–Wilk test (better for small samples, n < 50) and Kolmogorov–Smirnov test (n ≥ 50); skewness and kurtosis between −1 and +1 suggest approximate normality.
  • Quick check: if the SD is less than half the mean (CV < 50%), data are often taken as roughly normal.
  • Why it matters: normally distributed data are compared with parametric tests; otherwise non-parametric tests are used.

Which test of significance should be used — t-test, chi-square or ANOVA?

Choosing a test depends on three things: the aim of the study, the type and distribution of the data, and whether observations are paired (same subjects measured twice) or unpaired (different subjects in each group). Tests that compare means are parametric; tests that compare medians, ranks or proportions are non-parametric.

Choosing a test (from Mishra et al., Ann Card Anaesth 2019)
QuestionNormal data (parametric)Non-normal / ordinal (non-parametric)
Compare means of 2 independent groupsUnpaired (independent samples) t-testMann–Whitney U test
Compare 2 paired measurements (before–after)Paired t-testWilcoxon test
Compare means of 3 or more groupsOne-way ANOVA (F test)Kruskal–Wallis H test
Repeated measures, 3 or more timesRepeated-measures ANOVAFriedman test
Correlation between 2 variablesPearson correlationSpearman rank correlation
Comparing proportions (categorical data)
SituationTest
Association between 2 categorical variables, independent groupsPearson chi-square test or Fisher exact test
Change in proportions, 2 paired groupsMcNemar test
Change in proportions, 3 or more paired groupsCochran Q test

What does a p value mean, and how does it relate to confidence intervals?

The null hypothesis states that there is no difference between groups. It is assumed true until the data give enough evidence to reject it. Before the study the researcher fixes the significance level, alpha — usually 0.05, a 5% chance of being wrong when claiming a difference.

  • A p value is the probability, under the statistical model, of getting a result equal to or more extreme than the one observed.
  • If p < alpha (for example p = 0.02 with alpha 0.05), the null hypothesis is rejected and the result is called statistically significant.
  • The p value is not the probability that the null hypothesis is true.
  • Confidence intervals are reported with or instead of p values; their width depends on the standard error and sample size — a smaller sample gives a wider, less precise interval.
  • Statistical vs clinical significance: a tiny difference can be statistically significant in a huge study yet not matter to the patient; clinical significance asks whether the size of the effect is important.

What are type I and type II errors, power and sample size?

Errors in hypothesis testing
Null hypothesis actually trueNull hypothesis actually false
Null rejected (difference claimed)Type I error (alpha) — false positiveCorrect decision (probability = power, 1 − β)
Null not rejected (no difference claimed)Correct decisionType II error (beta) — false negative
  • Type I error: rejecting the null hypothesis and claiming a difference that does not exist. Its probability is set in advance as alpha (the significance level).
  • Type II error: declaring no difference when one really exists; its probability is beta.
  • Power = 1 − β: the probability of correctly rejecting a false null hypothesis. 80% is the usual minimum target (Z = 0.84 at 80% power, 1.28 at 90%).
  • Power depends on the significance level, sample size and effect size; a small study with low power is prone to type II error.
  • Sample size must be larger when the SD is larger, when the outcome is categorical rather than continuous, and when cluster or other complex sampling adds a design effect.
Medical Statistics - Part 10: Type 1 and Type 2 ErrorsAMBOSS explainer on type I and type II errors, significance level and statistical power.Video: AMBOSS: Medical Knowledge Distilled · 8:25 · Watch on YouTube · Loads from YouTube (privacy-enhanced mode) only when you press play.

Frequently asked questions

What is the difference between sensitivity and positive predictive value?
Sensitivity is the proportion of people with the disease who test positive, a divided by a plus c in the 2×2 table. Positive predictive value is the proportion of people who test positive who really have the disease, a divided by a plus b. Sensitivity does not change with prevalence, but PPV rises when the disease is more common.
Which measures of a test are affected by prevalence?
Positive and negative predictive values depend on prevalence. When a disease is common, the test is better at ruling it in and worse at ruling it out, so PPV rises and NPV falls. Sensitivity, specificity and likelihood ratios are properties of the test itself and are not affected by how common the disease is.
What percentage of values lie within 2 SD of the mean?
In a normal distribution about 95.4 percent of observations lie within the mean plus or minus 2 standard deviations. About 68.2 percent lie within 1 SD and 99.7 percent within 3 SD. Exactly 95 percent lie within 1.96 SD, which is why 1.96 is the Z value used for a 5 percent significance level.
Which test compares means of three or more groups?
One-way analysis of variance, or ANOVA (the F test), compares means of three or more independent groups when the data are normally distributed. Repeated-measures ANOVA is used when the same subjects are measured several times. If the data are skewed, the Kruskal–Wallis H test or, for repeated measures, the Friedman test is used instead.
When is the chi-square test used?
The Pearson chi-square test, or the Fisher exact test, is used to test the association between two categorical variables in independent groups, that is, to compare proportions. For proportions measured twice in the same subjects, such as before and after an intervention, the McNemar test is used, and the Cochran Q test for three or more paired groups.
What is a type II error and how can it be reduced?
A type II error is concluding that there is no difference when a real difference exists, a false negative. Its probability is beta, and power equals one minus beta. Power rises with a larger sample size, a larger effect size and a higher significance level, so adequately sized studies with about 80 percent power reduce type II errors.
What is the difference between standard deviation and standard error?
Standard deviation measures how widely individual values spread around the mean. Standard error is the standard deviation of sample means and shows how precisely the sample mean estimates the population mean. It is calculated as SD divided by the square root of the sample size, so it shrinks as the sample becomes larger.

Sources

  1. StatPearls — Diagnostic Testing Accuracy: Sensitivity, Specificity, Predictive Values and Likelihood Ratios (NCBI Bookshelf)
  2. StatPearls — Type I and Type II Errors and Statistical Power (NCBI Bookshelf)
  3. StatPearls — Statistical Significance (NCBI Bookshelf)
  4. StatPearls — Hypothesis Testing, P Values, Confidence Intervals, and Significance (NCBI Bookshelf)
  5. Mishra P et al. Descriptive Statistics and Normality Tests for Statistical Data. Ann Card Anaesth 2019 (PMC)
  6. Mishra P et al. Selection of Appropriate Statistical Methods for Data Analysis. Ann Card Anaesth 2019 (PMC)
  7. Suresh KP, Chandrashekara S. Sample size estimation and power analysis for clinical research studies. J Hum Reprod Sci 2012 (PMC)

For exam preparation and education only — not a substitute for clinical judgement or local guidelines. How we write and review these pages: editorial policy.

Revise Biostatistics with questions

Kinase: NEET-PG & INICET has previous-year papers, a subject-wise QBank and Grand Tests with explanations — on Android, iOS and the web.