How are epidemiological study designs classified?
Every study tests a link between an exposure (a risk factor, drug or behaviour) and an outcome (usually a disease). The first split is whether the investigator assigns the exposure. If the investigator assigns it, the study is experimental; if the investigator only observes what people already do, it is observational.
| Group | Design | Main purpose |
|---|---|---|
| Observational — descriptive | Case report, case series, cross-sectional (prevalence) survey | Describes disease by time, place and person; generates hypotheses |
| Observational — descriptive/ecological | Ecological study | Compares group-level data between populations |
| Observational — analytic | Case-control study | Tests hypotheses: disease → past exposure |
| Observational — analytic | Cohort study (prospective or retrospective) | Tests hypotheses: exposure → future disease |
| Experimental | Randomised controlled trial (RCT), field and community trials | Proves the effect of an intervention |

What do cross-sectional and ecological studies tell you?
A cross-sectional study is a snapshot: exposure and outcome are measured at a single point in time, usually by survey. It has no follow-up, so it is simple and cheap and is the standard way to measure prevalence. Because exposure and outcome are collected together, it cannot establish cause and effect and is regarded as the weakest observational design for causation.
An ecological study uses aggregate (group-level) data — for example national salt intake versus national stroke rates. Its results apply only at the population level. Assuming that a group-level association also holds for individuals is the ecological fallacy, a form of confounding unique to this design.
A newer hybrid, the case-crossover study, suits transient triggers (for example heavy exertion before a heart attack): each case serves as its own control, comparing the exposure just before the event with an earlier control period.
How do case-control, cohort and randomised controlled trials compare?
| Feature | Case-control | Cohort | RCT |
|---|---|---|---|
| Starting point | Disease present (cases) vs absent (controls) | Exposed vs unexposed, all disease-free | Eligible people randomly allocated |
| Direction | Backward (retrospective) | Forward (prospective) or historical (retrospective cohort) | Forward |
| Measure of association | Odds ratio | Relative risk (incidence measured directly), attributable risk | Relative risk, risk difference |
| Best for | Rare diseases, many exposures, outbreaks | Rare exposures, many outcomes of one exposure | Testing interventions; proving causation |
| Main bias | Recall bias, control selection, confounding | Selection bias, loss to follow-up | Breaks in randomisation, refusals, drop-outs |
| Cost and time | Cheap, quick | Expensive, long for rare/slow outcomes | Most expensive |
| Incidence? | Cannot be calculated | Calculated directly | Calculated directly |
Case-control studies let you study a rare disease without following thousands of people for years: you simply collect existing cases and comparable controls. You can use more than one control per case (2:1 or 4:1) to increase power. They establish association, not causation.
The hardest part of a case-control study is choosing controls. Ideally cases and controls share age, sex, general health, history and environment, so the only systematic difference is the disease. If cases come from across the country but controls from one small community, any exposure tied to place (sunburn, diet, water) will look falsely linked to the disease. Unmeasured variables related to both exposure and outcome create confounding, which matching and stratified analysis try to control.
Cohort studies classify people by exposure first, so incidence can be calculated directly in exposed and unexposed groups — that is why relative risk is their measure of effect. Recall bias is very low and several outcomes can be studied at once, but they are more prone to selection bias and become very costly when outcomes are rare or slow.
In an RCT the researcher randomly assigns subjects to an experimental group and a control group (no treatment, placebo or standard care). Randomisation avoids confounding and minimises selection bias, so the groups differ only by the intervention. That is why the RCT is the gold standard design.

How are odds ratio, relative risk and attributable risk calculated?
Set the data in a 2 × 2 table: rows = exposed / unexposed; columns = disease / no disease. Cell a = exposed with disease, b = exposed without disease, c = unexposed with disease, d = unexposed without disease. For a broader primer on rates, tests and errors, see biostatistics high-yield.
Odds ratio = (a/b) ÷ (c/d) = ad / bc
Case-control studies. OR > 1: exposure linked to more disease; OR < 1: protective; OR = 1: no association.
Relative risk = [a / (a + b)] ÷ [c / (c + d)]
Cohort studies and trials: incidence in exposed ÷ incidence in unexposed.
Attributable risk (risk difference) = incidence in exposed − incidence in unexposed
Absolute excess risk due to the exposure.
Attributable risk per cent (attributable proportion) = (risk in exposed − risk in unexposed) ÷ risk in exposed × 100
Share of disease in the exposed group that is due to the exposure; assumes a single causal factor.
Vaccine efficacy = (risk in unvaccinated − risk in vaccinated) ÷ risk in unvaccinated = 1 − RR
Efficacy = ideal conditions (trial); effectiveness = field conditions.
| Lung cancer | No lung cancer | |
|---|---|---|
| Smokers | 17 (a) | 83 (b) |
| Non-smokers | 1 (c) | 99 (d) |
Here RR = (17/100) ÷ (1/100) = 17, while OR = (17/83) ÷ (1/99) ≈ 20.5. The two measures diverge because the outcome is common in the exposed group; when a disease is rare, the odds ratio approximates the relative risk. The 95% confidence interval for this OR (about 2.7 to 158) excludes 1, so the association is statistically significant; a CI that includes 1 means no significant association.
What are the main types of bias in epidemiological studies?
| Bias | What happens | Typical design / fix |
|---|---|---|
| Selection bias | Study population does not represent the target population | Cohort studies, hospital-based studies; proper sampling and controls |
| Berkson (admission) bias | Hospital cases compared with non-hospital controls; hospital patients are sicker and unrepresentative | Hospital-based case-control studies; choose appropriate controls |
| Recall bias | Cases remember and report past exposures more than controls | Case-control studies; shorten exposure-outcome interval, use records |
| Observer bias | Assessor's knowledge of group alters outcome recording | Trials; blinding of investigators |
| Hawthorne effect | Subjects change behaviour because they know they are observed | Any study; reduce or hide observation |
| Lead-time bias | Earlier detection makes survival look longer without changing death | Screening studies; compare mortality rates instead of survival |
| Length-time bias | Screening picks up slow, indolent cases more often, overestimating survival | Screening studies |
| Confounding | A third variable linked to both exposure and outcome distorts the association | Randomisation, matching, stratified/multivariable analysis |
| Publication bias | Positive results more likely to be published | Trial registration and archiving of results |
What is blinding and why does it matter in trials?
Blinding (masking) means concealing group allocation from people involved in the trial — participants, data collectors, intervention providers and even data analysers. It reduces observer bias (differential assessment of subjective outcomes) and placebo-driven behaviour changes in participants.
| Level | Who does not know the allocation |
|---|---|
| Single-blind | Usually the participant |
| Double-blind | Participant and the investigator/outcome assessor |
| Triple-blind | Participant, investigator and the data analyst |
| Open-label | Nobody is blinded — e.g. surgery versus medical therapy where the scar shows |
Randomisation and blinding are not the same. Randomisation decides who gets what and protects against confounding and selection bias at the start. Blinding protects the measurement of outcomes during and after the trial. A trial can be randomised but open-label when the intervention cannot be hidden.
How do you pick the right design in a question?
- Is an intervention being assigned by the investigator? Yes → experimental (RCT, field trial, community trial). No → observational.
- Is everything measured at one time point? Yes → cross-sectional (gives prevalence).
- Is the unit a population, not an individual? Yes → ecological.
- Does the study start with people who already have the disease? Yes → case-control (odds ratio).
- Does it start with exposed and unexposed disease-free people? Yes → cohort (relative risk, attributable risk); retrospective if both exposure and outcome are drawn from past records.
Study design links directly to prevention and disease natural history — see levels of prevention, spectrum of disease and health indicators.