What do screening tests and validity measure?
Screening identifies people who may have an unrecognised condition and need further assessment. A positive screen is not automatically a confirmed diagnosis. Test validity asks whether the test correctly distinguishes people with and without the target condition when compared with a suitable reference standard. Sensitivity and specificity describe that distinction; predictive values answer the question from the perspective of someone who has received a result.
The essential skill is choosing the correct denominator. Sensitivity begins with the people who truly have disease. Specificity begins with those who truly do not. Positive predictive value begins with positive test results, and negative predictive value begins with negative results. The numerator then asks how many people within that chosen group have been classified correctly. Build the table before using a formula.
A test result, a disease status and a clinical decision are separate ideas. The test may be positive, the reference standard may show disease is absent, and the sensible next step may still be confirmation. Keeping these ideas separate prevents the common assumption that every positive result must be a true positive. It also explains why screening programmes can generate many false alarms for uncommon conditions.
How should a two-by-two table be built?
| Test result | Disease present by reference standard | Disease absent by reference standard |
|---|---|---|
| Positive | a = true positive, TP | b = false positive, FP |
| Negative | c = false negative, FN | d = true negative, TN |
Letters are only a convention. Some textbooks rotate the table, so memorising a letter without reading its label is risky. A true positive is positive by both the index test and the reference standard. A false positive is positive by the index test but disease-free by the reference. A false negative is missed disease, and a true negative is correctly excluded disease. Reconstruct these labels even if the table orientation changes.
Complete the margins before calculating. The diseased column totals TP plus FN. The non-diseased column totals FP plus TN. The positive row totals TP plus FP; the negative row totals FN plus TN. These totals make the four formulas almost self-explanatory. If a stem provides only sensitivity and the number with disease, first calculate the true positives, then obtain false negatives by subtraction.
The reference standard also deserves attention. A study with an imperfect reference, selective verification or a highly selected patient group may give misleading estimates. Examination tables usually treat the stated reference as the truth for calculation, but interpreting a published test study requires asking how that truth was established. Validity estimates are properties measured under particular study conditions.
How are sensitivity and specificity calculated?
| Measure | Formula | Plain-language question |
|---|---|---|
| Sensitivity | TP / (TP + FN) | Among diseased people, how many test positive? |
| Specificity | TN / (TN + FP) | Among non-diseased people, how many test negative? |
| False-negative rate | FN / (TP + FN) = 1 − sensitivity | Among diseased people, how many are missed? |
| False-positive rate | FP / (TN + FP) = 1 − specificity | Among non-diseased people, how many trigger a false alarm? |
A sensitive test misses relatively few diseased people. A specific test produces relatively few false positives among people without disease. These statements describe different errors and do not imply that one test must always be better than another. A test can have high sensitivity but low specificity, or the reverse. Its usefulness depends on the consequences of missing disease and the consequences of an unnecessary follow-up.
SnNout links a highly sensitive test with the potential value of a negative result, while SpPin links a highly specific test with the potential value of a positive result. Treat both as prompts rather than mathematical guarantees. The actual change in disease probability also depends on the likelihood ratio and the pretest probability. A negative result does not override an extremely convincing clinical picture.
How do PPV and NPV differ from test validity?
| Measure | Formula | Interpretation |
|---|---|---|
| Positive predictive value | TP / (TP + FP) | Probability of disease among people with positive results |
| Negative predictive value | TN / (TN + FN) | Probability of no disease among people with negative results |
PPV answers the everyday question, “My test is positive: how likely is the disease?” Sensitivity answers a different question, “If I have disease, how likely is a positive test?” Reversing the direction of a conditional probability is a major source of error. A test can detect most diseased people while its positive results contain a substantial proportion of false positives.
NPV asks how often a negative result correctly represents absence of disease. It does not equal sensitivity, specificity or the fraction of all people who tested negative. A large number of true negatives can make NPV high in a low-prevalence population even if the test misses an appreciable fraction of the relatively few diseased people. Always return to the negative-result row.
A screening result should therefore be interpreted in its setting. The same assay used in an asymptomatic population and in a specialist referral clinic may have different predictive values because the populations contain different proportions of disease. This is why a reported PPV from one clinical service should not be carried unchanged into every other service.
How can all four measures be calculated from one example?
Consider a constructed practice example, not a measured population: one thousand people are tested. The reference standard identifies one hundred with disease. The test detects eighty of them and gives ninety false-positive results among the nine hundred without disease. Fill every cell before calculating. The remaining twenty diseased people are false negatives, and the remaining eight hundred and ten non-diseased people are true negatives.
| Test result | Disease present | Disease absent | Row total |
|---|---|---|---|
| Positive | 80 | 90 | 170 |
| Negative | 20 | 810 | 830 |
| Column total | 100 | 900 | 1000 |
- Sensitivity = 80 / 100 = 80%.
- Specificity = 810 / 900 = 90%.
- PPV = 80 / 170 ≈ 47.1%.
- NPV = 810 / 830 ≈ 97.6%.
- False-positive rate = 90 / 900 = 10%; false-negative rate = 20 / 100 = 20%.
The striking result is that fewer than half the positive results represent disease despite reasonably high sensitivity and specificity. There are many more non-diseased than diseased people, so even a modest false-positive rate creates numerous false positives. The test still detects most diseased people, which is exactly what sensitivity measures. It does not make every positive screen a diagnosis.
What changes when disease prevalence rises or falls?
Holding sensitivity and specificity constant, higher prevalence increases PPV and decreases NPV. Lower prevalence has the opposite effect. Prevalence changes the balance between diseased and non-diseased people available to generate true and false results. It does not directly change the conditional fractions used to define sensitivity and specificity in this simplified examination model.
Repeat the constructed example with a higher disease prevalence: among one thousand people, five hundred have disease. With sensitivity of eighty percent and specificity of ninety percent, there are four hundred true positives, one hundred false negatives, fifty false positives and four hundred and fifty true negatives. PPV is now four hundred divided by four hundred and fifty, approximately 88.9%; NPV is four hundred and fifty divided by five hundred and fifty, approximately 81.8%.
The assay characteristics were held fixed; the population changed. This is a direct application of the formulas, not a claim about a particular real-world screening programme. In practice, disease severity, case mix, threshold choice and measurement conditions can also change apparent sensitivity and specificity. State the fixed-test assumption when answering a theoretical prevalence question.

How do cut-offs and ROC curves affect performance?
For a marker where higher values suggest disease, lowering the positive threshold labels more people positive. This tends to increase sensitivity and decrease specificity. Raising the threshold tends to increase specificity and decrease sensitivity. State the direction of the marker first: a test where lower values indicate disease requires the corresponding reversal of the threshold logic.
A receiver operating characteristic, or ROC, curve plots sensitivity on the vertical axis against the false-positive rate, one minus specificity, on the horizontal axis. Each point represents a different threshold. A curve closer to the upper-left region has better discrimination across thresholds. The diagonal represents chance-level discrimination. ROC analysis compares discrimination; it does not supply PPV without considering the population.

Choosing a cut-off requires a clinical purpose. Screening may prioritise avoiding missed disease, whereas a confirmatory use may prioritise reducing false positives. Neither priority should be interpreted as a universal rule that sensitivity alone determines the best screening test. A programme must consider the harms, availability of confirmation and whether identifying the condition can improve outcomes.
How do serial and parallel testing alter the result?
In serial testing, the final result is positive only when both tests are positive. This usually favours specificity at the expense of sensitivity: a diseased person missed by either test may be missed by the combined rule. In parallel testing, either positive test makes the final result positive. This usually favours sensitivity but accepts more false positives.
The terms describe decision rules, not simply the timing of blood draws. Two samples taken on the same day can still be interpreted with a serial positive rule. Conversely, tests done sequentially can use an either-positive rule. Read how the final classification is assigned. If an examination asks for combined percentages, multiplication formulas require the stated assumption of conditional independence; correlated tests need more careful treatment.
- Disease known first: use sensitivity or specificity.
- Result known first: use PPV or NPV.
- Prevalence changes with fixed test characteristics: PPV and NPV change in opposite directions.
- Higher disease-marker threshold: fewer positive labels, greater specificity and lower sensitivity.
- Both positive required: serial rule; either positive sufficient: parallel rule.