Screening Tests: Sensitivity and Specificity — Exam Revision

Written & medically reviewed by the Kinase Medical Team · Last reviewed

Quick Answer

Sensitivity detects disease among diseased people; specificity excludes disease among non-diseased people. PPV is the probability of disease after a positive result, while NPV is the probability of no disease after a negative result. Draw a two-by-two table first. With fixed test characteristics, increasing prevalence raises PPV and lowers NPV.

What do screening tests and validity measure?

Screening identifies people who may have an unrecognised condition and need further assessment. A positive screen is not automatically a confirmed diagnosis. Test validity asks whether the test correctly distinguishes people with and without the target condition when compared with a suitable reference standard. Sensitivity and specificity describe that distinction; predictive values answer the question from the perspective of someone who has received a result.

The essential skill is choosing the correct denominator. Sensitivity begins with the people who truly have disease. Specificity begins with those who truly do not. Positive predictive value begins with positive test results, and negative predictive value begins with negative results. The numerator then asks how many people within that chosen group have been classified correctly. Build the table before using a formula.

A test result, a disease status and a clinical decision are separate ideas. The test may be positive, the reference standard may show disease is absent, and the sensible next step may still be confirmation. Keeping these ideas separate prevents the common assumption that every positive result must be a true positive. It also explains why screening programmes can generate many false alarms for uncommon conditions.

Sensitivity and Specificity Explained Clearly (Biostatistics)Clinical explanation of sensitivity, specificity and interpreting test results.Video: MedCram - Medical Lectures Explained CLEARLY · 12:15 · Watch on YouTube · Loads from YouTube (privacy-enhanced mode) only when you press play.

How should a two-by-two table be built?

Rows contain test results; columns contain actual disease status
Test resultDisease present by reference standardDisease absent by reference standard
Positivea = true positive, TPb = false positive, FP
Negativec = false negative, FNd = true negative, TN

Letters are only a convention. Some textbooks rotate the table, so memorising a letter without reading its label is risky. A true positive is positive by both the index test and the reference standard. A false positive is positive by the index test but disease-free by the reference. A false negative is missed disease, and a true negative is correctly excluded disease. Reconstruct these labels even if the table orientation changes.

Complete the margins before calculating. The diseased column totals TP plus FN. The non-diseased column totals FP plus TN. The positive row totals TP plus FP; the negative row totals FN plus TN. These totals make the four formulas almost self-explanatory. If a stem provides only sensitivity and the number with disease, first calculate the true positives, then obtain false negatives by subtraction.

The reference standard also deserves attention. A study with an imperfect reference, selective verification or a highly selected patient group may give misleading estimates. Examination tables usually treat the stated reference as the truth for calculation, but interpreting a published test study requires asking how that truth was established. Validity estimates are properties measured under particular study conditions.

How are sensitivity and specificity calculated?

Multiply proportions by one hundred when a percentage is requested
MeasureFormulaPlain-language question
SensitivityTP / (TP + FN)Among diseased people, how many test positive?
SpecificityTN / (TN + FP)Among non-diseased people, how many test negative?
False-negative rateFN / (TP + FN) = 1 − sensitivityAmong diseased people, how many are missed?
False-positive rateFP / (TN + FP) = 1 − specificityAmong non-diseased people, how many trigger a false alarm?

A sensitive test misses relatively few diseased people. A specific test produces relatively few false positives among people without disease. These statements describe different errors and do not imply that one test must always be better than another. A test can have high sensitivity but low specificity, or the reverse. Its usefulness depends on the consequences of missing disease and the consequences of an unnecessary follow-up.

SnNout links a highly sensitive test with the potential value of a negative result, while SpPin links a highly specific test with the potential value of a positive result. Treat both as prompts rather than mathematical guarantees. The actual change in disease probability also depends on the likelihood ratio and the pretest probability. A negative result does not override an extremely convincing clinical picture.

How do PPV and NPV differ from test validity?

Predictive values start from the result already known to the patient
MeasureFormulaInterpretation
Positive predictive valueTP / (TP + FP)Probability of disease among people with positive results
Negative predictive valueTN / (TN + FN)Probability of no disease among people with negative results

PPV answers the everyday question, “My test is positive: how likely is the disease?” Sensitivity answers a different question, “If I have disease, how likely is a positive test?” Reversing the direction of a conditional probability is a major source of error. A test can detect most diseased people while its positive results contain a substantial proportion of false positives.

NPV asks how often a negative result correctly represents absence of disease. It does not equal sensitivity, specificity or the fraction of all people who tested negative. A large number of true negatives can make NPV high in a low-prevalence population even if the test misses an appreciable fraction of the relatively few diseased people. Always return to the negative-result row.

Statistics with Crayons - Sensitivity, Specificity, Positive & Negative Predictive ValuesUniversity teaching using simple illustrations to distinguish validity measures from predictive values.Video: Penn Dental Medicine · 6:40 · Watch on YouTube · Loads from YouTube (privacy-enhanced mode) only when you press play.

A screening result should therefore be interpreted in its setting. The same assay used in an asymptomatic population and in a specialist referral clinic may have different predictive values because the populations contain different proportions of disease. This is why a reported PPV from one clinical service should not be carried unchanged into every other service.

How can all four measures be calculated from one example?

Consider a constructed practice example, not a measured population: one thousand people are tested. The reference standard identifies one hundred with disease. The test detects eighty of them and gives ninety false-positive results among the nine hundred without disease. Fill every cell before calculating. The remaining twenty diseased people are false negatives, and the remaining eight hundred and ten non-diseased people are true negatives.

Invented counts for practising the verified formulas
Test resultDisease presentDisease absentRow total
Positive8090170
Negative20810830
Column total1009001000
  • Sensitivity = 80 / 100 = 80%.
  • Specificity = 810 / 900 = 90%.
  • PPV = 80 / 170 ≈ 47.1%.
  • NPV = 810 / 830 ≈ 97.6%.
  • False-positive rate = 90 / 900 = 10%; false-negative rate = 20 / 100 = 20%.

The striking result is that fewer than half the positive results represent disease despite reasonably high sensitivity and specificity. There are many more non-diseased than diseased people, so even a modest false-positive rate creates numerous false positives. The test still detects most diseased people, which is exactly what sensitivity measures. It does not make every positive screen a diagnosis.

What changes when disease prevalence rises or falls?

Holding sensitivity and specificity constant, higher prevalence increases PPV and decreases NPV. Lower prevalence has the opposite effect. Prevalence changes the balance between diseased and non-diseased people available to generate true and false results. It does not directly change the conditional fractions used to define sensitivity and specificity in this simplified examination model.

Repeat the constructed example with a higher disease prevalence: among one thousand people, five hundred have disease. With sensitivity of eighty percent and specificity of ninety percent, there are four hundred true positives, one hundred false negatives, fifty false positives and four hundred and fifty true negatives. PPV is now four hundred divided by four hundred and fifty, approximately 88.9%; NPV is four hundred and fifty divided by five hundred and fifty, approximately 81.8%.

The assay characteristics were held fixed; the population changed. This is a direct application of the formulas, not a claim about a particular real-world screening programme. In practice, disease severity, case mix, threshold choice and measurement conditions can also change apparent sensitivity and specificity. State the fixed-test assumption when answering a theoretical prevalence question.

Overlapping healthy and diseased test-value distributions divided by a cut-off, with diagrams for sensitivity, specificity and predictive values.
Sensitivity and specificity start from disease status; predictive values start from the observed test result.Image: Original by Luigi Albert Maria, CC BY-SA 4.0

How do cut-offs and ROC curves affect performance?

For a marker where higher values suggest disease, lowering the positive threshold labels more people positive. This tends to increase sensitivity and decrease specificity. Raising the threshold tends to increase specificity and decrease sensitivity. State the direction of the marker first: a test where lower values indicate disease requires the corresponding reversal of the threshold logic.

A receiver operating characteristic, or ROC, curve plots sensitivity on the vertical axis against the false-positive rate, one minus specificity, on the horizontal axis. Each point represents a different threshold. A curve closer to the upper-left region has better discrimination across thresholds. The diagonal represents chance-level discrimination. ROC analysis compares discrimination; it does not supply PPV without considering the population.

ROC plot showing true-positive rate against false-positive rate, with curves approaching the upper-left corner and a diagonal random-classifier line.
Read the horizontal axis as one minus specificity, not specificity. Each position on a ROC curve corresponds to a threshold.Image: cmglee, MartinThoma, CC BY-SA 4.0

Choosing a cut-off requires a clinical purpose. Screening may prioritise avoiding missed disease, whereas a confirmatory use may prioritise reducing false positives. Neither priority should be interpreted as a universal rule that sensitivity alone determines the best screening test. A programme must consider the harms, availability of confirmation and whether identifying the condition can improve outcomes.

How do serial and parallel testing alter the result?

In serial testing, the final result is positive only when both tests are positive. This usually favours specificity at the expense of sensitivity: a diseased person missed by either test may be missed by the combined rule. In parallel testing, either positive test makes the final result positive. This usually favours sensitivity but accepts more false positives.

The terms describe decision rules, not simply the timing of blood draws. Two samples taken on the same day can still be interpreted with a serial positive rule. Conversely, tests done sequentially can use an either-positive rule. Read how the final classification is assigned. If an examination asks for combined percentages, multiplication formulas require the stated assumption of conditional independence; correlated tests need more careful treatment.

  • Disease known first: use sensitivity or specificity.
  • Result known first: use PPV or NPV.
  • Prevalence changes with fixed test characteristics: PPV and NPV change in opposite directions.
  • Higher disease-marker threshold: fewer positive labels, greater specificity and lower sensitivity.
  • Both positive required: serial rule; either positive sufficient: parallel rule.

Frequently asked questions

What is the easiest way to avoid formula errors?
Use labels before letters. Sensitivity uses all diseased people; specificity uses all non-diseased people. PPV uses all positive results, and NPV uses all negative results. Place the correctly classified people from that group in the numerator. This works even when the table has been rotated.
Does high sensitivity imply high PPV?
No. Sensitivity describes positive results among people already known to have disease. PPV describes disease among people who tested positive. False positives can outnumber true positives when disease is uncommon. The measures reverse the direction of the condition and should never be treated as interchangeable probabilities.
Which predictive value increases with prevalence?
PPV increases when prevalence rises, provided sensitivity and specificity are held constant. NPV decreases under the same assumption. More disease creates more true positives and fewer true negatives in a fixed population. The statement is a theoretical comparison of the same test characteristics across different prevalence settings.
Is a highly sensitive test always the best screening test?
Avoiding missed disease often makes sensitivity valuable for screening, but it is not the only requirement. Specificity, consequences of false results, access to confirmation and benefit from identifying disease also matter. A negative result from a highly sensitive test should still be interpreted against the pretest probability.
What does lowering a test cut-off do?
For a marker in which high values indicate disease, lowering the positive cut-off increases the number labelled positive. Sensitivity generally increases while specificity decreases. Raising the cut-off reverses that trade-off. If low values indicate disease, work through the direction of the result rather than applying the rule mechanically.
What are the standard ROC axes?
The vertical axis is sensitivity, also called the true-positive rate. The horizontal axis is one minus specificity, or the false-positive rate. Different points represent different thresholds. Curves nearer the upper-left corner show better discrimination. ROC performance alone does not give the predictive value in a particular population.
How do serial and parallel testing differ?
Serial testing requires both results to be positive for a final positive classification, tending to increase specificity. Parallel testing accepts either positive result, tending to increase sensitivity. Read the stated classification rule rather than relying on the order of testing. Numerical combination formulas also require suitable independence assumptions.

Sources

  1. StatPearls — Diagnostic Testing Accuracy
  2. PMC — Understanding and using sensitivity, specificity and predictive values
  3. PMC — Bayes’ Rule for Clinicians
  4. PMC — Statistics review 13: Receiver operating characteristic curves

For exam preparation and education only — not a substitute for clinical judgement or local guidelines. How we write and review these pages: editorial policy.

Revise Screening Tests: Sensitivity and Specificity with questions

Kinase: NEET-PG & INICET has previous-year papers, a subject-wise QBank and Grand Tests with explanations — on Android, iOS and the web.