Knowledge IVD Development How to Verify IVD Reference Intervals with a Small Sample Size? 20-Sample Verification Guide
Author avatar

Tech Team · CamelBio

Updated 1 month ago

How to Verify IVD Reference Intervals with a Small Sample Size? 20-Sample Verification Guide


The most pragmatic way to test an IVD manufacturer’s reference interval locally is with a small-sample verification. By measuring just 20 specimens from healthy individuals in your own population, you can confirm whether the manufacturer’s proposed limits are fit for your instrument and your patients. The rule is straightforward: if no more than 2 out of those 20 results fall outside the published upper or lower limits, the reference interval is considered transferable.

The 20-sample protocol gives you a rapid, statistically grounded yes-or-no answer. However, it cannot detect a situation where your true local range is narrower than the manufacturer’s range. To build real confidence, you must supplement the one-time check with routine data mining of your outpatient population and ensure that the donor population and your own are demographically comparable.

The 20-Sample Protocol: A Practical Verification Method

This method, endorsed by CLSI guidelines, balances statistical power with the real-world constraints of a clinical laboratory. It answers the question: "Is the manufacturer’s reference interval grossly inappropriate for my setup?"

Step 1: Selecting Your Local Reference Sample

Collect specimens from 20 healthy reference individuals. These individuals must be representative of your target patient population—matching in age, sex, ethnicity, and environmental exposures where relevant. The selection process is critical; if your local population differs substantially from the manufacturer’s study group, even a perfect 20-sample test may be meaningless.

Ensure preanalytical conditions (patient preparation, specimen collection, handling, and storage) are identical to those specified in the IVD kit’s instructions. Any deviation can introduce bias that clouds the verification.

Step 2: Running the Test and Applying the 10% Rule

Measure each sample on your instrument using the same reagent lot and calibration as for routine patient testing. Apply the “2 out of 20” criterion: count how many results fall outside the manufacturer’s reference limits. If the number is two or fewer, the transferred interval is verified.

A result of three or more outliers signals a problem—the interval cannot be accepted for local use. This usually triggers an investigation into potential causes: population mismatch, analytical bias, or preanalytical issues. In such cases, a full de novo reference interval study is often required.

Why 20? The Statistical Rationale

Twenty is the smallest sample size that reliably detects a large population shift. Statistically, if 95% of the parent population truly falls within the reference limits, the probability of observing 0, 1, or 2 individuals outside the limits in a random sample of 20 is about 92%. So you have a high chance of “passing” a truly appropriate interval.

Conversely, if the actual proportion of healthy people outside the limits is 15% or more (a meaningful mismatch), the chance of seeing 3 or more outliers in 20 samples exceeds 50%. The test is therefore reasonably sensitive to gross incompatibility, but not to subtle differences.

Beyond the 20-Sample Check: Bolstering Confidence

Passing the 20-sample hurdle is encouraging but insufficient on its own. The deep need is to know that the reference interval will perform safely for every routine result, not just for the first 20 you test.

The Hidden Limitation: Narrower Local Ranges

A successful 20-sample verification does not guarantee that your local population’s healthy range is as wide as the manufacturer’s. If your population is biologically more homogeneous, the true central 95% could be narrower, meaning the quoted limits might allow abnormal patients to appear “normal.” The 20-sample method lacks the statistical resolution to detect this kind of tight-range problem.

Data Mining: The Continuous Audit Tool

Leverage your existing outpatient data as a cost-free ongoing monitor. Calculate the median of a stable, healthy outpatient population (e.g., routine bicarbonate or electrolyte results) and track it over time. Compare that median to the midpoint of the transferred reference interval.

If the median begins to drift or if more than 2.5% of outpatient values consistently fall below or above the limits, your reference interval may no longer be appropriate. This kind of data mining can flag analytical shifts or population changes long before a formal re-verification is triggered. When doubt arises, a fresh 20-sample study can quickly confirm or refute the need for adjustment.

Verifying Demographic and Preanalytical Consistency

Even before you run the 20-sample protocol, you must critically compare the manufacturer’s reference population to your own. Look for differences in age distribution, biological sex ratio, ethnicity, and lifestyle/environmental factors. If the populations are not comparable, no amount of local verification with 20 samples can make the transferred interval valid.

Similarly, confirm that your analytical method and preanalytical workflow exactly mirror those used during the manufacturer’s reference interval study. Any difference—a different platform generation, a modified onboard dilution, a change in sample tube additive—can create a systematic bias that invalidates the transfer.

Understanding the Trade-offs

The 20-sample protocol is a screening tool, not a thorough validation. Its power lies in speed and simplicity, but it carries blind spots:

  • It cannot detect a local range that is narrower than the manufacturer’s.
  • It assumes that your 20 individuals are perfectly representative, which is never fully true.
  • A “pass” does not guarantee long-term stability; it only confirms the situation at a single point in time.

Data mining is powerful but reactive. It depends on the volume and integrity of your historical test results, and it requires statistical discipline to set appropriate alert thresholds. Misinterpreting normal seasonal variation as a shift can lead to unnecessary investigations.

Population and method comparability are non‑negotiable. If you lack demographic data on the manufacturer’s reference population, you must err on the side of caution. A de novo reference interval study—though resource-intensive—is the only safe route when comparability is in serious doubt.

Making the Right Choice for Your Laboratory

Every clinical laboratory walks a line between regulatory compliance and operational practicality. The best path depends on your specific goal.

  • If your primary focus is rapid implementation of a standard test in a well‑matched adult population: Use the 20‑sample verification as your first‑line tool. Then, immediately institute an outpatient median tracker to catch any future drift, so you never rely on a single, static check.
  • If your population or workflow differs significantly from the manufacturer’s baseline (e.g., pediatric, distinct ethnicity, or unique preanalytical conditions): Do not rely on the 20‑sample protocol. Plan for a full de novo reference interval study, or work with a reference laboratory to establish appropriate limits.
  • If you want cost‑effective ongoing assurance without re‑verification studies: Implement disciplined data mining. Let your historical LIS data continuously monitor the health of your reference intervals, and use a 20‑sample spot check only as a rapid confirmation tool when an alert fires.

By layering the simple 20‑sample verification with data‑driven surveillance and a hard‑nosed assessment of population fit, you can confidently deploy manufacturer‑provided reference intervals while actively protecting your patients from hidden mismatches.

Summary Table:

Verification Strategy Sample Size Primary Purpose Key Criterion / Action Key Limitation
20-Sample Protocol 20 healthy individuals Rapid check of reference interval transferability Verified if $\le$ 2 outliers out of 20 results Cannot detect a narrower local range
Outpatient Data Mining LIS historical data Continuous long-term surveillance & drift detection Monitor outpatient medians and % out-of-bounds Reactive; requires statistical discipline
De Novo Study $\ge$ 120 reference donors Full statistical determination of local reference limits Non-parametric calculation of central 95% interval Highly resource- and time-intensive

Achieving reliable assay performance and seamless regulatory compliance requires precision at every step of assay validation. CamelBio provides diagnostic manufacturers, labs, and research institutes with one-stop access to high-quality IVD raw materials, technical services, and consulting—covering every stage from concept to clinic. Whether you are validating new reagent kits or establishing local reference intervals, our technical experts are here to support your laboratory workflows. Contact CamelBio today to streamline your diagnostic development and clinical validation!


Leave Your Message