Knowledge IVD Manufacturing What criteria should IVD manufacturers follow when selecting an APS model? Master the Milan Hierarchy
Author avatar

Tech Team · CamelBio

Updated 1 month ago

What criteria should IVD manufacturers follow when selecting an APS model? Master the Milan Hierarchy


Model selection is not a box-ticking exercise—it’s the strategic foundation of your assay’s clinical validity. IVD manufacturers and clinical laboratory consultants should select a model for establishing analytical performance specifications (APS) by applying the Milan consensus hierarchy, which pivots on three pivotal criteria: (1) the measurand’s direct role in clinical decision-making, (2) the availability of validated biological variation data and a steady-state biological behavior, and (3) the existence of robust clinical outcome or proficiency testing data. A structured decision workflow—first evaluating disease specificity and clinical consequence, then biological steady-state, and finally data availability—ensures specifications are neither arbitrary nor overly permissive but precisely matched to the assay’s medical purpose.

The right model emerges from the intersection of the biomarker’s clinical job, its biological behavior, and the evidence you can stand behind. For a measurand that directly drives a binary clinical decision, Model 1 (clinical outcomes) is mandatory. For most general analytes with stable homeostasis and valid biological variation data, Model 2 provides the most universally applicable yardstick. When neither outcome studies nor reliable biological variation data exist, Model 3 (state of the art) offers a defensible fallback—but only as a temporary anchor while better evidence is generated.

Understanding the Three Anchor Models of the Milan Hierarchy

The Milan consensus provides a three-tiered framework that translates a laboratory’s analytical capability into clinically meaningful specifications. Each model answers a different question about the assay’s performance limits.

Model 1: Based on Clinical Outcomes

This model is reserved for measurands that play a direct, central role in a specific clinical decision. It sets specifications by measuring how analytical error changes the rate of misdiagnosis or patient harm. For example, cardiac troponin (cTn) goals are defined not by biology alone but by the acceptable misclassification rate for acute myocardial infarction—a CV of 6% at the 99th percentile limit keeps false-negative results below 0.5%.

Model 2: Based on Biological Variation

When a biomarker maintains a steady-state concentration in healthy individuals and valid biological variation data exist, performance goals are derived from the ratio of analytical noise to biological signal. The classic specification for imprecision is CVA ≤ 0.5 × CVI (within-subject biological variation). This minimizes the analytical contribution to total variation, ensuring that a change in a patient’s result reflects a true physiological shift rather than assay noise. Glucose, for instance, sets imprecision ≤2.9%, bias ≤2.2%, and total error ≤6.9% based on this model.

Model 3: Based on State of the Art

This model applies when neither outcome data nor valid biological variation data are available—often for novel biomarkers or rare analytes. Here, APS are pegged to the highest technically achievable performance observed in top-tier proficiency testing (PT) or external quality assessment (EQA) schemes. It is a recognition that, for now, the assay must simply be as good as the best existing platform, with the understanding that specifications will tighten as evidence matures.

The Decision Framework: Three Criteria That Drive Model Selection

Selecting the right model is a sequential, criteria-driven process. Manufacturers and consultants should walk through each gate before locking in specifications.

Criterion 1: Does the Measurand Directly Dictate a Clinical Decision?

Ask whether the assay result stands alone as a quantitative trigger for a high-stakes action—initiating a therapy, confirming a diagnosis, or ruling out a condition. Measurands like troponin, HbA1c, or blood glucose answer “yes” here. For these, Model 1 is non-negotiable. The analytical specification must be derived from how much performance drift a clinician can tolerate before patient management changes inappropriately. If the answer is “no” or “only supportive,” proceed to the next criterion.

Criterion 2: Is the Biomarker in a Biological Steady State with Valid Variation Data?

A measurable analyte in a homeostatic steady state in the reference population can be described by within-subject (CVI) and between-subject (CVG) biological variation. If high-quality biological variation data (ideally validated via the BIVAC checklist) are available, Model 2 becomes the primary tool. This model works exceptionally well for general chemistry analytes—electrolytes, lipids, enzymes—where the goal is to keep the analytical signal clean enough to detect genuine physiological changes. Critically, the CVI value must come from a properly controlled study (homogeneous cohort, strict preanalytical control, duplicate analyses, ANOVA outlier exclusion). Data from poorly designed studies can produce misleading specifications.

Criterion 3: What Evidence Is Actually Available?

When neither clinical outcome studies nor robust biological variation data exist, the decision tree defaults to Model 3 (state of the art). Specifications are then set at the current technical capability limit—the performance achieved by the best 10–20% of laboratories in EQA programs or described in peer-reviewed validation studies. While this model lacks the direct clinical grounding of the others, it serves an essential purpose: it prevents acceptance of inferior performance and pushes the entire field to improve. It is also the typical starting point for novel biomarkers that lack extensive clinical history.

Beyond the Model: Translating APS into Validation Parameters

Once the model is selected, the resulting specifications must be translated into quantifiable assay validation criteria. The model gives the boundary; you fill it with these core parameters.

The Five Non-Negotiables of Analytical Validation

  • Trueness and Bias: Alignment with certified reference materials or a reference method, quantified as systematic error.
  • Precision: Imprecision (repeatability and reproducibility) expressed as CV% must fit within the model-derived budget.
  • Analytical Measurement Range (AMR) / Linearity: The concentration interval where the assay remains within the APS limits for bias and imprecision.
  • Limit of Detection (LoD) and Quantification: The lowest concentration reliably distinguished from blank matrix, critical for low-level cutoff decisions.
  • Analytical Specificity and Interference: The model’s total error budget must remain intact under realistic sample conditions, including potential interfering substances.

Ruggedness—The Hidden Connective Tissue

An assay that meets specifications on paper but fails under real-world stress is clinically useless. Evaluations of lot-to-lot consistency, operator variability, and instrument-to-instrument reproducibility must be built into the validation plan. The model’s goals are only meaningful if they hold up across time, users, and supply chains.

Navigating the Trade-offs and Common Pitfalls

No model is perfect. Understanding their limitations is essential to avoid overspecifying (wasting resources) or underspecifying (risking patient harm).

The Trap of “Model 2 for Everything”

Biological variation data are not a universal fix. Many biomarkers have significant preanalytical variation (e.g., posture, exercise, hemolysis) that eclipses analytical imprecision. Applying Model 2 without accounting for these confounders leads to specifications that are artificially tight or impossible to meet. Also, biomarkers with pulsatile secretion or rapid systemic clearance (e.g., some hormones) do not maintain a steady state, violating a core assumption of the model.

The Danger of Over-Reliance on State-of-the-Art

Model 3 is a lifeboat, not a permanent residence. An assay that merely matches current commercial performance may still be clinically unacceptable if the entire market is underperforming relative to patient needs. This model must be coupled with a plan to transition to Model 1 or 2 as soon as outcome data or standardized biological variation studies become available. Without that roadmap, IVD manufacturers risk locking themselves into mediocrity.

Ignoring the Decision Workflow Sequence

Jumping directly to Model 2 because biological variation data exist, without first ruling out a direct clinical decision role, is a critical error. Disease specificity and clinical consequence always take precedence. A measurand like troponin can be described biologically, but its APS must come from misclassification rates, not CVI, because the clinical decision (rule-in/rule-out MI) is too consequential.

How to Apply This Framework to Your Assay

The Milan hierarchy is not a rigid rule but a dynamic sequence for aligning analytical rigor with clinical truth. Use these goal-based paths to choose wisely.

  • If your primary focus is establishing specifications for a novel biomarker with no existing clinical outcome data: Start with Model 3 to set an achievable benchmark from the best available EQA data, but simultaneously launch prospective studies to gather the clinical outcome or biological variation evidence needed for a model upgrade.
  • If you have access to rigorously validated biological variation data and the measurand is in a stable homeostatic state: Adopt Model 2 directly, using the CVA ≤ 0.5×CVI formula, and validate that preanalytical factors do not overwhelm the analytical budget.
  • If the assay result directly drives a binary, high-stakes clinical decision (diagnosis, initiation of treatment, or risk stratification): Insist on Model 1. Derive your imprecision and bias limits from direct clinical outcome studies or indirect decision models that calculate acceptable misclassification rates at the medical decision limit.

Ground your selection in the biomarker’s clinical role first, its biological behavior second, and the available evidence third—that sequence will always produce analytical performance specifications that protect patients and satisfy regulatory scrutiny.

Summary Table:

Model Selection Criterion / Trigger Core Foundation & Logic Typical Applications
Model 1: Clinical Outcomes Assay result directly drives a high-stakes, binary clinical decision Tolerable misclassification rates and clinical harm limits Cardiac Troponin, HbA1c, Blood Glucose
Model 2: Biological Variation Biomarker is in steady state with validated biological variation data Ratio of analytical noise to biological signal ($CV_A \le 0.5 \times CV_I$) General chemistry, electrolytes, lipids, routine enzymes
Model 3: State of the Art Neither outcome studies nor biological variation data are available Highest technically achievable performance in EQA/PT schemes Novel biomarkers, rare analytes, emerging diagnostics

Accelerate Your Diagnostic Assay Development from Concept to Clinic

Establishing defensible Analytical Performance Specifications (APS) requires strict methodological rigor and uncompromising assay stability. CamelBio provides diagnostic manufacturers, labs, and research institutes with one-stop access to IVD raw materials, technical services, and consulting—covering every stage from concept to clinic.

Whether you are defining specifications for a novel biomarker or optimizing high-precision reagents to meet strict clinical outcome goals, our team is ready to empower your development process.

Contact CamelBio Today to Support Your Assay Validation


Leave Your Message