Setting analytical performance specifications is not a guessing game—it’s a strategic, evidence-based process. As an IVD assay developer, you need a clear, defensible target for imprecision, bias, and total error before you ever run a validation study. The most robust framework for this is the Milan hierarchy, a three-tiered model that ties your assay’s technical goals directly to clinical need, human biology, or the best available technology.
The Milan hierarchy gives IVD developers a structured pathway: set specifications based on clinical outcome impact (Model 1), biological variation (Model 2), or current state-of-the-art performance (Model 3). Whichever model you choose, those targets then govern every critical validation parameter—trueness, precision, measurement range, detection limit, and specificity—ensuring your new test earns regulatory trust and delivers clinically reliable results.
The Milan Hierarchy: A Strategic Framework for Specification Setting
The Milan consensus models were developed to answer one deceptively simple question: How good does an assay need to be? Without a clear answer, reagent design, raw material sourcing, and validation efforts drift. The hierarchy gives you a decision logic, ordered from most clinically powerful to most practically expedient.
Model 1 – Clinical Outcome-Driven Specifications
This is the top of the hierarchy. Specifications are derived directly from the impact of analytical error on patient diagnoses and health outcomes.
You ask: What rate of diagnostic misclassification is clinically acceptable? For example, high-sensitivity cardiac troponin (cTn) assays often target a coefficient of variation (CV) ≤ 6% at the 99th percentile upper reference limit. Why? Because modeling shows this keeps false-positive and false-negative classification rates around 0.5%—a level considered safe for emergency triage of myocardial infarction.
Model 1 can use large-scale outcome studies or sophisticated indirect decision models. It’s the strongest evidence because it links a lab number directly to what happens to the patient. However, it’s also the most data‑intensive to build.
Model 2 – Biological Variation-Based Specifications
When outcome data aren’t available, the biology of the measurand itself becomes your blueprint. Model 2 bases goals on the natural within-subject (CVI) and between-subject (CVG) biological variation of the analyte.
Two core formulas govern countless IVD developments:
- Desirable imprecision: CVA ≤ 0.5 × CVI
This ensures analytical noise adds less than about 12% to the total test result variability, preserving the ability to detect true patient changes. - Allowable bias: B ≤ 0.25 × √(CVI² + CVG²)
This keeps results harmonized across labs and enables shared reference intervals.
Take a glucose assay. With a known CVI of ~5.7% and CVG of ~6.9%, Model 2 sets strict but achievable targets: imprecision ≤ 2.9%, bias ≤ 2.2%, and total analytical error ≤ 6.9%. These numbers directly shape calibrator formulation and reagent lot release criteria.
Developers often apply a three-tier performance classification within Model 2: minimum (0.75 × CVA), desirable (0.5 × CVA), and optimum (0.25 × CVA) levels. This creates a practical roadmap for continuous improvement from early prototype to final commercial kit.
Model 3 – State-of-the-Art Specifications
When neither outcome data nor reliable biological variation estimates exist, you turn to Model 3. It sets specifications based on the highest performance currently achieved by existing commercial assays or technical standards on the market.
Model 3 is inherently reactive—it asks, “What’s the best anyone can do today?”—rather than “What’s clinically necessary?”. It’s a legitimate fallback, especially for novel analytes or esoteric tests, but it should always be treated as a starting point, not the final ambition.
Translating Specifications into Measurable Performance Criteria
Setting targets is only half the battle. You must translate them into the five core validation parameters that regulators and clinicians scrutinize:
- Trueness (Accuracy/Bias)
- Precision (Repeatability & Reproducibility)
- Analytical Measurement Range (AMR)
- Limit of Detection (LoD)
- Analytical Specificity & Interference
Imprecision and Bias Targets Directly from the Models
Models 1 and 2 give you hard numbers for CV% (imprecision) and %bias—the most direct translation. For instance, a Model 2 specification of CVA ≤ 2.9% immediately defines your precision validation acceptance criteria under CLSI EP05 protocols. A bias limit of ≤ 2.2% drives your method comparison and trueness studies using certified reference materials.
Beyond Bias and Imprecision – Completing the Performance Profile
The other parameters—LoD, AMR, specificity—are still informed by the framework, even if indirectly.
A clinical-outcome model (Model 1) might dictate a LoD low enough to rule out disease at a specific decision cut-off. A biological variation model (Model 2) might justify a wide AMR to cover all physiological and pathological concentrations, while a state-of-the-art approach (Model 3) benchmarks your AMR against the leading competitor’s assay insert. Specificity and interference goals, meanwhile, are heavily shaped by the intended patient population (e.g., percent of hemolyzed samples, common medications). These must be validated using CLSI EP07, EP14, and EP37 standards to prove your reagent’s ruggedness across real-world matrix conditions.
Understanding the Trade-offs and Limitations
No single model is perfect. Your job as a developer is to navigate the tensions.
Model 1 is the clinical gold standard, but it’s rarely feasible early in development. It requires large, expensive outcome studies or pre-existing medical decision literature. For novel biomarkers, this data simply doesn’t exist.
Model 2 is elegant and evidence-based, but not universal. Biological variation databases (like the EFLM BV database) cover many common analytes, but for rare or newly discovered markers, reliable CVI and CVG estimates may be missing or derived from small, non-diverse populations. Applying a generic three-tier model without validating the underlying BV data can lead to over‑ or under‑engineered assays.
Model 3 is always available, but it’s a moving, often clinically blind target. Chasing the “best on market” can give you precise but clinically meaningless performance if the entire market already exceeds what patients need. Worse, it can lock you into a performance arms race that delays launch without adding value.
The three-tier performance model (minimum, desirable, optimum) introduces flexibility but also ambiguity. Developers must clearly define which tier aligns with their intended use claims and regulatory pathway. A minimum specification is rarely acceptable for a high-risk cardiac marker, yet it might be appropriate for a wellness screening test.
Making the Right Choice for Your Assay
Your selection should be driven by the clinical question your assay answers and the maturity of the available evidence.
- If your assay directly influences emergency or critical care decisions: Pursue Model 1. Invest in the literature or partner with clinicians to define an acceptable misclassification rate and build your entire reagent performance around that error budget.
- If biological variation data is robust and internationally accepted for your analyte: Adopt Model 2 as your primary framework. It yields objective, defendable goals for imprecision, bias, and total error that simplify lot release and regulatory submissions.
- If you’re developing a first‑in‑class biomarker with no outcome or BV data: Start with Model 3 to benchmark against the closest available technology. Treat this as an interim milestone, and commit to transitioning to Model 2 or 1 as clinical evidence accumulates.
- If meeting commercial cost‑per‑test targets is critical: Use Model 3 to identify the performance floor. Then, strategically push toward Model 2’s desirable tier only if it demonstrably improves diagnostic accuracy and market differentiation.
- Regardless of the model, always validate the five core parameters: Use CLSI and ISO protocols (EP05 for precision, EP09 for method comparison, EP17 for LoD, EP07 for interference, EP25/EP26 for stability) to prove your assay meets the chosen specifications across multiple reagent lots and operators.
By anchoring your analytical specifications in the Milan hierarchy, you stop chasing arbitrary numbers and start building a rigorous, clinical‑logic‑based foundation for every diagnostic kit you develop.
Summary Table:
| Milan Model Tier | Specification Basis | Key Target / Formula | Recommended Application |
|---|---|---|---|
| Model 1: Clinical Outcome | Impact of analytical error on patient health & diagnosis | Misclassification rates (e.g., cTn CV ≤ 6% at 99th percentile) | High-risk emergency or critical care markers |
| Model 2: Biological Variation | Natural within-subject ($CV_I$) and between-subject ($CV_G$) variation | Imprecision: $CV_A \le 0.5 \times CV_I$ Allowable Bias: $B \le 0.25 \times \sqrt{CV_I^2 + CV_G^2}$ |
Standard clinical chemistry & immunoassay analytes |
| Model 3: State-of-the-Art | Performance achieved by top existing commercial assays | Benchmarking against market-leading specifications | First-in-class biomarkers or novel target analytes |
Accelerate Your IVD Assay Development with CamelBio
Translating analytical performance specifications into robust, regulatory-ready diagnostic tests requires high-quality reagents and technical precision. CamelBio provides diagnostic manufacturers, clinical laboratories, and research institutes with one-stop access to premium IVD raw materials, technical services, and consulting—covering every stage from concept to clinic.
Whether you are selecting raw materials for prototype assays or optimizing lot release criteria for commercial production, our team is here to support your success. Contact CamelBio today to discuss your diagnostic reagent needs with our experts!