1. Start with the failures that threaten the result

Identify inadequate reference response, missing concentration coverage, excessive replicate variation, unsupported model behavior, incomparable reference and test shapes, poor controls, or a result outside the validated range. Add a criterion only when its outcome protects an intended decision.

Evaluate reference and control performance before relying on the test comparison, then data and replicates, model evidence, comparability, and result conditions. A high R² cannot establish every role, and a narrow CV does not prove accuracy or comparability.

Analytical roleIllustrative protected failure
Reference/controlReference response or recovery does not support the run
Data/replicatesToo little coverage or excessive within-level variation
ModelConvergence, parameters, or residual behavior is unacceptable
ComparisonReference and test do not support one potency relationship
ResultEstimate or uncertainty falls outside the reportable condition

2. Define each rule as an executable record

Save the rule name and version, protected failure, sample scope, source observations, calculation and summary level, comparison operator, inclusive limit or range, units, missing-value behavior, and explanatory text.

Calculate and store the observed value separately from the PASS, FAIL, or not-evaluable comparison. A label such as “precision check” is incomplete unless a reviewer can identify the exact interval-relative-to-estimate calculation and limit.

FieldWorked meaning
CriterionRelative-potency range or interval precision
Observed valueCalculated from the retained model result
RuleInclusive method-defined range or maximum
OutcomePASS, FAIL, or not evaluable
Overall combinationAll required criteria must pass

3. Carry observed values through one overall decision

In this illustrative rule set, the selected model passes its adequacy terms, the relative-potency estimate is 80.0%, and its confidence interval is 62.0% to 98.0%. The point estimate passes the configured 50.0% to 150.0% reportable range.

Interval-relative-to-estimate width is (98.0 − 62.0) ÷ 80.0 × 100 = 45.0%. Because the configured maximum is 40.0% inclusive, precision fails. The retained potency remains 80.0%, but one required failure makes overall suitability FAIL. These synthetic limits illustrate execution; they are not recommended universal criteria.

Formula(98.0 − 62.0) ÷ 80.0 × 100 = 45.0%
Illustrative current-model suitability combination
CriterionObserved valueConfigured ruleOutcome
Model adequacyAll required terms evaluable and passingEvery required adequacy term passesPASS
Relative-potency range80.0%50.0% to 150.0%, inclusivePASS
Interval precision45.0%≤ 40.0%FAIL
Overall suitabilityOne required failureEvery required criterion passesFAIL

4. Preserve the outcome before disposition

Display the observed value, configured rule, version, and original outcome. The method and quality procedure determine whether the failed run is invalidated, investigated, repeated, or otherwise handled; software should not soften FAIL to “review.”

If an authorized exception or override process exists, retain the failure, rationale, user, timestamp, decision, and applicable procedure. An override is a separate disposition record, not a recalculated pass.

5. Establish and lock criteria before routine use

Use development, robustness, validation, historical, and deliberately challenged data to understand distributions, failure detection, and false-failure burden across relevant analysts, instruments, lots, days, and sample levels.

Approve the criterion set with the method version. A changed limit, scope, calculation, or missing-value rule creates a new version and requires impact assessment; it does not silently alter prior reports.

6. Report criteria and overall suitability together

For every rule, show its name and version, scope, observed value, units, comparison and limit, PASS, FAIL, or not-evaluable status, and explanation. Show the overall combination rule and outcome beside the potency result.

Keep source observations and diagnostic fits traceable from the report. The laboratory selects and validates scientifically appropriate criteria for its intended workflow.

Limits and non-universal thresholds

No universal relative-potency range or interval-precision maximum applies to every bioassay. Criteria must match the assay, response level, model, intended use, and approved procedure.

Provenarium can evaluate configured criteria deterministically and retain their observed values and outcomes, but it does not select or validate customer acceptance criteria. Suitability also remains distinct from sample disposition and any investigation.

Frequently asked questions

What is a sensible R² threshold for 4PL?

There is no universal threshold, and R² alone is insufficient. Evaluate assay-specific diagnostics, reference behavior, coverage, precision, and intended use.

Should every test sample have the same CV criterion?

Not necessarily. Variability can depend on response level, sample type, and design. Define criteria with justified scope and units.

What should be saved for each suitability result?

Save the rule and version, scope, observed value, units, comparison and limit, PASS, FAIL, or not-evaluable outcome, missing-value behavior, and the source observations used to calculate it.

Primary references