Define what each check protects against

A criterion should correspond to a failure mode that threatens the reportable result: inadequate reference response, excessive replicate variation, unsupported asymptotes, incomparable curve shapes, poor controls, or a result outside the validated range.

Avoid redundant collections of familiar metrics. A high R² can coexist with systematic model error, and a narrow CV does not prove accuracy or comparability.

Organize criteria by analytical role

Evaluate the reference or control before relying on the test comparison. Then assess data coverage, replicate behavior, model diagnostics, curve relationship, and final result conditions.

CategoryIllustrative checks
Reference/controlResponse window, midpoint range, control recovery
Data/replicatesMinimum included levels, replicate SD or CV, missingness
ModelConvergence, parameter plausibility, residual or lack-of-fit evidence
ComparisonParallelism/equivalence or shared-parameter suitability
ResultReportable range, uncertainty or precision where defined

Define enough detail to reproduce every check

For every rule, save its name and version, the samples it applies to, the measured value, summary level, comparison, limit, units, handling of missing values, and explanatory text. A label such as “CV check” is incomplete if a reviewer cannot tell which replicates were included or which SD convention was used.

Calculate and store the observed value separately from the PASS or FAIL comparison. Reports can make the outcome easy to scan, but the number, rule, and units must remain readable without relying on color.

A failed criterion should remain a failure

Display the observed value, expected rule, and FAIL outcome. Do not translate it to “review” simply to soften the interface. The method and quality procedure determine whether the run is invalid, investigated, repeated, or otherwise handled.

If an authorized override workflow exists, preserve the original failure, rationale, user, timestamp, decision, and applicable procedure. An override should not recalculate history as a pass.

Establish criteria using representative and challenged data

Use development, robustness, validation, and historical performance to understand distributions and failure behavior. Assess sensitivity and false-failure burden across analysts, instruments, lots, days, and sample levels relevant to routine use.

Lock the approved criteria with a method version. Changes should create a new version and trigger impact assessment rather than silently modifying earlier reports.

Frequently asked questions

What is a sensible R² threshold for 4PL?

There is no universal threshold, and R² alone is insufficient. Evaluate assay-specific diagnostics, reference behavior, coverage, precision, and intended use.

Should every test sample have the same CV criterion?

Not necessarily. Variability can depend on response level, sample type, and design. Define criteria with justified scope and units.

What should be saved for each suitability result?

Save the rule and version, scope, observed value, units, comparison and limit, PASS or FAIL outcome, missing-value behavior, and the source observations used to calculate it.

Primary references