A candidate produces the intended effect in an animal model. What exactly do we know at that point?
Less than it seems. We know that an administration was followed by an effect. We do not yet know whether the candidate reached its target, whether it bound to it, or whether the observed effect is explained by the postulated mechanism rather than another. These three questions are demonstrated separately, and the pharmaceutical industry’s experience shows that neglecting them carries a measurable cost.
This note sets out what has to be established for a preclinical result to support a decision, and two methodological requirements that follow — how to establish a mechanism, and how to choose a readout that decides between competing hypotheses.
Three pillars, and what their absence costs
An analysis of Phase II decisions across forty-four programs run at Pfizer looked for what distinguished the surviving candidates from the others. The first finding is expected: most failures come down to a lack of efficacy. The second much less so — in 43% of cases it was not possible to conclude whether the mechanism had been adequately tested [1].
Let us pause on that figure. It does not bear on the quality of execution, but on what was known at the end of those programs: was the candidate ineffective, or had it never reached its target at a sufficient concentration? A result obtained under those conditions does not teach what is needed. It indicates that the compound, as administered, did not produce the expected effect — without saying why, and therefore without allowing the target to be abandoned on an informed basis.
The authors draw from this three elements to be demonstrated jointly, which they call the pillars of survival: exposure at the site of action, target binding, and the expression of pharmacological activity consistent with the first two. None replaces the others, and it is their conjunction that gives a result its value for decision-making.
Demonstrating exposure
The requirement is more precise than it appears. What has to be established is not that a dose was administered, but that a concentration of free drug, unbound to plasma proteins, was reached at the site of action, at a level exceeding the compound’s pharmacological potency, and for the intended duration [1]. Every departure from that statement introduces doubt about the adequacy of the exposure.
A practical difficulty arises at once: direct measurement at the site of action is often experimentally inaccessible, and blood sampling is used instead [1]. That approximation is legitimate, provided it is treated as such — the plasma concentration is not the tissue concentration, and the gap between the two depends on the molecule as much as on the tissue.
The temporal dimension is frequently neglected. A recent review notes that preclinical models are typically used at a single time point, which leaves aside the dynamic nature of the response — whereas a minimum duration of exposure is required for target binding to translate into a therapeutic response [2]. A candidate that reaches the required concentration for two hours and one that maintains it for two weeks do not raise the same question, and a single-time-point design does not distinguish them.
One precondition underlies all of this: the model must make the question possible. Where a candidate is directed against a human protein, exposure remains measurable in an animal expressing only the murine homolog — the compound can still be assayed. But target binding and the activity that depends on it cease to be measurable: two pillars out of three collapse, and the first does not suffice. This is what motivated, in the field of hemostasis, the construction of fully humanized models [4]. We develop this point in our guide to in vivo preclinical models.
Establishing that the effect really comes from there
Observing the expected effect after administration does not demonstrate that the postulated mechanism is operating. The result is compatible with the hypothesis; it does not establish it. An effect may arise from an unanticipated pathway, from a non-specific property of the compound, or from an indirect consequence of the treatment.
The methodological answer lies in a single experiment: remove the element assumed to mediate the effect, and check that the effect disappears. This is a loss-of-function control, and it turns a compatibility into a demonstration.
A recent example illustrates the logic well. A single-domain antibody designed to extend the circulating half-life of von Willebrand factor produces the expected effect; but it is the disappearance of that effect in animals lacking the FcRn receptor that establishes recycling mediated by that receptor as the cause [3]. Without that experiment, the explanation would have remained one plausible hypothesis among others.
The reasoning is identical to that of target validation, where perturbation alone distinguishes a driver from a passenger — a logic we set out in relation to CRISPR-Cas9 screens. Candidate evaluation does not dispense with it.
Choosing the readout that separates the hypotheses
A second requirement, less often stated, concerns the choice of readout. It starts from a simple observation: two different mechanisms can produce exactly the same observation.
If the circulating level of a protein rises after treatment, two explanations present themselves — the organism produces more of it, or it eliminates less. Measuring the level, however precisely, does not settle the matter: it is equally compatible with both hypotheses.
A criterion that separates them is therefore needed. In the example above, the ratio of propeptide to von Willebrand factor antigen plays that role: it is a clearance marker, and its reduction points toward slower elimination [3], whereas increased production would leave it unchanged, both species being secreted together. The measurement is no more precise than the previous one; it is simply discriminating.
Hence a principle that goes beyond the case: a readout is not chosen for its availability, its convenience or its precision, but for its capacity to decide between competing hypotheses. Which presupposes having made those hypotheses explicit before designing the measurement — for a criterion chosen after the fact is rarely discriminating by chance.
What a negative result says, and does not say
Between a positive and a negative result there is an asymmetry that on its own justifies the effort of measuring exposure.
A positive result without documented exposure remains fragile: the concentration that produced the effect is unknown, so the dose cannot be transposed, nor compared with another compound, nor used to anticipate what will happen in humans.
A negative result without documented exposure is of another kind: it is uninterpretable. A candidate that does not act cannot be distinguished from one that never reached its target in sufficient quantity, and the two situations call for opposite decisions — abandoning the molecule in one case, reformulating its administration in the other. This is exactly the situation described above for 43% of the programs analyzed.
The consequence is a design rule: exposure measurements are planned before the experiment, because they alone make a failure informative. Adding them after a disappointing result does not recover the information lost.
Tolerability, the other side
Evaluation is not confined to efficacy. The question that decides passage into the clinic is not only “does the candidate act”, but “how far from the poorly tolerated dose”. A narrow margin between the effective dose and the toxic dose compromises a program as surely as an absence of effect.
The two sides gain from being related to comparable exposures, failing which one sets against each other situations that are not comparable. This part of the evaluation falls under specific regulatory frameworks, whose detail lies beyond the scope of this note.
How Inovarion can support you
Inovarion supports preclinical evaluation projects from design through to interpretation: defining the study design in light of the questions asked, choosing the model and verifying its relevance to the candidate, exposure measurements, designing mechanistic controls and readouts, then analysis. Our teams have contributed to the development of fully humanized models intended precisely to make evaluable what was not. The question we ask first is rarely technical: it is what decision the result will have to support.
Publications
Field references
- Morgan P, Van Der Graaf PH, Arrowsmith J, et al. Can the flow of medicines be improved? Fundamental pharmacokinetic and pharmacological principles toward improving Phase II survival. Drug Discovery Today, 2012;17(9-10):419-424. PubMed
- Hughes EA, Davenport LL, Mochel JP, Douglass EF. Gaps and paths forward in cancer pharmacology and translational research. Frontiers in Pharmacology, 2026;17:1779049. DOI
- Peyron I, Casari C, McCluskey G, et al. Blood, 2025;146(21):2597-2607. DOI
Inovarion contribution
- McCluskey G, et al. A fully humanized von Willebrand disease type 1 mouse model as unique platform to investigate novel therapeutic options. Haematologica, 2025;110(4):923-937. PubMed
updated July 2026