← Forefront

Living Knowledge · Methodology

How Forefront assesses confidence

Forefront scores individual scientific claims against the evidence it can retrieve and read. This page explains exactly how that number is produced, what it establishes, and — at least as importantly — what it does not. It is written in three levels: read the first if you only want to know how to read the number, the second for the science, the third for the equations.

Engine methodology v1.10.0Snapshot build 9Validation suite v1.2.0
Level 1

In 60 seconds

What an assessment is, and what it is not.

What a Confidence assessment is

It summarises one specific claim
An assessment covers a single named claim — one intervention, one population, one outcome — and never a drug in general. Two claims about the same drug are scored separately and can disagree.
It is a weighted share of the evidence we could read
A normalised weighted-evidence share, not a probability. Two claims with the same mass ratio score identically regardless of disease, drug or endpoint. The percentage is the share of graded evidence mass that supports the claim, measured against a fixed burden of proof assumed before anything is observed.
It is deterministic
The scoring calculation is a pure function of the assembled evidence: the same evidence always produces the same number, on every machine and every run. Nothing about the score is sampled, randomised, or generated at display time.
It shows its working
Every assessment lists the trials and publications that moved it, the direction each was read as, and why. A number you cannot audit is not a scientific instrument.

What it is not

Not a probability that the claim is true
This is the most important thing to understand about the number. It is not the probability the claim is true, and it is not a calibrated probability, a likelihood, or a posterior. An assessment of 70% does not mean the claim has a 70% chance of being correct. It means that, of the evidence we could read and grade, a particular weighted share pointed toward support.
A heuristic, not a statistical estimator
The engine is a weighted-evidence heuristic. Its constants were chosen to behave sensibly across a set of reference claims, not derived from a statistical model of how evidence relates to truth. Nothing in it estimates a quantity that exists in the world; it summarises the evidence in front of it under a fixed set of rules.
Not calibrated against real-world outcomes
Constants were fitted by coordinate descent to expert-assigned target bands on a small synthetic anchor set. No calibration against observed outcomes has been performed: no Brier score, reliability curve or ground-truth comparison exists.
Not a confidence interval
The range shown alongside the number is an uncertainty range. It is not a confidence interval and not a credible interval: it has no sampling distribution, no likelihood and no coverage guarantee. It is a deterministic heuristic that widens with thin, contested or concentrated evidence and narrows as resolving evidence accumulates.
Not a study-quality or risk-of-bias assessment
An engineering heuristic encoding the conventional evidence hierarchy. This is not a study-quality assessment: a double-blind placebo-controlled trial and an open-label randomised trial receive identical weight. Nothing in the engine performs a risk-of-bias appraisal, and no reviewer has appraised the studies behind an assessment.
Not a systematic review, and not medical advice
An assessment is an automated summary of retrieved literature and registry records. It is not a systematic review, not a clinical guideline, not a regulatory judgement, and must not be used to make clinical or treatment decisions.
Not a measure of all the evidence that exists
The engine can only weigh what the pipeline retrieved and could read. Evidence that was never published, never posted results, or could not be parsed is simply absent — and an absence lowers an assessment in the same way that genuinely weak evidence does.

How to read it

A higher percentage
More of the readable, graded, de-duplicated evidence pointed toward support — relative to evidence that contradicted the claim, well-powered evidence of no effect, and the fixed prior burden of proof.
A wider range
The evidence is thin, internally contested, or concentrated in a single detected source group. A wide range is a statement about how much the estimate could move, not about how likely the claim is.
A low percentage is not a refutation
Because supporting mass appears only in the numerator, a claim with no readable support scores near zero whether it was tested and found wanting or barely studied at all. Where that is the case the assessment is labelled insufficient readable evidence rather than headlined as a score.
Level 2

The scientific method

How evidence is weighed, matched and discounted.

Study design

Each piece of evidence is placed in a design tier, and that tier is the dominant term in how much it weighs. A meta-analysis counts for more than a randomised trial, which counts for more than a cohort study, which counts for more than a case series. Trial design is read from the registry's design fields — allocation, intervention model, study type — and not from the trial's phase, because a randomised masked trial is a randomised trial even when its phase is unrecorded.

An engineering heuristic encoding the conventional evidence hierarchy. This is not a study-quality assessment: a double-blind placebo-controlled trial and an open-label randomised trial receive identical weight.

Design tiers and their base weights
DesignBaseEffective weight
meta analysis32.4
systematic review2.21.76
rct1.61.28
cohort0.80.64
case control0.60.48
case series0.250.2
animal0.150.12
in vitro0.080.064
in silico0.050.04
expert opinion0.050.04

Limitation. Design tier is the only quality-like input the score has. A double-blind placebo-controlled trial and an open-label randomised trial of the same size receive identical weight, because blinding and pre-registration are never populated on production evidence.

Direction

Evidence moves an assessment only if the engine can establish which way it points. There are four outcomes, and the fourth is the common one.

  • Supports — Counts for the claim. It raises the score, and it narrows the uncertainty range, because there is now more evidence that resolves the question either way.
  • Contradicts — Counts against the claim. It lowers the score, and it widens the uncertainty range, because the evidence is now in dispute with itself.
  • No effect — A study large enough to have detected an effect, which found none. It counts against the claim, and it does two things to the range at once: it narrows it, because the question is better resolved than before, and it widens it where the claim also has real support, because a null and a positive result disagreeing is a genuine dispute.
  • Neutral — On topic, but the direction could not be read from the record. It contributes nothing at all: it does not move the score and it does not affect the range. It is kept only so the evidence counts stay honest about what was found and not used.

Limitation. Most retrieved records are neutral. A study existing is never treated as the study agreeing: where direction cannot be established from the record, the study contributes nothing to the number and appears only in the coverage figures.

Relevance to the claim

Evidence that does not quite match the claim is discounted rather than discarded. The engine compares three things: whether the study was treating or preventing, whether the enrolled population is the one the claim names, and whether it falls inside the claim's age band.

Unknown or mainstream scope receives no discount, so an unclassifiable study is never penalised for being unclassifiable.

Limitation. Relevance is coarse. Dose, duration, follow-up length, study setting, geography and ancestry are not modelled at all, and evidence is never transferred between different diseases.

Correlation and de-duplication

Repeated results from the same source are not repeated evidence. The engine handles this in two distinct steps, and claims only what it can actually detect.

First, de-duplication. Publications reporting the same underlying experiment collapse to a single strongest contribution, identified by registry number or trial acronym. Five papers about one trial count about once.

Second, correlation grouping. What survives is clustered by whatever shared influence the record reveals — a commercial sponsor, a research team, an institution, a publication venue. Inside a cluster each further piece of evidence contributes less than the last, so a single source cannot keep raising an assessment however much it publishes.

A detected correlation cluster (sponsor, research team, institution, venue, or shared trial). A floor on the correlation present, never proof of independence — undetected correlation stays undetected.

What we can detect is bounded, and it is bounded in a direction worth knowing about — every limit below causes evidence to be treated as less related than it really is, never more:

  • Who funded a publication is read from grant records. Company-funded trials are rarely written up as grants, so much industry-funded literature ends up looking unattributed and is treated as its own separate source.
  • The map from a trial programme to the company behind it is applied to reviews and summaries, but not to primary research papers.
  • A paper that names a trial only by its study name is not matched to the same trial's registry entry, so those two accounts of one experiment are not recognised as the same experiment.
  • Shared investigators, shared patient cohorts, shared contract research organisations and declared conflicts of interest are not detected at all.

Limitation. Grouping is a floor on the correlation present, never proof of independence. The limits above are the ones we know about; correlation we cannot see from the record stays undetected, and the engine never describes evidence as independent on the strength of having failed to link it.

Extraction confidence

When the engine reads a result it also records how much that one reading of that one result should count as directional evidence. This is the quantity most often misread, so it is worth stating precisely.

How much one reading of one reported result should count as directional evidence. Four things set it, and only the first is legibility in the narrow sense: where in the record the direction was read from (a title counts for less than a results section); whether the between-arm comparison was statistically significant; how probative the comparator was (placebo, standard care, none); and how many participants are known to have produced the result. Not study quality, not risk of bias, not statistical power, and not confidence that the result is correct.

A reading graded high, medium or low carries 1, 0.6 or 0.3 of its otherwise-computed weight.

Limitation. Extraction confidence says nothing about whether a study was well conducted, and it is not a statement that the extracted reading is correct. It bounds how far one reading of one reported result is allowed to carry an assessment.

Sample size

A larger study is measured more precisely. It is not thereby more supportive. The engine encodes exactly that distinction: sample size enters in two bounded places, and neither is a general multiplier on evidence weight.

First, the uncertainty range. Each resolving experiment carries a bounded size adjustment anchored at 1,000 participants, moving by 0.25 per ten-fold change in size and clamped to [0.5, 1.5]. These are mass-weighted across the resolving experiments and scale the range. Larger studies narrow it; smaller ones widen it.

Second, a ceiling on small supporting readings. A supporting read drawn from fewer than 30 participants can rise no higher than medium extraction confidence, and below 10 no higher than low. This is a ceiling, not a multiplier: it can only lower a reading that was above it, it is idempotent where a reading already sits below it, and it never touches contradicting or null readings.

Where the count is unknown it stays unknown. An absent sample size is neutral — an adjustment of exactly 1.0, and no ceiling applies. Unknown is never treated as small.

Limitation. The two places read different counts, deliberately and imperfectly. The ceiling judges the smallest population genuinely known about a reading; the well-powered-null gate still reads the recorded registry enrolment. Publications carry no structured participant count at all, so neither the size adjustment nor the ceiling can apply to them.

Uncertainty

The range is computed from the same evidence as the score, by a deterministic rule.

A deterministic heuristic half-width: the reciprocal of accumulated resolving mass — scaled by a bounded study-size adjustment — plus a contradiction term, plus a flat single-source penalty. Study size enters here and nowhere else. It has no sampling distribution and no coverage guarantee: two five-million-participant trials still yield a range of several points, where a statistical interval would be a fraction of one.

No sampling distribution, no likelihood, no posterior, no coverage guarantee. It is symmetric by construction and clamped, so a 0% assessment displays a range that extends only upward.

  • Narrowed by accumulating supporting mass.
  • Narrowed by accumulating well-powered null mass.
  • Narrowed by evidence spread across two or more source groups.
  • Widened by thin evidence.
  • Widened by contradiction (support vs harm).
  • Widened by all decisive evidence sitting in one source group.

Limitation. The two kinds of disagreement do not widen the range by the same amount. A study finding no effect and a study finding harm now both register as disagreement, on the same scale — but a null also counts as evidence that resolves the question, so a support-versus-null split still reads as somewhat more settled than an equal support-versus-harm one. That is deliberate rather than an oversight: a null answers the question, a contradiction disputes the answer.

Level 3

Technical methodology

Equations, constants, versioning and limitations.

Equations

The score
confidence = supportMass / (supportMass + contradictMass + nullMass + prior)

Supporting mass appears in the numerator and the denominator; contradicting mass and well-powered null mass appear in the denominator only. The prior is a fixed pseudo-mass of counter-evidence, so the expression saturates and can never reach 1.

The uncertainty range
band = clamp(baseBand / (1 + massScale * (supportMass + nullMass)) + contradictionWeight * contradictionFraction + contradictionWeight * nullDisagreement + concentrationPenaltyIfSingleGroup, min, max)

A half-width in confidence points, clamped to [0.02, 0.45]. The resolving-mass term is scaled by the mass-weighted study-size adjustment below.

The study-size adjustment
clamp(1 + 0.25 · log10(n / 1000), 0.5, 1.5)

Computed per resolving experiment, then mass-weighted across them after de-duplication. Unknown n yields exactly 1.0. It scales the range and enters the score nowhere.

Effect magnitude
clamp(1 + scale * log2(effect / clinicallyMeaningfulThreshold), min, max)

Applies only to a direction read from a trial's own reported results. It is never applied to a publication, to a result showing no effect, or to a partial supporting read — one where the effect is real but bounded, such as a subgroup-only finding.

The correlation discount
Within a source group, edges are sorted descending and the i-th contributes weight * discount^i.

The geometric series converges, so a single source group can never contribute more than its ceiling however many studies it publishes.

Constants

These are the values the engine runs on in production. They are read directly from the modules that define them, so this table cannot disagree with the code.

ConstantValueMeaning
prior1.4A fixed pseudo-mass of counter-evidence assumed before anything is observed. It sets the burden of proof and guarantees confidence never reaches 1.
designScale0.8Uniform multiplier on every design tier. Shifts the overall burden of proof without changing the hierarchy.
correlationDiscount0.55Within one detected source group, the i-th strongest contribution is scaled by this factor raised to the i-th power.
baseBand0.4The uncertainty half-width before any evidence narrows it.
massScale0.6How fast accumulating resolving mass narrows the range.
contradictionWeight0.2How much a fully contested support-versus-harm split widens the range.
concentrationPenalty0.09Flat widening applied while all decisive evidence sits in fewer than two detected source groups.
unreplicatedDecayPerYear0.08Annual decay on ageing unreplicated sub-trial supporting evidence. Never applied to contradicting or null evidence, so an assessment cannot drift upward with time.
extractionConfidenceWeighthigh 1 · medium 0.6 · low 0.3What one reading's extraction confidence is worth as a multiplier on its weight.
NULL_POWER_MIN300Recorded enrolment at or above which a non-significant result becomes decisive evidence of no effect rather than absence of evidence.
TINY_TRIAL_N / SMALL_TRIAL_N10 / 30Populations below which a supporting reading's extraction confidence is capped at low, and at medium, respectively.
sizeAdjustmentanchor 1,000 · 0.25 per decade · clamp [0.5, 1.5]The bounded study-size adjustment, applied to the uncertainty range only.

Assumptions

  • The conventional evidence hierarchy is a usable proxy for evidential strength, encoded as fixed design tiers.
  • Supporting evidence is the only thing that can raise the number. A claim with no readable support therefore scores near zero whether it has been tested and found wanting or hardly studied at all — which is why such assessments are labelled rather than headlined.
  • A well-powered non-significant result is evidence of no effect rather than absence of evidence, above a fixed enrolment threshold.
  • Correlation that cannot be established from the record is treated as absent, which systematically understates how correlated the evidence really is.
  • Regulatory approvals and safety actions are tracked on a separate axis and never enter evidential confidence.
  • Evidence the pipeline never retrieved, and results that were never posted, are indistinguishable from evidence that does not exist.

Word labels

The word cut-offs are an editorial default, not calibrated against any labelled set.

  • Very high 90%
  • High 75%
  • Moderate 50%
  • Low 30%
  • Very low 0%

Determinism and the role of AI

The calculation is always deterministic
Scoring is a pure function of the assembled evidence. No language model takes part in weighting, grouping, de-duplication, the uncertainty range, or the final number — not sometimes, not as a tie-breaker, not at all. Given the same evidence the engine returns the same result every time, which is why a stored assessment can be recomputed and checked against the number that was served.
Reading a trial result is also deterministic
Results reported to a clinical-trials registry arrive as structured tables — arms, participants, measures, statistics — and are parsed by fixed rules. No language model reads them, and none is consulted about what they mean.
Reading a journal abstract may optionally use a language model
Publications are prose, and prose is harder. A rule-based extractor runs first on every on-topic publication. A deployment may additionally enable a language-model fallback, which is offered only the leftovers: an on-topic publication, with an abstract, where the rule-based pass came back with no direction or low confidence. It is never offered a trial result, a regulatory record, or a publication the rule-based pass already read decisively.
The fallback can recover a reading, never overturn one
It must answer in the same constrained form as the rule-based reader, and any error, timeout or malformed answer leaves the rule-based result standing. A reading it recovers then enters the score through the same relevance, extraction-confidence and small-sample machinery as any other, and is subject to the same ceilings. It is not trusted more for having been recovered.
Whether a particular assessment used it is not currently shown
This is a real gap and we would rather state it than let the section imply otherwise. An assessment does not record which of its readings, if any, came from the fallback, so the page you are reading cannot tell you. Some assessments are produced by a path that never invokes it at all. What is always true is the first point above: however a reading was obtained, the arithmetic applied to it was deterministic.

Versioning

Three version axes are kept deliberately separate, so a reader can tell why a number moved. The engine methodology version is the scoring algorithm: a change to it reinterprets existing evidence. The snapshot build is an operational cache generation, never a methodology version. The validation suite version is the benchmark harness that gates changes. A confidence number that moved because new evidence arrived is a different event from one that moved because the methodology changed, and the versions are what let you tell them apart.

Methodology history
Last changed 2026-08-25 · v1.10.0

The scoring algorithm has 13 recorded versions, counting the original. Every change that can move a score is recorded against the version it produced, together with what moved and why — so a number that changed because the methodology changed can always be told apart from one that changed because new evidence arrived.

Scope

What the engine does not do

Capabilities a reader might reasonably assume are present, which are not. 7of them are defined inside the engine but never receive any data from real evidence, so they contribute the same constant to every assessment. A capability that never runs is not a capability, and each one’s consequence is spelled out below rather than left to be inferred from its presence in the code.

It does not assess study quality or risk of bias.
Design tier is the only quality-like signal in the score. Blinding and pre-registration are defined in the engine but never populated on production evidence, so they contribute a constant to every study.
It does not detect retractions.
No retraction check exists at any stage of the pipeline, so a retracted study continues to contribute.
It does not resolve a meta-analysis against the trials inside it.
A synthesis and its constituent trials can both contribute when both are in the corpus, which double-counts that evidence.
It does not distinguish surrogate endpoints from clinical ones.
Two identical trials, one measuring a blood marker and one measuring a clinical outcome, receive the same weight.
It does not read participant counts from publications.
Europe PMC exposes no structured count, and mining abstract prose would be inference rather than measurement. Publications therefore receive no size adjustment and no small-support ceiling.
It does not establish that any evidence is independent.
Grouping detects the correlation the record reveals and nothing more. Shared investigators, shared cohorts, shared contract research organisations and conflict-of-interest disclosures are not detected at all.
It does not let regulatory approval raise evidential confidence.
Approvals and safety actions are tracked on a separate axis, because approval reflects consensus and actionability, not proof of a mechanism.
It does not calibrate against what later turned out to be true.
No Brier score, reliability curve or ground-truth comparison exists. The constants were fitted to expert-assigned target bands on a small synthetic anchor set.
It does not give credit for replication.
Whether one study repeats another is not established, so a replicated finding is not strengthened for having been replicated, and it is not protected from the slow discount that ageing unreplicated evidence receives.
It does not prefer the original report of a study over later ones.
Where several accounts of one experiment compete to represent it, the engine has no record of which came first, so it keeps whichever weighs most rather than the primary report.
It does not measure how much evidence exists in the world.
Coverage figures describe the share of retrieved, eligible evidence that was incorporated — never a share of the literature.

Not modelled at all

  • dose
  • duration / follow-up
  • indication transfer between diseases
  • study setting
  • geography or ancestry
  • primary vs secondary endpoint (for publications)
Known limitations

Where the engine is knowingly not yet right

Every limitation the engine records as open is listed here, each one written out as what it means scientifically and what it does to a number. Each is also pinned by a test asserting the current behaviour, so none can quietly disappear into a passing test run, and a newly recorded limitation cannot ship without appearing on this page. They are published rather than held back: a limitation you can read is worth more than a number you have to trust.

  1. KG-09A very small study's result can escape the small-sample ceiling

    A supporting result read from very few participants is capped, so it cannot carry a claim as far as a larger one. That cap needs a participant count, and journal articles do not report one in a form we can read reliably. Where one experiment appears both as a registry record and as an article, the engine keeps whichever version weighs most — which can be the article, and the article cannot be capped.

    What this means for a number. In principle a result from a handful of participants could count as though its size were simply unknown rather than known to be small. No claim currently on the site is affected, and nothing detects the situation automatically, so we publish the limitation rather than wait for a case to appear.

  2. KG-08Whether a null result counts as decisive is judged on enrolment, not on who was analysed

    A study finding no effect counts as evidence of no effect only above a fixed number of participants; below it, the result is treated as absence of evidence instead. That threshold is checked against the number of people enrolled in the study, rather than the number actually analysed for the specific outcome — and enrolment is always the larger, more forgiving of the two.

    What this means for a number. Some null results are currently treated as decisive on the strength of an enrolment figure larger than the group the result was actually measured in. Judging them on the analysed group instead could only remove such results, never add them. Separately, the bigger constraint on null evidence is not this threshold at all — far more studies are set aside because no outcome could be read from them.

  3. KG-01A marker and a clinical outcome carry the same weight

    Evidence is weighed by how a study was designed, not by what it measured. A trial measuring a blood marker and a trial measuring something a patient experiences — survival, a heart attack, a hospital admission — are weighed identically when their designs match. Changing a marker is not the same kind of finding as changing an outcome, and the engine does not yet make that distinction.

    What this means for a number. A claim resting mainly on marker evidence can read as strongly as one resting on outcome evidence. The evidence list on each assessment names what every contributing study measured, so the distinction is visible to a reader even though the score does not make it.