ortHOTalk home

How the score is calculated

Every paper on this site carries two numbers. Rigor asks can I believe this? and impact asks does it matter? They are combined into one heat score that ranks the week:

heat = impact × credibility(rigor) − recency
credibility(r) = 0.35 + 0.65 (r/100)0.7

Multiplication rather than an average, because rigor is a credence and impact a magnitude: their product is expected-value-shaped, which is what a ranking should sort on. The floor of 0.35 means weak methods discount a finding rather than annihilating it, and credibility reaches exactly 1.00 at rigor 100 so nothing is ever inflated. The recency term takes off 1.5 points a week after the first week, capped at 12. A retraction zeroes all three numbers, and nothing else can.

Three rules shape everything below. Design sets a starting point and a ceiling, not a verdict. Gates fire only on observed defects, never on absence. Signals add up inside a domain, and domains cap rather than sum. The last is the shape of RoB 2, ROBINS-I and AMSTAR-2, and it is the answer to the 1999 demonstration that additive quality scales give contradictory answers about the same trials.

Drive it yourself

This runs the actual scoring module, not a simplified copy of it. Set a design and an appraisal and watch the numbers move. Three things are worth trying: switch a criterion between no and not reported and watch the score refuse to drop; drag the confidence interval from a decisive null to a demonstrated benefit; and turn on a spin flag to watch a cap clamp the score rather than nudge it.

PThe paper

Design family
Patients randomised
Centres
Journal tier

RHow it was done

Primary outcome
Outcome type
Allocation concealmentneutral when silent
Blinded outcome assessmentneutral when silent
Intention-to-treat analysis
Prospective registrationneutral when silent
Power calculation
Adjustment for confounding
Spin: conclusions the primary outcome does not support

IWhat it found

Confidence interval, in MCID units
Outcome tier
Relation to existing guidance
Topic, matched against the recommendation table
How many patients this affects

Heat · importance discounted by believability

Strong81.3roughly 7587
Rigor
98
Impact
82

heat = impact 82 × credibility(98) 0.991 = 81.3

coverage 1.00 · uncertainty half-width ±6

A precise null: the 95% CI -1.33 to 2.84 rules out a difference as large as the MCID the paper cites of 5 in either direction.

Zone Z5 · magnitude score 63.9

What each criterion is worth right now

Each row re-runs the scorer with that one field changed. Note that not reported never scores below no.

CriterionYesSilentNo
Allocation concealment989283
Blinded outcome assessment989081
Random sequence generation989588
Intention-to-treat analysis989488
Prospective registration989292
Aligned time zero989898
Competing risks handled1009898
Harms reported989898

Allocation concealment was actually performed in 100% of surveyed trials but reported in 41%, and power calculations performed in 76% but reported in 16%. Reporting is a poor proxy for conduct, so silence contributes exactly zero. Only an affirmed defect subtracts.

The widget exposes the criteria that move a score most, not all 65 fields of the appraisal, and the confidence-interval slider is read directly in MCID units. Everything else is the shipped code: the same computeHeat the weekly pipeline calls.

Rigor, in full

Rigor starts from a design prior drawn from a 47-family table and is held under a ceiling from the same row. The starting points are compressed and the ceilings wide, so execution decides where in the corridor a paper lands. A sham-controlled trial starts at 68.6 with no ceiling below 100; an adjusted registry cohort starts at 37.4 with a ceiling of 80 and therefore has 42.6 points of headroom; a retrospective case series starts at 8.8 under a ceiling of 42.

The ceiling is structural rather than a judgement about quality. A cross-sectional study cannot establish temporality however well it is conducted, and a cadaveric construct cannot speak to patient outcomes. Preclinical families are exempt from the spread that separates the clinical hierarchy, because their limitation belongs to impact, not to rigor: good bench science should read as good science answering a different question.

Each credited criterion is then summed and shrunk toward that prior by A / (A + 1.5), where A counts the criteria that could actually be assessed. A paper appraised on two criteria sits near its design prior; one appraised on twelve is scored on its own execution. Journal tier is the only proxy left in the model, worth at most 6 points, and it decays to exactly zero once 80% of the applicable criteria have been read.

D1 · Allocation and comparison validity
Concealed allocation+6
Allocation not concealedon a subjective outcome; −5 on an objective one−11
Random sequence generation+3
Quasi-random allocation−8
Sham-controlleda bonus, and does not enlarge the coverage denominator+8
Adjustment for confoundingnone 0, matching +3, multivariable +5, propensity score +8, IPTW +11, g-methods or IV +130 to +13
Exposure defined after cohort entryimmortal time bias, and it caps rigor at 25−14
Comparatorplacebo or sham +6, current standard +4, outdated −10−10 to +6
D2 · Measurement integrity
Blinded outcome assessmenton a subjective outcome; +4 on an objective one+8
Unblinded assessment of a subjective outcomethe largest single criterion penalty in the model−11
Registry coding validated+3
Independently adjudicated outcomes+4
Registry capture 95% or better+5

Blinding carries the heaviest penalty because the evidence says it should. Unblinded assessors exaggerate odds ratios by about 36% and standardised mean differences by 68%, and in orthopedics specifically the gap is a standardised mean difference of 0.76 unblinded against 0.25 blinded.

D3 · Completeness
Intention-to-treat analysis+4
Per-protocol analysis only−7
Follow-up under half the window the question needsand it caps rigor at 60. Arthroplasty survivorship needs 60 months; fracture union needs 6−12
Follow-up at twice the threshold or better+4
Loss to follow-up above 30%above 20%, −5; at or below 20%, +2−10
D4 · Reporting integrity
Adequately powered+7
Underpowered−10
No power calculation reportednot scored at all: it holds on 78% of the corpus, so it is an offset rather than a discriminator0
Prospectively registereda bonus and never a penalty, because only about a quarter of orthopedic surgical trials report it+6
Headline is not the primary outcome−8
No correction for multiple comparisons−4
D5 · Precision
A positive result from under 50 patientsunder 100, −4−8
A positive result from a single centre−5
Fragility index at or below the subspecialty medianabove it +1, at three times it +3. Medians: arthroplasty 1, sports and spine 2, trauma 3, foot and ankle 6−5
More patients lost than the result can absorb−8
Reverse fragility index of 5 or morefires only on a non-significant result, where a high value means the null is robust+4

The size and single-centre penalties fire only when the primary outcome was significant, because small-study and single-centre inflation are directional: they exaggerate positive findings and cannot manufacture a null.

Synthesis · the AMSTAR-2 critical domains
Pools randomised trialsmixed designs −4, observational −8, uncontrolled case series −14+6
I² of 75% or more, unexplained50% or more unexplained −9, explained by subgroup −4, under 50% +3−18
GRADE certaintyhigh +8, moderate +3, low −5, very low −10−10 to +8
Two or more critical failuresone critical failure caps at 72cap 55
Hard caps · every one fires on an observed defect, never on silence
Immortal time bias25
Data leakage between training and test sets30
Two-gate sampling, or conclusions the results do not support40
A pilot making an efficacy claim, or an unvalidated prediction model45
An effectiveness claim with no adjustment for confounding50
A conclusion asserting a benefit the primary outcome did not show55
The outcome window was not reached for the question asked60

Impact, in full

Impact starts from a flat 45 for every paper, with no design prior at all. Design enters only at the bottom, as an absolute ceiling. The dominant term is not the author's conclusion but the confidence interval, normalised into units of the minimal clinically important difference and signed so that positive means benefit.

That choice is deliberate. Conclusion strength is exactly what spin inflates: 44.8% of orthopedic trials with a non-significant primary outcome contain spin, and spun papers are cited more, so scoring the rhetoric would rank the most spun papers highest by construction. The authors calling their own work practice-changing is worth two points, and it does not even count toward the shrink denominator.

The zone classifier · L is the bound nearest the null, U the far bound, E the point estimate
Z1 · the whole interval clears the MCIDa demonstrated clinically important effect82 to 100
Z2 · the point estimate clears it, the interval reaches below62 to 82
Z5 · a precise null, both bounds inside the MCIDthe question is answered: an important difference is ruled out in either direction56 to 69
Z3 · significant, spanning trivial and important48 to 62
Z4 · significant and clinically trivialreal, but noise wearing a p-value30 to 48
Z6 · straddles zero and still allows more than the MCIDinconclusive, not negative20 to 40
Z7 · a demonstrated harma harm of a given size scores exactly what a benefit of that size scoresthe benefit ladder, mirrored
Z8 · not scoreable, or declined for a stated reason50

Where a null sits, and why. A trial that demonstrated a benefit patients can feel ranks slightly above an otherwise identical trial that returned a decisive null. Holding a multicentre trial's design, blinding, registration and guideline collision fixed and changing only the effect estimate, the demonstrated benefit scores impact 88 and the decisive null 82. The margin is small on purpose: both answer their question, and one of them also changes what we do.

What that re-ranking deliberately did not touch is more important than the re-ranking itself. Z5's floor of 56 sits above Z4's ceiling of 48 by construction, so a decisive null still beats a statistically significant but clinically trivial result by a wide margin, and no amount of p-value polishing lifts a trivial finding past it. A decisive null also still outranks every observational design. A null is not weak evidence. It is just not, on its own, a reason to change an operation.

The MCID divisor
The MCID the paper itself citesit is the bar the authors were judged againstalways wins
A published library valueKOOS JR spans 0.5 to 36.6 across sixteen derivation methods and mHHS 7.2 to 16.8, so picking one would be an editorial act wearing the costume of arithmetic. Where they disagree the magnitude scores neutral and the card says soonly where thresholds agree
A ratio measurea 25% relative reduction
A standardised mean differencerather than Cohen's arbitrary 0.5the median anchor-based important change
Outcome tier · applied multiplicatively, so a large effect on a surrogate cannot out-score a modest one on revision
Mortality or major morbidity1.00
Revision or reoperation0.97
Patient-reported function or pain0.95
Clinician-measured0.80
Radiographic or physiologic surrogatean affirmative claim on a surrogate primary also caps the magnitude at 55, because such findings are inflated roughly fourfold0.65
Laboratory or biomechanical0.50
Guideline collision · by the strength of the recommendation actually matched
Contradicts itStrong +20, Moderate +18, Limited +14, Consensus +12+12 to +20
Fills a gap it statespays most where the guideline concedes the evidence is limited+8 to +18
Confirms itconfirming a weak recommendation is worth more than confirming a strong one, because a Strong recommendation needs two or more concordant high-quality studies. That row is what stops the model punishing replication+2 to +10
A relation we cannot tie to a real recommendationit is then a claim about the literature, not a collision with guidance× 0.4
The rest of the impact sum
A powered, registered null that settles the question+5
First evidence on the questionsettles conflicting trials +5, first randomised evidence +4, adds to already consistent evidence −2+6
Collision and novelty togethercontradicting a recommendation and being first are partly the same claim, so they share one budgetcapped at +24
Comparatorsham +9, current standard +6, accepted alternative +3, outdated −3, none −5−5 to +9
How many patients this affectsfloored at +3 for a limb- or life-threatening problem, so prevalence weighting cannot bury sarcoma or a mangled extremity−2 to +7
Needs no new equipmentrequires equipment most centres lack, −4+4
Not reproducible from the description−12
Harms reportednot reported, −2+2
Five or more centres+4
Absolute impact ceilings · whatever the execution
Target trial emulation78
Prospective or adjusted registry cohortthe anchor: a 17-centre prospective cohort halving reoperation is clearly important and clearly not a landmark72
Meta-analysis of observational studies70
Retrospective comparative cohort65
Cross-sectional55
Cadaveric constructanimal and finite element 3540
Case report, editorial, narrative review25
Trials and syntheses of trialsuncapped

These ceilings are not a statement that a given cohort is poor. The top of the impact scale is reserved for a finding that overturns a guideline or becomes one, and a single observational study structurally cannot do either: GRADE starts observational evidence at low certainty, and a Strong recommendation needs two or more concordant high-quality studies. Reaching above impact 85 also requires that at least 60% of the applicable criteria could actually be assessed. A sparse paper can outrank a fully appraised bad one, which is correct because we do not know it is bad, but it cannot win on ignorance.

Why you see a band, not a decimal

Scores are computed continuously and displayed as one of five bands: exceptional from 85, strong from 68, sound from 48, limited from 30, weak below that. Composite critical-appraisal agreement runs around an intraclass correlation of 0.65 to 0.70, which supports four to six distinguishable strata and not a decimal place. Every card also carries an explicit uncertainty half-width of 6 + 14 × (1 − coverage): about ±6 when the whole paper could be appraised, and ±20 when there was no full text to read. A bare 78 out of 100 asserts a precision the underlying judgement cannot support.

Two things are deliberately absent. Citation counts, because at one or two weeks old they measure indexing speed rather than importance. And reader-relevance signals such as Canadian authorship or subspecialty, because they say nothing about how good a paper is; they inform editorial triage instead. Heat is a screening tool, not a verdict. Cold papers can still be great papers.