ortHOTalk home

About ortHOTalk

Orthopedic literature moves fast, and clerkship moves faster. This site does the scanning for you: every week it pulls newly published papers from the core orthopedic journals, reads the ones it can reach in full, scores each on rigor and impact, and publishes plain-language summaries pitched at medical students and junior residents.

Where the papers come from

OpenAlex
metadata, abstracts, open-access links
PubMed
publication types, retraction notices
Europe PMC
full text from the PMC deposit
Unpaywall
the remaining open-access copies

Only papers from a curated whitelist of orthopedic and major general-medicine journals make the cut.

The weekly pipeline

Once a week, an automated run:

  1. pulls everything the whitelisted journals published in the last week from OpenAlex, and drops anything already in the archive;

  2. adds PubMed's publication-type tags (that's where “RCT” and “meta-analysis” come from) and discards anything retracted;

  3. fetches and reviews each paper's full text, from open-access sources such as Europe PMC and Unpaywall;

  4. for every paper it can actually read, has an AI model (Claude) write the summary and fill in a structured critical-appraisal checklist from the full text;

  5. turns that appraisal into the rigor and impact scores, ranks the week by heat, and hands the top of the list to editorial review.

The rule the whole pipeline enforces: no paper receives a heat score until its full text has been read and appraised.

Papers whose full text isn't available yet simply wait, unscored, and are retried every week; open-access copies often appear weeks after publication.

Rigor

Rigor asks: can you believe the result? Every paper starts from a prior based on its study design (a randomized trial or meta-analysis can exclude more bias than a retrospective series, so it starts higher) and carries a ceiling from the same table that no execution quality removes: a cross-sectional study cannot establish temporality, a cadaveric construct cannot speak to patient outcomes. The starting points are deliberately narrow and the ceilings deliberately wide, because execution is supposed to decide, not the label. Journal tier is worth a few points at most and decays to nothing once most of the methods have been read directly, so reputation is never counted twice.

What the appraisal checks

  • Loss to follow-up
  • A-priori power
  • Allocation concealment
  • Blinded assessment
  • Intention-to-treat
  • Prospective registration
  • Crossover between arms
  • Primary outcome
  • Fragility index
  • Follow-up length
  • Competing risks
  • Conflicts of interest

The appraisal then moves the score on the things a journal club would check: loss to follow-up, an a-priori power calculation, allocation concealment, blinded outcome assessment, intention-to-treat analysis, prospective registration, crossover between the arms as randomised, whether the headline result is actually the study's primary outcome, and whether follow-up was long enough for the clinical question being asked. For significant dichotomous results we also compute the fragility index (how many patients would have to flip outcomes for significance to vanish) and flag the classic tell where more patients were lost to follow-up than the result can absorb. Meta-analyses and systematic reviews are further judged on what they pool (a meta-analysis of retrospective series is not Level I evidence), whether publication bias was assessed, and the AMSTAR-2 critical domains: one critical failure caps rigor at 64, and two or more cap it at 42. I² and the review's reported GRADE certainty do not receive conduct points: they adjust ranking credibility once, separately. Author royalties on the device under study cost a little; a sponsor-run analysis on top of those conflicts costs more.

When a paper simply doesn't report something, that scores as neutral, never negative: absence of evidence about the methods is not evidence of bad methods. A property test now holds the whole model to that contract: observing a strength can only raise a score, observing a defect can only lower it, and silence moves nothing. An earlier build shrank the summed appraisal toward the design prior as the appraisal thinned, and because that multiplier grew as criteria were added it could let an observed defect raise the score, so it was removed in favour of one constant. A thinly appraised paper now sits near its prior simply because few criteria moved it. Sparseness widens the uncertainty band and subtracts the excess uncertainty from the live rank, and a paper read on under 60% of its applicable criteria has its clinical importance held at 85, so it cannot win the week on ignorance.

Clinical importance & editorial priority

Clinical importance asks: even if it's true, does it matter? Every paper starts from the same flat 45, and the dominant adjustment is where the confidence interval falls once it is measured in units of the minimal clinically important difference (MCID) of the outcome instrument. Orthopedic outcomes are dominated by patient-reported scores with published MCIDs, and a statistically significant three-point change on a scale whose MCID is ten is noise wearing a p-value. Every extracted effect now carries its printed unit, so a risk difference of 0.05 cannot be mistaken for five percentage points. Risk ratios and odds ratios are converted with their own formulas using the control-arm rate. Hazard ratios decline rather than being treated as risks unless time-specific baseline survival is available.

44.8%

of orthopedic trials with a non-significant primary outcome contain spin, and spun papers are cited more often.

0vs25

points: the authors calling their own work practice-changing, against what the measured effect can move.

Clinical importance is deliberately not built on how strongly the authors phrase their conclusion, because that is exactly what spin inflates: 44.8% of orthopedic trials with a non-significant primary outcome contain spin, and spun papers are cited more often. The authors calling their own work practice-changing is worth 0 points. What the data show is worth up to 25.

Beyond that, patient-important outcomes score above radiographic surrogates, reporting harms earns a small one, and a weak or missing comparator costs points. A trial that screened four thousand patients to randomise sixty has answered the question for a sliver of the people you treat, however cleanly it answered it, and that narrowness costs points as well. Guideline collision and novelty now live on a separate editorial-priority axis: they may change what should be read this week without pretending the treatment effect is larger. Because those signals overlap, they share a budget capped at 24.

Every design carries an absolute clinical-importance ceiling no execution removes. A well-run propensity-matched registry cohort stops at 72, a flawless cadaveric study at 40, and even a flawless multicentre trial at 92: the top of the scale is reserved for synthesis, because a clinical-importance score of 99 means a finding that overturns a guideline or becomes one, and a Strong recommendation needs two or more concordant high-quality studies. None of this is a statement that a given paper is poor: it says a single study cannot on its own close a question, and that a bench study has demonstrated a mechanical property rather than a patient benefit. The week's open questions are handed to the appraisal, which records which one a paper answers if it answers any; that record earns nothing in the score.

Impact ceilings by design

Meta-analysis of trial data100
Multicentre trial92
Propensity-matched registry72
Cadaveric / bench40

The top of the scale is reserved for synthesis: an impact of 99 means a finding that overturns a guideline or becomes one.

The heat score

The operator

heat = [clinical + 0.5(editorial − 45)] × credibility(rigor)

Credibility rises from a floor of 0.35 at rigor 0 to 1.00 at rigor 100. Multiplication is the right operator because rigor is a credence and clinical importance a magnitude, so their product is expected-value-shaped, which is what a ranking should sort on. The 0.35 floor is why weak methods discount a finding rather than annihilate it.

1.000.35RIGOR 0100

A small recency decay comes off the top so this week's papers outrank last month's leftovers: 1.5 points a week past the first, and never more than 12 in total. A retraction zeroes the score outright; nothing else can.

Here is what that produces. Rigor runs across, clinical importance runs up, and every point on one contour scores the same heat. These are the worked archetypes the scoring code is tested against, plotted at the coordinates it actually emits.

heat 30heat 48heat 68heat 85020406080100020406080100RIGOR (CAN I BELIEVE IT?)CLINICAL IMPORTANCECEILING 72 · OBSERVATIONAL COHORTSCEILING 40 · PRECLINICALI · Demonstrated benefit95/78 → heat 79.7B · Moseley100/73 → heat 79A · PROFHER95/72 → heat 73.8F · Registry cohort71/65 → heat 54.6D · Meta-analysis of RCTs77/56 → heat 42.5C · Small positive trial43/50 → heat 28.6H · Cadaveric construct35/40 → heat 17.4E · Meta-analysis of cohorts35/45 → heat 7.8G · Case series9/42 → heat 11.6A trial that demonstrated a benefit patientscan feel (I) edges out the same trial run to adecisive null (A): clinical 78 against 72.A well-executed registry cohort (rigor 71)outscores a badly executed randomisedtrial (43). Design sets the starting pointand the ceiling. Execution decides.
randomised trialsynthesisobservationalpreclinical

Two things in that picture are the whole argument. The contours bend up to the left, so the same heat costs far more clinical importance when the methods are weaker: a paper cannot buy its way to the top on a dramatic claim. And a well-executed registry cohort lands above a badly executed randomised trial, which is the point of scoring execution rather than the design label.

A trial that demonstrated a benefit patients can feel sits above the same trial run to a decisive null. The margin is deliberately small. A null that rules out an important difference in both directions has answered the question, and it still outranks every observational study and every result that is statistically significant but too small to matter.

The score is deterministic, and every card shows the factors that moved it: the same reasoning the ranking used, not a post-hoc rationalisation. If you want to drive the arithmetic yourself, the scoring page runs this exact code in your browser.

Two things are deliberately absent. Citation counts: everything here is a week or two old, and at that age a citation count measures indexing speed, not importance. And reader-relevance signals like Canadian authorship or subspecialty: they say nothing about how good a paper is, so they feed the model that learns which papers actually get approved, never the score. Either way, heat is a screening tool, not a verdict: cold papers can still be great papers.

Summaries & the name

About the summaries. Summaries are generated by an AI model (Claude) from the paper's full text and reviewed by Pukhraj Gaheer, Medical Student, Queen's University before publishing. They are original paraphrases, not excerpts, and every card links to the actual paper via its DOI. Always read the full paper before citing it, and nothing on this site is medical advice.

The name. Orthopedics, and the talk around it: what gets argued about at rounds and forwarded on a Friday, which is more or less the whole job. Say it out loud and the three letters in the middle give away the rest: ortHOTalk, the week's hot topics. Not a coincidence.

Behind the site

Hi, I'm Pukhraj Gaheer, a second-year medical student at Queen's University. Before medicine I completed a Master's in Health Research Methodology at McMaster with a focus on multi-omics and statistical genetics. If this site seems unusually opinionated about methodology, that degree is why.

My research interests are: the multi-omics of bone healing, and cohort study analyses, the large, patient, unglamorous datasets that need cleaning and analyses.

I've also been programming for the past six years, which is how a reading habit turned into a website. I wanted a weekly, critically appraised scan of the orthopedic literature, couldn't find one, and decided to just build it. Away from the library I'm usually on a golf course or playing guitar, one of which occasionally goes well.

Portrait of Pukhraj Gaheer
Gaheer, P.Queen's Medicine