Appraisal methodology · CMO review

How Dose-Response appraises research

The exact methodology behind our critical appraisals, documented from the live prompts, output fields and editorial rules. Prepared for Dr Phil Cox's review as Chief Medical Officer.

Two appraisal agents, one standard. Both are governed by a single written constitution (the six HARD RULES below) and both run a draft → senior-review two-pass before any output is trusted. The difference is operational: the weekly agent produces the curated bulletin (human-approved before publishing); the on-demand agent appraises any paper a member chooses and returns it to them directly.
1 · Weekly appraiser 2 · On-demand appraiser The six rules Human oversight
Agent 1 · batch, curated

The Weekly Appraiser — "The Five"

The engine behind the weekly bulletin. Harvest candidate papers → appraise → senior-review → the strongest few are composed into the issue that goes to web and the mailing list.

Runs
Weekly (Mon 06:00), in batch
Model
Claude Sonnet 5 (via the Claude Max subscription, no metered cost)
Input
Title, journal, date, authors, abstract + citation metrics. Appraised from the abstract and metadata.
Passes
Appraiser draft → Reviewer edit → QA monitor → human Editor Desk

Pass 1 — the Appraiser's brief

"Your job is APPRAISAL, not summary. A busy clinician reads this between patients. Be rigorous but plain-spoken. Always separate statistical significance from clinical significance... Be willing to say a paper is underpowered, overhyped, or shouldn't change practice. Never invent findings, numbers, or citations; if the abstract doesn't support a claim, say the evidence is unclear. State uncertainty rather than smoothing it over."

What it must produce for every paper

FieldWhat it captures
the_paperOne line: study design, n, population.
what_they_foundThe headline result with the actual numbers (effect sizes, CIs, p-values, absolute differences). No spin.
study_designExact design and where it sits on the evidence hierarchy (narrative vs systematic review/meta-analysis; RCT vs cohort vs case series; prospective vs retrospective), n, population/setting.
methodologyHow it was done and how trustworthy; specific, named sources of bias (selection, performance, detection, attrition, publication, funding/COI). For reviews: search strategy, quality grading (GRADE/risk-of-bias tool), pooled vs narrative synthesis. For primary studies: randomisation, allocation concealment, blinding, power, follow-up, confounder adjustment.
appraisalClinically meaningful or only statistically significant? Magnitude worth acting on? Does the design support the causal claim, or only an association?
the_gapThe single most important limitation — what you'd need before trusting it in clinic.
landmark_contextFoundational work it builds on, extends, or contradicts. Never a fabricated citation.
practice_implicationDoes it change what a clinician does on Monday, or not?
verdictExactly one of the five verdicts (below).
confidenceHigh · Moderate · Low.
quality_scoreInteger 0–100: overall methodological quality & clinical usefulness.
one_line_takeawayThe "so what", ≤160 characters.
categories3–7 controlled-vocabulary tags (fixed taxonomy, no invented slugs).
reflection_prompts2–3 CPD questions; at least one asks the reader to picture a real patient.
linkedin_hookAn opening line for a LinkedIn post (specific, not clickbait). Unique to the weekly agent.

Verdict scale: Practice-changingWorth knowingWatch this spaceOverhypedNot yet

Pass 2 — the Reviewer (senior editor)

Every draft is handed to a second agent whose job is to challenge it. It edits in place, logs every change, and emails an audit report of what it changed, flagged, or rejected.

"You are a DUAL-REGISTERED Physiotherapist (HCPC/CSP) and UKSCA-accredited Strength & Conditioning coach with 24 years across elite sport and MSK clinic... You correct verdicts that overstate or understate the evidence, add missing caveats, make every claim match what the study actually shows, and make the practice implication something a clinician could act on Monday. You are decisive: approve, revise with concrete edits, or reject. You never rubber-stamp and never invent findings."

The Reviewer can override the verdict, confidence and quality score; adds an s_and_c_relevance rating (so a paper is judged for physios and S&C coaches); adds a clinical_application paragraph; and may, sparingly, name a specific validated in-clinic test (from the Benchmark testing battery) to make a finding measurable. It sets each appraisal to approved, edited or rejected.

Agent 2 · on-demand, member-triggered

The On-Demand Appraiser — "Appraise this paper"

When a member pastes a DOI, PMID, title or full text and asks "appraise this," this agent produces an appraisal on demand, to the same standard as the bulletin.

Runs
At the edge, on request (POST /api/appraise)
Model
Claude Sonnet 4.5, via the metered Anthropic API (kept separate so member demand never competes with the bulletin)
Input
DOI / PMID / title (abstract fetched free from Europe PMC), or pasted / PDF full text (which stays private and is never cached)
Passes
Appraiser draft → inline Reviewer edit
Efficiency
A paper appraised once is cached and reused for every future user — instant, free, no quota used

The Appraiser's brief

"You write to the same standard as our weekly reviewed bulletin... Your job is APPRAISAL then APPLICATION, not summary: separate statistical from clinical significance, name the effect size and whether it actually matters, be decisive and willing to say a paper is underpowered or overhyped, and never invent numbers or citations. Then translate it into what a practitioner does — for BOTH the physiotherapy clinic AND the S&C / performance floor."

Fields — same core as the weekly agent, plus one

It produces the same beats (the_paper, what_they_found, study_design, methodology, appraisal, the_gap, landmark_context, practice_implication, verdict, confidence, quality_score, one_line_takeaway, categories, reflection_prompts), with two differences:

FieldDetail
clinical_applicationMandatory here. 3–5 sentences, practitioner-to-practitioner: who it applies to, what changes in real practice, across both the physio and S&C lenses. Surgical, drug and diagnostic papers are not exempt; never stop at "refer on"; any %, score, imaging finding or cut-off must come with how the clinician ascertains it.
linkedin_hookDropped — not needed for a member's own appraisal.

Output is produced through a forced structured tool, so the field structure is always valid. A second inline Reviewer pass (ported from the weekly engine) then improves the draft; if it fails, the solid draft is kept.

The constitution

The six HARD RULES

Both agents are held to these, and an automated monitor checks live output against them.

  1. Dual-lens application — every appraisal must say what changes in real practice, for physios and S&C coaches. Never just "refer on."
  2. Ascertainment — no naked thresholds. Any %, score, imaging finding or cut-off must come with how the clinician establishes it (or a note that it needs referral/imaging).
  3. Evidence honesty — a narrative review is never dressed up as graded evidence; statistical significance is never sold as clinical significance.
  4. No fabrication — never invent findings, numbers, or citations.
  5. Voice — UK spelling, plain sentences, no em-dashes.
  6. Verdict discipline — one of the five verdicts, and it must be justified by the appraisal.
Governance

Human oversight

QA monitor

Audits recent output against Rules 1–6 — a deterministic check (missing fields, invalid verdict, banned punctuation) plus a second Claude pass scoring each rule pass/flag. It flags drift; it never silently changes anything.

Editor Desk (human-in-the-loop)

For the weekly bulletin, a person loads the reviewed appraisals, edits any card, assigns bulletin slots (lead / the five / exclude), drops weak papers, and clicks "Approve & compose issue." Only the approved issue publishes.

For your review, Phil: on-demand appraisals go straight to the member with no Editor Desk in the loop. The safeguards carrying quality there are the forced structured output, the two-pass review, the ascertainment rule, and the QA monitor sampling live output. Worth your view on whether on-demand appraisals of higher-stakes (surgical / pharmacological / diagnostic) papers warrant an additional guardrail.