How Dose-Response appraises research
The exact methodology behind our critical appraisals, documented from the live prompts, output fields and editorial rules. Prepared for Dr Phil Cox's review as Chief Medical Officer.
The Weekly Appraiser — "The Five"
The engine behind the weekly bulletin. Harvest candidate papers → appraise → senior-review → the strongest few are composed into the issue that goes to web and the mailing list.
Pass 1 — the Appraiser's brief
"Your job is APPRAISAL, not summary. A busy clinician reads this between patients. Be rigorous but plain-spoken. Always separate statistical significance from clinical significance... Be willing to say a paper is underpowered, overhyped, or shouldn't change practice. Never invent findings, numbers, or citations; if the abstract doesn't support a claim, say the evidence is unclear. State uncertainty rather than smoothing it over."
What it must produce for every paper
| Field | What it captures |
|---|---|
| the_paper | One line: study design, n, population. |
| what_they_found | The headline result with the actual numbers (effect sizes, CIs, p-values, absolute differences). No spin. |
| study_design | Exact design and where it sits on the evidence hierarchy (narrative vs systematic review/meta-analysis; RCT vs cohort vs case series; prospective vs retrospective), n, population/setting. |
| methodology | How it was done and how trustworthy; specific, named sources of bias (selection, performance, detection, attrition, publication, funding/COI). For reviews: search strategy, quality grading (GRADE/risk-of-bias tool), pooled vs narrative synthesis. For primary studies: randomisation, allocation concealment, blinding, power, follow-up, confounder adjustment. |
| appraisal | Clinically meaningful or only statistically significant? Magnitude worth acting on? Does the design support the causal claim, or only an association? |
| the_gap | The single most important limitation — what you'd need before trusting it in clinic. |
| landmark_context | Foundational work it builds on, extends, or contradicts. Never a fabricated citation. |
| practice_implication | Does it change what a clinician does on Monday, or not? |
| verdict | Exactly one of the five verdicts (below). |
| confidence | High · Moderate · Low. |
| quality_score | Integer 0–100: overall methodological quality & clinical usefulness. |
| one_line_takeaway | The "so what", ≤160 characters. |
| categories | 3–7 controlled-vocabulary tags (fixed taxonomy, no invented slugs). |
| reflection_prompts | 2–3 CPD questions; at least one asks the reader to picture a real patient. |
| linkedin_hook | An opening line for a LinkedIn post (specific, not clickbait). Unique to the weekly agent. |
Verdict scale: Practice-changingWorth knowingWatch this spaceOverhypedNot yet
Pass 2 — the Reviewer (senior editor)
Every draft is handed to a second agent whose job is to challenge it. It edits in place, logs every change, and emails an audit report of what it changed, flagged, or rejected.
"You are a DUAL-REGISTERED Physiotherapist (HCPC/CSP) and UKSCA-accredited Strength & Conditioning coach with 24 years across elite sport and MSK clinic... You correct verdicts that overstate or understate the evidence, add missing caveats, make every claim match what the study actually shows, and make the practice implication something a clinician could act on Monday. You are decisive: approve, revise with concrete edits, or reject. You never rubber-stamp and never invent findings."
The Reviewer can override the verdict, confidence and quality score; adds an s_and_c_relevance rating (so a paper is judged for physios and S&C coaches); adds a clinical_application paragraph; and may, sparingly, name a specific validated in-clinic test (from the Benchmark testing battery) to make a finding measurable. It sets each appraisal to approved, edited or rejected.
The On-Demand Appraiser — "Appraise this paper"
When a member pastes a DOI, PMID, title or full text and asks "appraise this," this agent produces an appraisal on demand, to the same standard as the bulletin.
The Appraiser's brief
"You write to the same standard as our weekly reviewed bulletin... Your job is APPRAISAL then APPLICATION, not summary: separate statistical from clinical significance, name the effect size and whether it actually matters, be decisive and willing to say a paper is underpowered or overhyped, and never invent numbers or citations. Then translate it into what a practitioner does — for BOTH the physiotherapy clinic AND the S&C / performance floor."
Fields — same core as the weekly agent, plus one
It produces the same beats (the_paper, what_they_found, study_design, methodology, appraisal, the_gap, landmark_context, practice_implication, verdict, confidence, quality_score, one_line_takeaway, categories, reflection_prompts), with two differences:
| Field | Detail |
|---|---|
| clinical_application | Mandatory here. 3–5 sentences, practitioner-to-practitioner: who it applies to, what changes in real practice, across both the physio and S&C lenses. Surgical, drug and diagnostic papers are not exempt; never stop at "refer on"; any %, score, imaging finding or cut-off must come with how the clinician ascertains it. |
| linkedin_hook | Dropped — not needed for a member's own appraisal. |
Output is produced through a forced structured tool, so the field structure is always valid. A second inline Reviewer pass (ported from the weekly engine) then improves the draft; if it fails, the solid draft is kept.
The six HARD RULES
Both agents are held to these, and an automated monitor checks live output against them.
- Dual-lens application — every appraisal must say what changes in real practice, for physios and S&C coaches. Never just "refer on."
- Ascertainment — no naked thresholds. Any %, score, imaging finding or cut-off must come with how the clinician establishes it (or a note that it needs referral/imaging).
- Evidence honesty — a narrative review is never dressed up as graded evidence; statistical significance is never sold as clinical significance.
- No fabrication — never invent findings, numbers, or citations.
- Voice — UK spelling, plain sentences, no em-dashes.
- Verdict discipline — one of the five verdicts, and it must be justified by the appraisal.
Human oversight
QA monitor
Audits recent output against Rules 1–6 — a deterministic check (missing fields, invalid verdict, banned punctuation) plus a second Claude pass scoring each rule pass/flag. It flags drift; it never silently changes anything.
Editor Desk (human-in-the-loop)
For the weekly bulletin, a person loads the reviewed appraisals, edits any card, assigns bulletin slots (lead / the five / exclude), drops weak papers, and clicks "Approve & compose issue." Only the approved issue publishes.