Radiology Peer Review vs Peer Learning: What Changed
Score-based peer review is giving way to peer learning across American radiology. What each model actually does, why the shift happened, what accreditation still requires, and the case selection problem neither model solves.
By the Radiological.ai team
July 2026 · 9 min read
Worklist
Structured report
DraftRun the assistant to draft this report for review.
Illustrative sample · not a real patient study, not a diagnosis
Decision support for qualified clinicians. Radiological.ai does not provide a diagnosis and is not a substitute for professional judgment.
The short answer: Peer review scores a colleague's earlier interpretation on a scale and produces a record for accreditation. Peer learning drops the score, collects instructive cases including near misses and good calls, and puts them in front of the group to learn from. American radiology has been moving from the first to the second for about a decade: a survey of ACR members found 53 percent of respondents using peer learning against 29 percent who were not. The reason is blunt. Score-based review has never been shown to measure competence or improve performance, and most groups found it produced paperwork and resentment in roughly equal measure. Peer learning is better education, but it does not remove the requirement to document a peer review program for accreditation, and it does not solve the harder problem underneath both models: finding the cases worth talking about.
Updated July 2026.
If you sit on a radiology quality committee, you have had a version of this conversation. Somebody points out that the peer review numbers have looked the same for four years, that the discrepancy rate is suspiciously close to zero, and that nobody has learned anything from the process since it started. Somebody else points out that the accreditation body asks for it, so it continues. That stalemate is what peer learning was invented to break, and it is worth understanding what actually changed and what did not.
What is peer review in radiology?
Peer review is the process in which one radiologist evaluates another radiologist's prior interpretation of an imaging study. In practice it usually happens opportunistically: you are reading a new exam, a prior study is available for comparison, and you record whether you agree with how that prior was read. That requirement for a comparison study is structural, and it quietly shapes everything. Only studies with priors are eligible, so the pool skews toward patients with ongoing imaging and away from the one-off emergency read where a miss is most consequential.
The dominant program in the United States is RADPEER, introduced by the American College of Radiology in 2002 and now used by more than 18,000 radiologists across more than 1,100 groups. Reviewers score the prior interpretation from 1, meaning concurrence, through 4, meaning a discrepancy that should have been made most of the time. The ABR accepted RADPEER as practice quality improvement in 2009, which locked it into maintenance of certification and gave every radiologist a personal reason to keep participating.
The sampling standard most groups work to is five percent of cases, selected at random, which is the documentation rate the Joint Commission has historically expected. Less than two percent of peer review is still done on paper. Everything else runs through software, either RADPEER itself or a module inside a reporting or workflow platform.
What is peer learning in radiology?
Peer learning keeps the case and throws away the score. Instead of grading a colleague, radiologists submit cases they think the group would benefit from seeing: a subtle finding that was called, a near miss, an unusual presentation, a systems failure that had nothing to do with interpretation. Those cases go to a conference, get discussed without attribution, and the themes that emerge feed back into protocols, templates and hanging protocols.
The philosophical shift came out of patient safety work well outside radiology, which has consistently found that blame-based quality systems suppress reporting while non-punitive ones surface more. The Institute of Medicine made the same argument for medicine generally: if you want to learn from error and near misses, you have to build a system where telling the truth is safe.
What is the difference between peer review and peer learning?
| Score-based peer review | Peer learning | |
|---|---|---|
| Primary purpose | Assurance and documentation | Education and improvement |
| Output | A score per case, aggregated per radiologist | A discussed case and a group takeaway |
| Attribution | Tied to the individual radiologist | Usually de-identified in conference |
| Case selection | Random sample, typically five percent, requires a prior | Submitted by colleagues, plus whatever the group can surface |
| Includes good calls | No | Yes, deliberately |
| Cultural effect | Frequently reported as punitive | Higher engagement when run well |
| Accreditation evidence | Direct and easy to produce | Requires deliberate documentation |
The last row is where most of the real-world friction sits, and it is why the majority of groups do not actually choose one model. They run peer learning for the education and keep a lighter documented process for the surveyor.
Why did radiology move away from score-based peer review?
Three reasons, all of them well documented in the literature.
The first is that it does not measure what it claims to measure. Score-based review has not been shown to establish a radiologist's competence, and the statistical reasons are unforgiving. Discrepancies are rare events, five percent samples are small, and the resulting per-radiologist numbers are dominated by noise. Ranking colleagues on them is close to meaningless.
The second is that it changes behavior in the wrong direction. When a score attaches to a named individual and lands in a credentialing file, the rational response is to score generously and to avoid flagging anything ambiguous. Published critiques describe exactly this, along with the resentment and fear it can build in a department. A quality system that people quietly route around is worse than no system, because it produces reassuring data.
The third is that it teaches nothing. A 3 on a case you read eight months ago, delivered with no discussion, does not change how you read next week. The cases that would change practice, the near misses and the instructive saves, have no place to go in a scoring model at all.
Does peer learning satisfy ACR accreditation and the Joint Commission?
This is the question that stops committees, and it deserves a careful answer rather than a comfortable one. Active participation in a peer review program is a requirement for ACR facility accreditation in CT, MR, nuclear medicine, PET, ultrasound, breast ultrasound and breast MRI. You are not obliged to use RADPEER specifically: facilities can be approved with an equivalent program in place. What you cannot do is run an undocumented process and describe it as peer learning when a surveyor asks.
The related trap is OPPE. Ongoing professional practice evaluation is the routine competency monitoring the Joint Commission and CMS expect for credentialed physicians, with focused evaluation, FPPE, as the follow-up when a concern arises or a new hire is being assessed. Peer review data feeds OPPE, but the ACR has been explicit that RADPEER is not designed to serve as a sole OPPE measure. Groups that lean on it as their only competency signal tend to discover the gap during a survey rather than before one. A defensible OPPE program pulls in turnaround times, communication of critical results, participation rates and complications alongside interpretive quality.
How do you choose cases for peer learning?
Here is the problem neither model solves. Peer learning improves what happens to a case once someone finds it. It does nothing about finding it. Voluntary submission is biased toward what individuals happen to notice and remember, and it systematically under-samples the misses nobody caught, which are exactly the cases with the most to teach. Random sampling, meanwhile, is a method for finding average cases.
Most groups end up with three sources: colleague submissions, cases surfaced by clinician complaints or adverse outcomes, and whatever the random sample throws up. All three are late, biased, or thin. Pulling anything better means joining data across the PACS, the RIS and the reporting platform, and anyone who has tried to wire those systems together knows the join is where these projects usually stall, because the case, the report and the outcome live in three databases that were never designed to be queried as one.
There is a newer option worth understanding, because it is genuinely different in kind. Software that reviews studies as they arrive can run across the entire volume rather than a sample, and shortlist the cases where its own flags do not line up with the signed report. Published work gives a realistic picture of what that funnel looks like: in a review of 25,104 chest radiographs, automated comparison surfaced discrepancies in 21.1 percent of cases, of which 0.9 percent were judged clinically relevant on external review and 0.1 percent were confirmed as clinically relevant by the institution's own radiologists. A smaller study of 2,573 CT pulmonary angiograms at a multisite teleradiology practice produced 136 potentially discrepant cases, which radiologists narrowed to 13 confirmed missed findings.
Read those numbers honestly and they say two things at once. The shortlist is noisy and needs human triage, so anyone selling this as automated error detection is overselling it. But a process that reliably delivers a double-digit number of genuinely instructive cases per year, drawn from the whole volume instead of a five percent sample, is a better feed for a peer learning conference than voluntary submission has ever been. That is the argument behind AI-assisted radiology peer review case selection: the software picks candidates, the radiologists decide what they mean.
One structural prerequisite is worth flagging. A discrepancy is only detectable if the report is consistent enough to compare against. Groups running free-text narrative reporting with wide individual variation find this much harder than groups on structured radiology reporting, which is one of the less obvious arguments for template discipline.
How do you run a peer learning conference people actually attend?
The failure modes are predictable. A conference that turns into a slow-motion blame session dies within a quarter. One that only ever shows fascinating zebras is entertaining and changes nothing. One with no follow-through teaches the group that submitting a case leads nowhere.
What tends to work: de-identify by default and make that a stated rule rather than a courtesy; include good calls and saves, not only errors, because a conference that is exclusively about failure trains people to submit nothing; assign an owner and a date to every action item that comes out of a case, and read the previous list at the start of the next session; and keep the scored process, if you still need one, administratively separate from the learning process so nobody has to wonder which one they are sitting in.
Distributed groups have a harder version of this problem. When reads are spread across contractors, time zones and overnight coverage, the conference is often the only moment the group is in the same room, and the cases that most need discussing are frequently the ones read by someone who is asleep. Groups in that position should think about peer learning as a governance question rather than a scheduling one, which is the same theme running through our teleradiology software page.
What should a group decide right now?
If you are still running scored review only, the move is not to tear it out. Keep whatever produces your accreditation evidence, start a peer learning conference alongside it, and let the scored process shrink to the minimum the surveyor needs. Groups that tried to replace one with the other in a single step generally spent a year arguing about documentation instead of discussing cases.
If you already run peer learning and attendance is drifting, the problem is almost never the format. It is the case supply. Fix where cases come from before redesigning the meeting. If you are evaluating software that claims to help with either, our radiology AI vendor evaluation questions cover what to ask, and the best AI radiology software roundup is an honest read on which vendors do what in this category.
The thing worth holding onto is that peer review was never really about scores, and peer learning is not really about conferences. Both are attempts to answer one question: what is this group getting wrong that it does not know it is getting wrong? Every serious improvement to either model has come from getting better at surfacing those cases, and that is still where the remaining upside is.
See Radiological.ai read a study
The assistant flags suspected findings for review, prioritizes the worklist so urgent studies surface first, and drafts the structured report into your template. You review, edit and sign every study.