
Algorithms now sit at the front door of care, telling patients which surgeon is “good” long before a referral reaches a clinic; whether that guidance is trustworthy depends less on data volume than on validation, risk adjustment, and whether clinicians can see and challenge how they were judged.
At a Glance
- Garner says it scores individual doctors using more than 60 billion claims covering 320 million patients, refreshed monthly and compared against local peers.
- Its model blends clinical outcomes and total cost of care with patient reviews and distance, using more than 500 specialty metrics across 80-plus fields.
- The company discloses high-level risk adjustment and case exclusions but withholds detailed formulas and weights, limiting outside replication.
- Physicians currently cannot view their own scores, fueling concerns about opacity and the limits of claims-only analytics.
What Garner is attempting: an outcomes-and-cost signal for individual physicians
Garner’s proposition is straightforward and ambitious: analyze an enormous corpus of insurance claims to estimate which doctors deliver better outcomes at lower total cost than their local peers, then steer patients accordingly. The company says it has assembled more than 60 billion medical claims representing over 320 million patients and transforms that data through more than 500 proprietary clinical metrics tailored to 80-plus specialties, refreshed monthly to stay current as claims mature. The output, it argues, is not popularity or bedside-manner scorekeeping; it is a performance signal rooted in outcomes and total cost of care, moderated by patient reviews and distance so recommendations reflect quality, value, and practicality for the user.
Mechanically, this is a claims-based quality engine. Claims capture encounters, diagnoses, procedures, and subsequent utilization—readmissions, reoperations, complications coded via follow-up services—at scale. Garner frames its comparisons as local and specialty-specific to reduce the confounding that comes from national, one-size-fits-all rankings. The company also states that it excludes the most complex cases when risk adjustment is unreliable, then adjusts remaining cases for comorbidities and demographics before attributing performance to an individual clinician. Only doctors who outperform local peers on both quality and total cost of care, with a positive review track record, are recommended.
Why claims, why proprietary metrics, and what that buys you
Claims are inescapably imperfect but uniquely comprehensive. They standardize data across payers and regions, allowing models to observe longitudinal episodes of care and downstream consequences. This is why health plans, employer coalitions, and several analytics firms have built measurement systems on claims foundations: even if claims miss clinical nuance, they provide a broad, auditable ledger of what happened next—who returned, who required revision surgery, who developed complications detectable in subsequent billing. Garner leans into that advantage and scales it, arguing that a bottom-up, doctor-level view is feasible with enough data density and careful risk adjustment. Its white-paper materials emphasize outcomes and total cost, a pairing that acknowledges what purchasers actually manage: both quality and spend.
Scope and cadence matter. A system that covers hundreds of millions of patients and refreshes monthly can spot patterns too faint for hospital report cards or consumer review sites. Local peer benchmarking also fits clinical reality; surgeons inhabit specific referral spheres, resource environments, and patient populations. Comparing a spine surgeon in Phoenix to peers in Phoenix is sounder than comparing her to a national composite where practice patterns diverge wildly.
Where the model’s credibility is earned—and where it is not yet
The strongest parts of Garner’s case are its scale, specialty-specific design, and explicit attention to risk. It says it excludes cases it cannot adjust reliably, risk-adjusts the remainder for comorbidities and demographics, and requires superiority on both outcome and total cost dimensions. Those are the right instincts. But instincts are not validation. The company’s public materials do not disclose weighting schemes, exclusion thresholds, model coefficients, calibration checks, or error rates. Without that, outside experts cannot reproduce a score, estimate false positives and false negatives, or gauge how much “quality” versus “cost” drives an individual doctor’s recommendation. Large data alone does not immunize a model from coding variation, referral bias, or misattribution; it only gives you statistical power to be wrong with confidence if the design is flawed.
Transparency for physicians is the other fault line. A surgeon writing in Reason described discovering he was being favorably routed patients by a company he had never engaged with, while being unable to see his own score or the method behind it. That is not an isolated anxiety; Garner’s own pages acknowledge that older systems have been criticized as black boxes that penalize doctors who take on the hardest cases. Garner asserts it avoids those traps, yet it has not provided the tools for the judged to interrogate their judgment. A single Trustpilot comment won’t carry methodological weight, but the theme—opacity around criteria—aligns with the constraints Garner itself sets by keeping formulas proprietary.
How risk adjustment and case exclusion should work in surgeon ratings
Two design choices decide whether claims-based, surgeon-level ratings add signal or amplify noise: attribution and risk adjustment. Attribution must link outcomes to the surgeon with clinical plausibility (index procedure ownership, episode windows, and shared-care logic that assigns complications appropriately). Risk adjustment must reflect preoperative risk, not postoperative sequelae, and it must avoid “adjusting away” true quality differences. Garner says it excludes the most complex cases it cannot adjust reliably; that can improve fairness for high-acuity surgeons but also narrows the applicability of the score for exactly the patients who most need guidance. The right compromise is transparent exclusion criteria, specialty-specific variables (e.g., frailty for orthopedics, tumor stage proxies for oncology), and published calibration that shows how well predicted risk aligns with observed outcomes in and out of sample.
Blending cost into the same recommendation layer is defensible so long as the two axes are separable and visible. Purchasers care about total cost of care; patients care about outcomes first. A composite that hides weightings invites confusion: is this surgeon “best” because revision rates are lower, or because post-acute spend is leaner? Garner’s patient-facing layer says a recommended doctor beats local peers on both, plus has positive reviews, but the lack of exposed weights leaves readers inferring, not knowing, which lever dominated a given recommendation.
What would settle the debate: concrete validation and fair process
This dispute does not require philosophical resolution; it requires evidence and process. Three steps would move the conversation from assertion to credibility. First, a peer-reviewed validation comparing Garner’s recommendations to hard outcomes—mortality where relevant, complications, readmissions, reoperations, infections, and patient-reported outcomes—across several high-volume surgical specialties. Second, a technical appendix or redacted white paper disclosing metric lists, inclusion/exclusion rules, risk variables, and calibration plots, sufficient for independent statisticians to critique, even if exact weights remain proprietary. Third, a physician portal enabling secure score review, case mix summaries, and an appeals workflow, with turnaround SLAs and audit logs. These are common in payer programs and directly address the fairness and feedback loop clinicians expect.
If Garner can show that its “better outcomes, lower cost” signal persists across specialties and markets after rigorous risk adjustment—and give surgeons a professional path to contest misattribution—it will have earned the authority its distribution already implies. If not, the black-box narrative will outrun the merits, and patients will conflate “best value” with “best surgeon” in ways that neither clinicians nor purchasers intend.
Bottom line for patients, clinicians, and purchasers
For patients, a claims-grounded navigator can be a useful starting point, not a final word. Use recommendations to frame questions—case volumes, complication rates in your risk profile, and how the surgeon manages revisions—rather than to substitute for them. For clinicians, the lack of score visibility is a genuine problem; ask your health systems and benefit partners whether they use Garner and what recourse exists to correct attributions. For purchasers, the decision is governance: do not outsource trust blindly. Require external validation, transparency on risk adjustment and exclusions, and an appeals process before you let any model become the default gatekeeper to specialty care.
Sources:
reason.com, garnerhealth.com, app.getgarner.com, garneremployerguide.com



