Category: Decision-support

  • The Machines Caught Up. On the Easy Calls.


    Two years ago I sat down with two hundred pairs of simulated patients and did something incredibly boring: I decided, one pair at a time, which of two patients would be seen first. No hedging, no “it depends,” no ties. Just 1 or 2, two hundred times. Then I asked three frontier models to do the same thing, and measured how often they agreed with me. I wanted to see what an LLM would decide, when faced with a seemingly trivial decision under scarcity (the scarce resource being a time slot in a clinic).

    They mostly didn’t. GPT-4 and Claude Sonnet 3.5 landed at a Cohen’s κ of 0.17 against my choices. Gemini Ultra managed 0.23. On the pairs I had flagged in advance as hard the ones where I actually had to think, the models scored κ between 0.01 and 0.11. Not far from chance. These results and the difficulty I had in aligning the models to my preferences led to the Human Values Project (see hvp.global). I did write up the study as a pre-print on arxiv: https://arxiv.org/abs/2409.18995

    I re-ran the whole thing this month. Same two hundred pairs, same gold standard (importantly: not made public), same task. Nine models instead of three, five independent runs of each instead of three. This time The result is both unsurprising and oddly surprising.

    What happened

    Everything got better. But where?

    Triage concordance, 2024 to 2026

    Overall concordance roughly tripled. The 2026 frontier (GPT-5.2 at κ 0.67, Gemini 2.5 Pro at 0.64, Grok 4.5 at 0.63, Claude Opus 4.8 at 0.62) sits where two years ago there was nothing but noise. By the conventional rubric these models have gone from “slight” agreement with an experienced clinician (forgive the presumption) to “moderate-to-substantial.” On a task nobody trained them for, with no correct answer to memorize. My gold standard has never been published, so they could not have looked up my answers.

    But the middle panel is revealing. On the easy pairs the models are at κ 0.92 to 0.94. That is near-perfect agreement. The right panel barely moved.

    Easy versus hard concordance by model

    Every model, without exception, is dramatically more concordant on the easy pairs than the hard ones. GPT-5.2: 0.92 easy, 0.37 hard. Grok 4.5: 0.94 easy, 0.28 hard. Claude Opus 4.8: 0.93 easy, 0.26 hard. The gap is telling.

    Why I do not call that a failure

    There is a temptation to read the right-hand section of the panel as the models’ report card, and to conclude that they remain unfit for the decisions that matter. But recall: the gold standard on those hard pairs is one (trained in the 80s and 90s) physician’s opinion on a Tuesday afternoon.

    On the easy pairs I was not expressing a value. I was reporting something close to a fact. It is the sort of judgment on which essentially every clinician would agree. That is precisely why κ climbs to 0.94: the models are converging on a consensus that genuinely exists. Concordance there is a sanity check, and they now pass it. That is real progress and I do not want to undersell it. Two years ago they were failing a test that should be easy.

    On the hard pairs I was expressing a preference. The 71-year-old with Alzheimer’s disease or the 43-year-old with breast cancer? I made a call. I could defend it. So could a thoughtful colleague who called it the other way. There is no reason to believe that my ordering is the right one, and a great deal of reason (five decades of decision science, and the sometimes heated experience of any tumor board) to believe there is no single right one to be had. When a model disagrees with me on those pairs it is not necessarily wrong. It may hold different values, or the same values with a different weighting.

    So κ = 0.30 on the hard pairs is not a measurement of model error. It is a measurement of disagreement with one person, on exactly the class of decisions where disagreement is the expected condition.

    Reproducibility is seems solved, if you pay for it

    The 2024 result that bothered me most was not the low concordance. It was that the models disagreed with themselves. Ask the same model the same question three times and you could get three different orderings, sometimes at negative κ with each other. I might be wary of a doctor with that level of inconsistency.

    Run-to-run consistency by model

    That problem is largely gone at the frontier. Across five runs, GPT-5.2 reproduces itself at κ 0.94, Claude Opus 4.8 at 0.92, Grok 4.5 at 0.89, Gemini 2.5 Pro at 0.85. These systems now have stable, repeatable preferences. If you ask twice, you get the same answer.

    I want to be careful about what that buys us. Consistency is a precondition for accountability. It is challenging to regulate, or negotiate with a system whose answers vary widely for the same decision. But a consistent triage policy that is systematically wrong for your patients is a worse problem than a noisy one, not a better one, because it will be wrong the same way every single time, for every patient, forever. We have traded variance for something much closer to institutional policy. That deserves the scrutiny we give policy (see my discussion on clinical guidelines for more on that https://ai.nejm.org/doi/full/10.1056/AIe2600680).

    Note also who is not in that club. The open-weight models reproduce themselves at κ 0.28 (Kimi K2-Thinking), 0.24 (Llama 3.3 70B), and 0.14 (GLM-4.7). On this task they remain closer than I’d like to weathervanes

    The open-weight models are not close

    This is the finding I least wanted and I found most surprising because I use these open-weight extensively when I want to run them on my own hardware (old 4 year old macs with lots of memory from when memmory was much cheaper). I also think it is important for the development of a healthy ecosystem. Llama 3.3 70B scored κ 0.03 overall. 0.13 on the easy pairs, and negative on the hard ones. GLM-4.7 reached 0.24, Kimi K2-Thinking 0.35. The best open-weight model in this set is roughly where the closed frontier was in 2024.

    For any health system contemplating a self-hosted model on privacy or cost grounds (a decision that may become increasingly compelling) this is a number to sit with. The open option is not merely a little behind on this task. It is a different animal. And the variance among open models is enormous, so “open-weight” is not a category that predicts anything. Test the specific model, on your own decisions.

    What alignment did, and what “alignment” means here

    Alignment is now an overloaded term. Let me be precise what I tested back in 2024.

    The alignment technique I re-ran is one specific, humble thing: in-context example alignment. The model is shown roughly sixty previously annotated patient pairs. Specifically a different set from the two hundred being scored. Each one is labeled with the choice an expert clinician made, and instructed to generalize from them rather than pattern-match. That is it. No fine-tuning, no RLHF, no system-prompt constitution. Few-shot exemplars in the context window.

    It is also only one of several probes from the original study. That paper also tested alignment by supplying population-level inequality exemplars, by forcing an abstract generalization (maximize quality-adjusted life years), and by perturbing the gold standard itself to see whether model rankings survived. Those were the conditions that produced the most unsettling results, For example the QALY instruction made every model worse. I have not re-run those other alignment approacjhes yet. Everything below concerns the gentlest, most favorable form of alignment I know how to apply.

    Where in-context alignment acts

    Alignment barely touched the easy pairs, for the obvious reason that there was nowhere left to go. Where it acted, it acted almost entirely on the hard ones: Grok 4.5 from 0.28 to 0.46, GPT-4o from 0.23 to 0.36, Kimi K2-Thinking from 0.12 to 0.35, Claude Sonnet 5 from 0.17 to 0.26.

    That is a genuinely encouraging pattern. The contested decisions are the ones a health system would actually want to shape, and they are the ones showing movement. Sixty examples of how one clinician weighs a trade-off measurably shifted how several models weigh it. I see that as positive evidence of personalization. You might see this as spineless decision-making on the part of the LLM.

    But look at the arrows pointing the wrong way. Claude Opus 4.8 went down on the hard pairs, 0.26 to 0.22, and down overall. GLM-4.7 fell from 0.12 to 0.06 and lost most of its already-poor self-consistency. GPT-5.2 and Gemini 2.5 Pro did not budge. The same sixty examples, in the same words, helped some models, did nothing to others, and actively degraded two. Whatever “aligning a model to our institution’s preferences” is going to mean operationally, it is not a switch, and it does not generalize across vendors. It has to be measured per model, which is the entire reason I proposed the Alignment Compliance Index in the first place.

    Grok 4.5 is the one to watch

    The new entrant is the surprise of this run. Grok 4.5 arrives at κ 0.63 unaligned . That is third best, statistically indistinguishable from Gemini and Opus. κ is 0.94 on the easy pairs and run-to-run consistency of 0.89. Respectable, if that were all.

    It is not all. Grok is the most compliant of the frontier models. Its concordance gain from in-context alignment, ΔC = +0.08, is the largest of any model tested, and it lands at κ 0.71 after alignment. That is the highest concordance any model achieved in any condition in this study, 2024 or 2026. On the hard pairs it improved more than any frontier model, 0.28 to 0.46.

    However: Grok’s overall ACI is a modest +0.03, because its run-to-run consistency slipped slightly as it moved, 0.89 to 0.84. It became more like me and slightly less like itself. Kimi K2-Thinking posts a much larger ACI of +0.40, but from a poor baseline and with the widest error bars in the set.

    Being both highly concordant and highly steerable is the combination that matters, in my opinion, for clinical deployment, and right now Grok has it. It is also the cheaper than the Opus and Gemini models tested.

    What this does not show

    One clinician. Two hundred simulated patients, not real ones. A single alignment technique, and the mildest of the several in the original study. Five runs is enough to stabilize the frontier models and not quite enough for Kimi and GLM, whose standard deviations remain large. And κ against my preferences is a measure of agreement with me, not of correctness. I’ve made the point twice because it is the one most likely to be dropped in the retelling.

    Where this goes

    The 2024 version of this experiment took weeks of copying and pasting into chat windows. The 2026 version is a script: nine models, two conditions, five runs, two hundred pairs, every response cached, the whole thing reproducible on a laptop-independent machine overnight for a few dollars. The measurement is no longer the bottleneck. The gold standard is. Which is exactly the argument for the Human Values Project: if we can now measure alignment compliance cheaply, continuously, and across every model a clinician might touch, then the binding constraint is no longer the instrument but the reference. Whose values we are measuring against?

    One clinician’s answer key was a useful way to prove the instrument works. It is a terrible way to decide whose values get installed in the machines our patients will meet. Our more recent publications are using larger samples (e.g. https://arxiv.org/html/2605.18738v1).

    For this reference to be a distribution we need thousands of clinicians and, just as importantly, thousands of patients, making these sam, categorical choices, with their disagreements preserved rather than averaged away. Making human values explicit is harder than measuring them. If you want to help go to hvp.global and sign up.


    Version note: figures and numbers generated 22 July 2026 from a 9-model × 2-condition × 5-run sweep over the same 200 patient pairs used in the 2024 study. Concordance is Cohen’s κ against a single clinician’s prioritizations; ACI = ΔC + λ·ΔP with λ = 1. The 2024 values are from Table 1 of the original paper; the 2024→2026 comparison is tier-matched within each vendor (GPT-4→GPT-5.2, Sonnet 3.5→Sonnet 5, Gemini Ultra→Gemini 2.5 Pro).

  • MODW4US

    Make Our Data Work for Us

    Why patients—and clinicians—need a Human Values Project for AI in healthcare

    Why call for “making our data work for us” in healthcare?

    Because our data already works—for many parties other than us.

    Clinical data is essential for diagnosis and treatment, but it is also routinely used to shape wait times, coverage decisions, and access to services in ways patients rarely see and cannot easily contest. Insurance status documented in hospital records has been associated with longer waits for care even when clinical urgency is comparable. Medicare Advantage insurers have been accused of using algorithmic predictions to deny access to rehabilitation services that clinicians believed were medically appropriate.

    This asymmetry is not new. Medicine has always involved unequal access to expertise and power. But it was quantitatively amplified by electronic health records—and it is now being scaled again by AI systems trained on those records.

    At the same time, something paradoxical is happening.

    As primary care becomes harder to access, visits shorter, and care more fragmented, patients are increasingly turning, cautiously but steadily, to AI chatbots to interpret symptoms, diagnoses, and treatment plans. Nearly half of Americans now report using AI tools for health-related questions. These systems are imperfect and sometimes wrong in consequential ways. But for many people, the alternative is not a thoughtful clinician with time to spare. It is no timely expert input at all.

    That tension—between risk and access, empowerment and manipulation—is where AI in healthcare now sits. And to be perfectly clear, I personally use AI chatbots all the time for second opinions, or extended explanation, about the care of family members and pets (!). It makes me a better patient and doctor.


    This post grows directly out of my recent Boston Globe op-ed, “Who is your AI health advisor really serving?”, which explores how the same AI systems that increasingly advise patients and clinicians can be quietly shaped by the incentives of hospitals, insurers, and other powerful stakeholders. The op-ed focuses on what is at stake at a societal level as AI becomes embedded in care. What follows here is more granular: how these alignment pressures actually enter clinical advice, why even small downstream choices can have outsized effects, and what patients and clinicians can do—today—to recognize, test, and ultimately help govern the values encoded in medical AI.
    [Link to Globe op-ed ]


    Where alignment actually enters—and why it matters

    In  my Boston Globe op-ed, I argued that as AI becomes embedded in healthcare, powerful incentives will shape how it behaves. Hospital systems, insurers, governments, and technology vendors all have understandable goals. But those goals are not identical to the goals of patients. And once AI systems are tuned—quietly—to serve one set of interests, they can make entire patterns of care feel inevitable and unchangeable.

    This is not a hypothetical concern.

    In recent work with colleagues, we showed just how sensitive clinical AI can be to alignment choices that never appear in public-facing documentation. We posed a narrowly defined but high-stakes clinical question involving a child with borderline growth hormone deficiency. When the same large language model was prompted to reason as a pediatric endocrinologist, it recommended growth hormone treatment (daily injections for years). When prompted to reason as a payer, it recommended denial and watchful waiting (which might be the better recommendation for non-growth-deficient children).

    Nothing about the medical facts changed. What changed was the frame—a few words in the system prompt.

    Scale that phenomenon up. A subtle alignment choice, made once by a hospital system, insurer, or vendor and then deployed across thousands of encounters, can shift billions of dollars in expenditure and materially alter health outcomes for large populations. These are not “AI company values.” They are downstream alignments imposed by healthcare stakeholders, often invisibly, and often without public scrutiny.


    Why experimenting yourself actually matters

    This is the context for the examples below.

    The point of trying the same clinical prompts across multiple AI models is not to find the “best” one. It is to calibrate yourself. Different models have strikingly different clinical styles—some intervene early, some delay; some emphasize risk, others cost or guideline conformity—even when the scenario is tightly specified and the stakes are high.

    By seeing these differences firsthand, two things happen:

    1. You become less vulnerable to false certainty.
      Each model speaks confidently. Seeing them disagree—systematically—teaches you to discount tone and attend to reasoning.
    2. You partially immunize yourself against hidden alignment.
      Using more than one model gives you diversity of perspective, much like seeking multiple human second opinions. It reduces the chance that you are unknowingly absorbing the preferences of a single, quietly aligned system.

    This kind of experimentation is not a substitute for clinical care. It is a way of learning how AI behaves before it is intermediated by institutions whose incentives may not be fully aligned with yours.


    Using AI with your own data

    To make this concrete, I took publicly available (and plausibly fictional) discharge summaries and clinical notes and posed a set of practical prompts (see link here) to several widely used AI models. The goal was not to evaluate accuracy exhaustively, but to expose differences in clinical reasoning and emphasis.

    Some prompts you might try with your own records (see the bottom of this post about getting your own records):

    • “Summarize this hospitalization in plain language. What happened, and what should I do next?”
    • “Based on this record, what questions should I ask my doctor at my follow-up visit?”
    • “Are there potential drug interactions among these medications?”
    • “Explain what these lab values mean and flag any that are abnormal.”
    • “Is there an insurance plan that would be more cost effective for me, given my medical history?”
    • “What preventive care or screenings might I be due for given my age and history?”
    • “Help me understand this diagnosis—what does it mean, and what are typical treatment approaches?”

    Across models, the differences are obvious. Some are conservative to a fault. Others are aggressive. Some emphasize uncertainty; others project confidence where none is warranted. These differences are not noise—they are signatures of underlying alignment.

    Seeing that is the first step toward using AI responsibly rather than passively.


    The risks are real—on both sides

    AI systems fail in unpredictable ways. They hallucinate. They misread context. They may miss urgency or overstate certainty. A plausible answer can still be wrong in ways a non-expert cannot detect.

    But here is the uncomfortable comparison we need to make.

    We should not measure AI advice against an idealized healthcare system with unlimited access and time. We should measure it against the system many patients actually experience: long waits, rushed visits, fragmented records, and limited access to specialists.

    The real question is not whether AI matches the judgment of a thoughtful physician with time to think. It is whether AI can help patients make better use of their own data when that physician is not available—and whether it does so in a way aligned with patients’ interests.


    Why individual calibration is not enough

    Learning to interrogate AI systems helps. But it does not solve the structural problem.

    Patients should not have to reverse-engineer the values embedded in their medical advice. Clinicians should not have to guess how an AI system will behave when trade-offs arise between cost, benefit, risk, and autonomy. Regulators should not have to discover misalignment only after harm occurs at scale. If AI is going to influence care at scale—and it already does—values can no longer remain implicit.

    This is where the Human Values Project (HVP) begins.

    The aim of HVP is to make the values embedded in clinical AI measurable, visible, and discussable. We do this by systematically studying how clinicians, patients, and ethicists actually decide in value-laden medical scenarios—and by benchmarking AI systems against that human variation. Not to impose a single “correct” value system, but to make differences explicit before they are locked into software and deployed across health systems. The HVP already brings together clinicians, patients, and policymakers across the globe.

    In the op-ed, I called for public and leadership pressure for truthful labeling of the influences and alignment procedures shaping clinical AI. Such labeling is only meaningful if we have benchmarks against which to measure it. That is what HVP provides.


    Conclusion

    Medicine is full of decisions that lack a single right answer. Should we favor the sickest, the youngest, or the most likely to benefit? Should we prioritize autonomy, cost, or fairness? Reasonable people disagree.

    AI does not eliminate those disagreements. It encodes them.

    The future of clinical AI depends not only on technical accuracy, but on visible alignment with values that society finds acceptable. If we fail to make those values explicit, AI will quietly entrench the priorities of the most powerful actors in a $5-trillion system. If we succeed, we have a chance to build decision systems that earn trust—not because they are flawless, but because their commitments are transparent.

    That is the wager of the Human Values Project.


    How to participate in the Human Values Project

    The Human Values Project is an international, ongoing effort, and participation can take several forms:

    • Clinicians:
      Contribute to structured decision-making surveys that capture how you approach difficult clinical trade-offs in real-world scenarios. These data help define the range—and limits—of reasonable human judgment.
    • Patients and caregivers:
      Participate in parallel surveys that reflect patient values and preferences, especially in situations where autonomy, risk, and quality of life are in tension.
    • Ethicists, policymakers, and researchers:
      Help articulate and evaluate normative frameworks that can guide alignment, without assuming a single universal standard.
    • Health systems and AI developers:
      Collaborate on benchmarking and transparency efforts so that AI systems disclose how they behave in value-sensitive clinical situations.

    Participation does not require endorsing a particular ethical framework or AI approach. It requires a willingness to make values explicit rather than implicit. Participants will receive updates on findings and early access to benchmarking tools. If you want to learn more or wish to participate, visit the site: https://hvp.global or send email to [email protected]

    If AI is going to help make our data work for us, then the values shaping its advice must be visible—to patients, clinicians, and society at large.



    For those wanting to go deeper, the following papers lay out some of the conceptual and empirical groundwork for HVP.

    Kohane IS, Manrai AK. The missing value of medical artificial intelligence. Nat Med. 2025;31: 3962–3963. doi:10.1038/s41591-025-04050-6
    
    Kohane IS. The Human Values Project. In: Hegselmann S, Zhou H, Healey E, Chang T, Ellington C, Mhasawade V, et al., editors. Proceedings of the 4th Machine Learning for Health Symposium. PMLR; 15--16 Dec 2025. pp. 14–18. Available: https://proceedings.mlr.press/v259/kohane25a.html
    
    Kohane I. Systematic characterization of the effectiveness of alignment in large language models for categorical decisions. arXiv [cs.CL]. 2024. Available: http://arxiv.org/abs/2409.18995
      
    Yu K-H, Healey E, Leong T-Y, Kohane IS, Manrai AK. Medical artificial intelligence and human values. N Engl J Med. 2024;390: 1895–1904. doi:10.1056/NEJMra2214183
    

    Getting your own data

    To try this with your own information, you first need access to it.

    Patient portals.
    Most health systems offer portals (such as MyChart) where you can view and download visit summaries, lab results, imaging reports, medication lists, and immunizations. Many now support exports in standardized formats, though completeness varies.

    HIPAA right of access.
    Under HIPAA, you have a legal right to a copy of your medical records. Providers must respond within 30 days (with a possible extension) and may charge a reasonable copying fee. The Office for Civil Rights has increasingly enforced this right.

    Apple Health and other aggregators.
    Under the 21st Century Cures Act, patients have access to a computable subset of their data. Apple Health can aggregate records across participating health systems, creating a longitudinal view you can export. Similar options exist on Android and via third-party services. I will expound on that in another post.

    Formats matter—but less than you think.
    PDFs are harder to process computationally than structured formats like C-CDA or FHIR, but for the prompts above, even a discharge summary PDF is enough.