The Work Beneath the Conversation: I Could Do That If I Had Time
Three trials tested whether the conversation is the intervention. Two came back null and one looked inside the machine. None of them supports the sentence — or our answer to it.
A hospitalist with eighteen patients and a two o'clock discharge target does not have forty-five uninterrupted minutes, and no amount of specialist expertise fixes a schedule built to prevent a conversation from happening.
So when a colleague says I could have that conversation myself, I just don't have the time, half of it is plainly true. The other half is a claim that time is the only thing between their skills and specialist palliative care.
We sometimes rebut with an inventory: chart review, prognostic reconciliation, the interdisciplinary work underneath the forty-five minutes anyone can see. Accurate, and beside the point. Nobody (okay maybe "most people") accused us of not working. They said the work is a conversation, and that an hour of protected time would close the gap.
I think we half-believe them, which is why the rebuttal comes out the way it does. The hallway sentence and our answer to it both put the intervention inside the meeting, then spend their energy on who deserves to run it. Three trials have tested that assumption. It doesn't hold.
We Taught It and Couldn't Measure the Impact
Curtis and colleagues randomized internal medicine and nurse practitioner trainees at two universities to thirty-two hours of simulation-based communication training, adapted from the workshop that improved oncology fellows' skills, or to usual education. Anthony Back is an author. This is not an outside critique of our communication work. It is the field auditing its own flagship curriculum.
The trainees learned it. A companion analysis confirmed they got better at giving bad news and responding to emotion on standardized-patient encounters.
Then their patients rated them, and their families rated them, and the nurses and attendings who watched them work rated them. And we couldn't measure a difference. Not on quality of communication, the primary outcome. Not on quality of end-of-life care.
Curtis's own conclusion questions whether the skills transferred, and — this is the part the field skipped — whether we can measure them at all. No minimal clinically important difference has ever been established for the instrument. The ratings clustered so hard at the top of the scale that the analysis required censored regression to handle it. Trained standardized patients rate communication more reliably than real ones do, and Fallowfield's oncologist trial found the same split years earlier: skills improved by trained raters, invisible to patients.
We taught the skill. We verified the skill. We cannot demonstrate that anyone on the receiving end experienced it.
We Delivered It Alone, and It Vanished
The other half of the sentence got its own trial. Carson and colleagues enrolled patients across four medical ICUs who had been ventilated at least a week with no expectation of weaning or dying soon — the population our own consensus documents had nominated for automatic specialist consultation. Families in the intervention arm got structured meetings led by a palliative care physician and nurse practitioner. Families in the control arm got whatever their ICU team already did.
Anxiety and depression at three months, the primary outcome: no difference, and the gap that did appear fell short of the smallest one that would have mattered. Length of stay, unchanged. Survival, unchanged. Family satisfaction and discussion of patient preferences both came back numerically better in the control arm.
Now look at what the protocol was. Symptom management sat outside it. Social work and chaplaincy participated as needed. ICU physicians attended fewer than one in ten first meetings and almost none of the second. Nothing provided continuity across services or units. Surrogates averaged fewer than two specialist meetings. And the specialist meetings did not replace the ICU team's own family meetings — those happened at the same rate in both arms.
So the protocol pulled the conversation out of specialist palliative care and delivered it by itself. Our clinicians showed up, with the time, ran the meeting, and left. That is the hallway sentence rendered as a study design. The authors say as much in their discussion: the intervention may have failed because it lacked continuity, symptom management, and the other disciplines.
Put the two trials side by side and they are one experiment run from opposite ends. Curtis put the skill into a generalist and removed the team. Carson put a specialist in the room and removed everything else the team does. Both isolated the conversation. Neither found anything.
Inside the Trial That Worked
Temel's second early integrated palliative care trial gave patients with newly diagnosed incurable lung or GI cancer monthly palliative care visits from within four weeks of diagnosis until death. It improved quality of life and mood. In 2018, Greer and colleagues went back into the data to ask how (this is actually my very favorite, semi-secret Temel publication).
The answer was coping. At three months, nothing. At six months, the arms separated — and here is the detail that gets left out. The palliative care patients barely moved. The usual care patients got worse. Almost the entire between-arm difference is the control group's decline. The intervention held the line.
Most of the effect on quality of life and mood ran through that change, and once coping entered the model the direct effect of palliative care went away. The mediated gain came to about a point on a 108-point scale. Statistically solid, clinically small.
The ingredients matter more than the size. When Greer separated the components, acceptance and positive reframing produced the effect. Active coping did not. What worked looks like cognitive reappraisal — helping someone revise the meaning of what is happening to them — delivered monthly across half a year.
Greer's team drew the implication themselves, in print: these data could specify which components need specialty palliative care and which could come from primary or mental health clinicians. Our own researchers put the substitution question on the table before anyone outside the field did.
The Only Thing That Moved
Both null trials produced one signal, and we have handled it inexpertly.
Carson found more post-traumatic stress symptoms in the intervention arm. The finding is fragile: exploratory, unadjusted for multiple comparisons, P = .0495, with a confidence interval whose lower bound sat a hair above zero. When the authors asked how many surrogates actually crossed the diagnostic threshold, the difference disappeared.
Curtis found patients of trained residents reporting more depressive symptoms — two points on a scale where five is the smallest change considered meaningful, and one comparison among many. The authors bound it carefully. Then they reach for the explanation that should stop us: patients who understand they have an incurable illness rate their physicians' communication lower. Increasing prognostic awareness may itself produce the worse score.
If a patient's depression score rises because a resident finally told them something true, our instrument reads comprehension as harm — and a program optimizing against that instrument selects for clinicians who leave patients undisturbed.
I'm probably overstating this for effect. Both signals are weak, and both are equally consistent with a hard conversation done poorly. The problem is narrower and worse: we still have no way to tell those outcomes apart, and we have been publishing in this space for twenty-five years.
Who Wasn't in These Trials
Carson excluded surrogates who did not speak English. Temel's trial required patients to manage in English or with minimal interpreter help, analyzed a sample more than nine in ten white, and screened out several hundred patients whose clinicians already thought they needed palliative care. Curtis got responses from fewer than half the patients approached, and the rate fell further among minority patients, among the very old, and among patients in hospice — barely a quarter of whom returned a survey. The authors say plainly that sicker patients participated less, which limited what they could learn about the patients most likely to need the conversation.
The mechanistic evidence for what specialist palliative care does describes English-speaking, mostly white patients, weighted toward those well enough to fill out a form, whose clinicians did not yet think they needed us.
That is a mechanism problem, not a limitations paragraph. If the active ingredient is months of acceptance and reappraisal, it moves through language, through cultural norms about disclosure and family authority, and through an unbroken relationship. Those are the three conditions most reliably destabilized for patients who need an interpreter, whose families decide collectively rather than through one surrogate, or whose coverage moves them mid-illness. Whether reappraisal works the same way with an interpreter is the most consequential open question in our evidence base.
What Follows
I have argued before that the generalist–specialist binary describes the world badly and a competency continuum describes it better. This evidence is the floor under that argument. What works is a relationship over time that includes symptom management and cognitive support, with enough continuity for both to compound. A meeting doesn't describe that. Neither does a tier of clinician.
Don't sell the meeting. Describe the longitudinal object to referrers, executives, and fellows. When a consult arrives asking for one family meeting, the answer is a conversation about what the case actually needs.
Treat discipline mix as dosing. Carson made social work and chaplaincy optional and got nothing. Greer found the working ingredient to be acceptance and reappraisal — trained expertise for clinical social workers and chaplains, and largely not so for physicians. If reappraisal is the mechanism, team composition stops being a staffing preference and becomes a dosing question. A physician-and-APP team is an under-dosed version of the intervention whose results we quote. I've written about what physicians are actually for on these teams. This is the empirical case for the same claim.
Fix the instruments before defending the outcomes. We are arguing about the value of a service using measures with unknown responsiveness, documented ceiling effects, and no ability to distinguish a patient who is upset because they finally understand from a patient who is upset because we handled it badly. That is a solvable research problem, and it is more urgent than another value-proposition deck.
None of this gives the hospitalist back the forty-five minutes. That fight is real.
The Objections That Land
Carson tested a bad version of palliative care. Yes, and that is the argument — with one addition. It was the version the field asked for. Chronic critical illness sat on consensus trigger lists before the trial ran. Carson tested our recommendation, using a palliative-lite model.
Curtis found something where it counted. Among patients who rated their own health as poor — the ones for whom this conversation is most relevant — communication ratings did improve. It is a post hoc subgroup and the authors flag it. It is also what you would expect if the training works and the primary outcome was diluted by patients for whom the conversation never came up. This is the objection I take most seriously.
Absence of demonstrated effect is not demonstrated absence. Carson's confidence interval runs from a modest benefit through a difference in the direction of harm. Both stay live.
Mediation is not causation. Greer says so outright. Coping and outcomes moved in the same window, so the arrow could point either way. Better quality of life may produce more acceptance.
Where I Might Be Wrong
These trials enrolled between 2007 and 2015. Practice, staffing, referral timing, and team sophistication have moved since.
Carson's control arm was unusually good — high satisfaction, high communication ratings, and more off-protocol palliative care consults than the intervention arm received. Against a control that strong, a null says more about the marginal value of one additional meeting than about the intervention.
Reading Carson's design as an isolation experiment is my interpretation. The authors offer it as one of three candidate explanations.
And the implication I like least comes straight out of the evidence I just spent two thousand words citing. If the mechanism is reappraisal over time, some real share of what we do may be deliverable by psycho-oncology or embedded behavioral health, cheaper and at greater scale. I don't want that to be true. It is testable, our own data generated it, and we should run the test before a payer runs it for us.
Final Thoughts
The hallway sentence describes our service exactly as we taught medicine to see it. We spent twenty-five years selling the conversation, everyone bought it, and the reasonable inference from what we sold is that an hour is all that separates the hospitalist from the consult team.
Our own trials describe something we have never learned to ask for: months of contact, a team with the right disciplines actually in the room, symptom work and cognitive work running together, and instruments honest enough to tell a patient who understands from a patient we hurt. That request makes for a harder conversation with a health system than a consult order ever did.
I am a palliative care physician, educator, and professional strategery expert. Known for turning rounds into rants and rants into teaching points. Rounds & Rants represents my views — not those of any institution or professional membership organization where I hold a role. I don't write on their behalf and they don't vet what I publish.
8/25/26 I adjusted the framing in the header of the second section and emphasized that the Curtis trial was as much, if not more, a failure of measurement as it was a failure of skills transference after the generous input from Dr. Back below in the comments.