Using Delphi to reveal policy trade-offs—not manufacture agreement
A policy can help one constituency and disadvantage another. A useful Delphi makes those consequences visible rather than hiding them behind a single “support” score.
Ask who benefits, not only whether the panel approves
An October 2026 Health Affairs study convened 39 experts in a two-round e-Delphi to assess 15 Medicare Advantage reforms. Participants rated the consequences separately for beneficiaries, the federal government, and health plans; favorable judgments differed across those three constituencies. Read the study.
The article illustrates an important design choice: the same proposal can have a different impact depending on whose interests are being considered. The practical workflow below is Calibrum’s guidance for making those differences interpretable, not a reproduction of the study’s instrument or a claim about which policies should be adopted.
Do not confuse who is rating with who is affected
Stakeholder subgroup analysis and constituency-specific impact ratings answer different questions. A researcher, patient advocate, or industry representative may assess the impact on every constituency, not just their own.
| Part | Question | Example |
|---|---|---|
| Panelist role | Who supplied this judgment? | Researcher, practitioner, service user, industry expert. |
| Impact dimension | Whose outcomes are being assessed? | Service users, public funders, delivery organizations. |
| Subgroup comparison | Do different panelist groups judge the same impact differently? | Service users and practitioners disagree about the predicted burden on users. |
Keep these fields separate in the questionnaire and export. Otherwise, an apparent difference between affected groups can be confused with a difference between the people providing the ratings.
Define the consequences and assumptions before rating
Specify the proposed change, baseline comparator, affected populations, setting, and time horizon. A reform compared with current practice is a different question from the same reform compared with another proposed policy.
Illustrative policy item
Proposal: introduce a standardized application process for a public service.
Ask for separate ratings of the expected impact on applicants, the funding authority, and delivery organizations. If implementation feasibility also matters, collect it as another dimension rather than folding it into the impact ratings.
- Use explicit anchors for negative, neutral, and positive impact.
- Define what an uncertainty or “unable to judge” response means; do not equate it with no impact.
- State whether cost, access, workload, equity, or another consequence should be considered separately.
- Request a brief rationale or the condition on which a rating depends.
- Allow panelists to identify an affected group or consequence missing from the draft.
Limit the number of dimensions to those that inform the decision. A matrix that asks about every imaginable effect can increase burden without producing a clearer result.
A mixed impact profile is not necessarily panel disagreement
A panel might consistently judge a proposal beneficial to service users and costly to providers. That is agreement about a trade-off. A different pattern occurs when panelists disagree about the effect on service users themselves.
Report the first as contrasting consequences across dimensions and the second as variation or disagreement within a dimension. Neither can be interpreted adequately from an overall average.
Produce an item-by-constituency evidence table
For every proposal, show the rating profile for each affected constituency, the valid-response count, distribution or spread, consensus classification, and the conditions described in comments. Keep panelist subgroup comparisons alongside that profile when the protocol requires them.
Distinguish proposals with broadly favorable effects, those with explicit benefit–burden trade-offs, and those with unresolved uncertainty. This is a descriptive reporting structure, not a substitute for the study’s predefined classifications.
Document whether weighting, stakeholder vetoes, required-group thresholds, a live discussion, or later multi-criteria prioritization influenced the final recommendation. If a meeting changes the decision, retain the rating result and meeting decision separately.
Expert judgments are not measured policy effects. Explain the evidence available to the panel, the recruitment limitations, and where later evaluation would be needed to test predicted consequences.
Use Surveylet to structure the judgments and preserve the record
Surveylet can collect separate rating questions, organize panelists by stakeholder group, provide controlled feedback, retain comments, and export results for analysis. Use consistent proposal identifiers so the effect ratings and explanations can be brought together without confusing the dimensions.
For a focused meeting, selected unresolved items can be taken into a Live Consensus Session. Facilitation should help clarify assumptions and consequences, not pressure the group to remove a legitimate objection.
The same structure can inform education policy, resource allocation, professional standards, infrastructure planning, or technology governance. See Delphi beyond healthcare for other uses.
Sources and next steps
- Khodyakov D, et al. Charting a Path Forward for Medicare Advantage: Expert Consensus on Policy Priorities and Trade-Offs. Health Affairs; October 5, 2026. Also available through the RAND publication record.
- RAND. Methodological Guidance for Conducting and Critically Appraising Delphi Panels; 2023.
Before inviting the panel, write down the decision the findings will inform, who is affected, which impacts need separate ratings, and how uncertainty or conflicting consequences will be handled. A defensible result can be a clear map of trade-offs rather than one universally endorsed proposal.
Make the trade-offs visible
We can help configure the questions, stakeholder structure, and feedback workflow around your approved policy-study design.
