Skip to content
Expert judgment for complex decisions

Using Delphi to reveal policy trade-offs—not manufacture agreement

A policy can help one constituency and disadvantage another. A useful Delphi makes those consequences visible rather than hiding them behind a single “support” score.

Ask who benefits, not only whether the panel approves

An October 2026 Health Affairs study convened 39 experts in a two-round e-Delphi to assess 15 Medicare Advantage reforms. Participants rated the consequences separately for beneficiaries, the federal government, and health plans; favorable judgments differed across those three constituencies. Read the study.

The article illustrates an important design choice: the same proposal can have a different impact depending on whose interests are being considered. The practical workflow below is Calibrum’s guidance for making those differences interpretable, not a reproduction of the study’s instrument or a claim about which policies should be adopted.

Do not confuse who is rating with who is affected

Stakeholder subgroup analysis and constituency-specific impact ratings answer different questions. A researcher, patient advocate, or industry representative may assess the impact on every constituency, not just their own.

Two separate parts of the analysis
PartQuestionExample
Panelist roleWho supplied this judgment?Researcher, practitioner, service user, industry expert.
Impact dimensionWhose outcomes are being assessed?Service users, public funders, delivery organizations.
Subgroup comparisonDo different panelist groups judge the same impact differently?Service users and practitioners disagree about the predicted burden on users.

Keep these fields separate in the questionnaire and export. Otherwise, an apparent difference between affected groups can be confused with a difference between the people providing the ratings.

Define the consequences and assumptions before rating

Specify the proposed change, baseline comparator, affected populations, setting, and time horizon. A reform compared with current practice is a different question from the same reform compared with another proposed policy.

Illustrative policy item

Proposal: introduce a standardized application process for a public service.

Ask for separate ratings of the expected impact on applicants, the funding authority, and delivery organizations. If implementation feasibility also matters, collect it as another dimension rather than folding it into the impact ratings.

  • Use explicit anchors for negative, neutral, and positive impact.
  • Define what an uncertainty or “unable to judge” response means; do not equate it with no impact.
  • State whether cost, access, workload, equity, or another consequence should be considered separately.
  • Request a brief rationale or the condition on which a rating depends.
  • Allow panelists to identify an affected group or consequence missing from the draft.

Limit the number of dimensions to those that inform the decision. A matrix that asks about every imaginable effect can increase burden without producing a clearer result.

A mixed impact profile is not necessarily panel disagreement

A panel might consistently judge a proposal beneficial to service users and costly to providers. That is agreement about a trade-off. A different pattern occurs when panelists disagree about the effect on service users themselves.

Report the first as contrasting consequences across dimensions and the second as variation or disagreement within a dimension. Neither can be interpreted adequately from an overall average.

Do not let positive effects silently cancel negative effects. A blended score can make a consequential harm disappear. Any weighting or prioritization model must be justified and visible; it is a separate decision step, not an automatic consequence of Delphi consensus.

Produce an item-by-constituency evidence table

For every proposal, show the rating profile for each affected constituency, the valid-response count, distribution or spread, consensus classification, and the conditions described in comments. Keep panelist subgroup comparisons alongside that profile when the protocol requires them.

Distinguish proposals with broadly favorable effects, those with explicit benefit–burden trade-offs, and those with unresolved uncertainty. This is a descriptive reporting structure, not a substitute for the study’s predefined classifications.

Document whether weighting, stakeholder vetoes, required-group thresholds, a live discussion, or later multi-criteria prioritization influenced the final recommendation. If a meeting changes the decision, retain the rating result and meeting decision separately.

Expert judgments are not measured policy effects. Explain the evidence available to the panel, the recruitment limitations, and where later evaluation would be needed to test predicted consequences.

Use Surveylet to structure the judgments and preserve the record

Surveylet can collect separate rating questions, organize panelists by stakeholder group, provide controlled feedback, retain comments, and export results for analysis. Use consistent proposal identifiers so the effect ratings and explanations can be brought together without confusing the dimensions.

For a focused meeting, selected unresolved items can be taken into a Live Consensus Session. Facilitation should help clarify assumptions and consequences, not pressure the group to remove a legitimate objection.

Confirm custom reporting requirements. An item-by-constituency table, automatic trade-off alert, weighted policy score, or custom visualization is not implied by collecting several dimensions. Agree the analysis and deliverables with the research team and Calibrum before launch.

The same structure can inform education policy, resource allocation, professional standards, infrastructure planning, or technology governance. See Delphi beyond healthcare for other uses.

Sources and next steps

Before inviting the panel, write down the decision the findings will inform, who is affected, which impacts need separate ratings, and how uncertainty or conflicting consequences will be handled. A defensible result can be a clear map of trade-offs rather than one universally endorsed proposal.

Make the trade-offs visible

We can help configure the questions, stakeholder structure, and feedback workflow around your approved policy-study design.