Skip to content
Delphi methodology

Two Delphi design errors that can undermine consensus

A Delphi study can use an expert panel, several rounds, and a predefined threshold yet still produce a weak conclusion. Two design choices deserve particular attention: deciding who may contribute during early rounds and asking panelists to judge several variables in one statement.

Consensus depends on what the panel was allowed to consider

Consensus is not only a calculation at the end of a study. It is shaped by who helped define the candidate items, how those items were written, what feedback participants received, and what alternatives remained available in later rounds.

A 2026 editorial, Delphi Consensus Is Limited by Deviations From Delphi Methodology, highlighted two practices that can damage interpretation: allowing only selected subgroups to contribute during early rounds and combining several clinical variables in a single item.

Error 1: restricting early contribution without a clear methodological reason

Early rounds often define the content that the full panel will later rate. If one stakeholder group generates or screens the item set while another group is invited only after the choices have narrowed, the later participants may be asked to judge a framework they had no opportunity to shape.

This does not mean every participant must perform every task. A steering group may legitimately synthesize evidence, remove exact duplicates, or prepare candidate wording. The risk arises when the design excludes a relevant perspective from content development and then treats the final vote as though all stakeholder groups had equivalent influence over the agenda.

Questions to resolve before launch

  • Which groups may propose new items or concerns?
  • Who may revise, merge, or remove candidate items?
  • Will the complete panel see the evidence and reasons for those decisions?
  • Can a later-entering stakeholder group identify missing content?
  • How will staged participation be justified and reported?

Staged participation can be useful—but it must be visible

How early-round participation changes interpretation
Design choicePotential valueMethodological safeguard
Open first round for the full panelBroadens item generation and reveals missing perspectives early.Use a transparent coding and consolidation process for qualitative input.
Evidence review or steering group drafts the initial listCreates an efficient, evidence-informed starting point.Explain who drafted and screened items, then allow the wider panel to propose additions or identify omissions.
Different groups enter at different stagesMay match distinct expertise or practical constraints.Predefine each group’s role and do not imply that every group shaped every stage equally.

Error 2: combining several judgments in one statement

A compound or double-barreled item asks participants to rate more than one proposition with one response. In clinical work, a single statement may combine population, disease stage, biomarker, treatment, timing, outcome, feasibility, and cost. Panelists can agree with some parts and disagree with others, leaving the rating impossible to interpret.

ProblemBetter design
“The intervention is effective, affordable, and feasible.”Rate effectiveness, affordability, and feasibility separately.
One recommendation covers two different patient groups or settings.Create separate items or use clearly defined scenario variations.
One response must represent both validity and feasibility.Use parallel rating dimensions with a predefined rule for how both results affect the decision.
A low rating on a compound item does not reveal what failed. The panel may reject the evidence, population, wording, implementation burden, or only one condition embedded in the statement. Separate questions produce results that can be interpreted and revised.

A practical pre-launch review

Review the questionnaire at two levels. First, inspect the study pathway: who contributes, who decides what advances, and whether every necessary perspective has a defined opportunity to influence the content. Second, inspect every item: one proposition, one clear population or scenario, one meaningful rating task, and a response scale that matches the question.

  • Pilot the instructions, items, scales, branching, and feedback with people similar to the intended panel.
  • Separate distinct constructs such as agreement, importance, clarity, validity, feasibility, and appropriateness.
  • Use controlled scenario variations when context must change systematically.
  • Predefine how new, revised, merged, split, and removed items will be documented.
  • Ask whether a panelist could reasonably agree with one part of an item and disagree with another.

Software can structure rounds and rating dimensions, but it cannot decide whether an item is scientifically valid. The study team remains responsible for content expertise, methodological review, and pilot testing.

Report the path, not only the final percentage

A reader should be able to see who contributed during each stage, how the initial list was developed, who could add or remove items, what changed between rounds, and why. For compound items that were split or rewritten, retain the wording history so the final result can be traced back to the original concern.

Report subgroup participation by round when stakeholder roles matter. If some groups entered later, explain the reason and the consequences for interpretation. If a steering group made decisions between rounds, document the criteria and decision record.

Surveylet can support separate stakeholder groups, multi-round questionnaires, distinct rating dimensions, controlled feedback, response tracking, and exports. These functions help preserve the study record, while the research team defines and defends the methodology.

What these errors have in common

Both errors hide the source of the result. Restricted early input can hide which perspectives never entered the questionnaire. Compound statements can hide which component caused agreement or rejection. A stronger Delphi design makes both the contributors and the judgment task explicit.

The goal is not to make every Delphi study identical. It is to ensure that the final classification means what the report claims it means.

Design each round so the result remains interpretable

Calibrum can help configure a Surveylet workflow around your approved questionnaire, stakeholder structure, rating dimensions, feedback plan, and round decisions.