Two Delphi design errors that can undermine consensus
A Delphi study can use an expert panel, several rounds, and a predefined threshold yet still produce a weak conclusion. Two design choices deserve particular attention: deciding who may contribute during early rounds and asking panelists to judge several variables in one statement.
Consensus depends on what the panel was allowed to consider
Consensus is not only a calculation at the end of a study. It is shaped by who helped define the candidate items, how those items were written, what feedback participants received, and what alternatives remained available in later rounds.
A 2026 editorial, Delphi Consensus Is Limited by Deviations From Delphi Methodology, highlighted two practices that can damage interpretation: allowing only selected subgroups to contribute during early rounds and combining several clinical variables in a single item.
Error 1: restricting early contribution without a clear methodological reason
Early rounds often define the content that the full panel will later rate. If one stakeholder group generates or screens the item set while another group is invited only after the choices have narrowed, the later participants may be asked to judge a framework they had no opportunity to shape.
This does not mean every participant must perform every task. A steering group may legitimately synthesize evidence, remove exact duplicates, or prepare candidate wording. The risk arises when the design excludes a relevant perspective from content development and then treats the final vote as though all stakeholder groups had equivalent influence over the agenda.
Questions to resolve before launch
- Which groups may propose new items or concerns?
- Who may revise, merge, or remove candidate items?
- Will the complete panel see the evidence and reasons for those decisions?
- Can a later-entering stakeholder group identify missing content?
- How will staged participation be justified and reported?
Staged participation can be useful—but it must be visible
| Design choice | Potential value | Methodological safeguard |
|---|---|---|
| Open first round for the full panel | Broadens item generation and reveals missing perspectives early. | Use a transparent coding and consolidation process for qualitative input. |
| Evidence review or steering group drafts the initial list | Creates an efficient, evidence-informed starting point. | Explain who drafted and screened items, then allow the wider panel to propose additions or identify omissions. |
| Different groups enter at different stages | May match distinct expertise or practical constraints. | Predefine each group’s role and do not imply that every group shaped every stage equally. |
Error 2: combining several judgments in one statement
A compound or double-barreled item asks participants to rate more than one proposition with one response. In clinical work, a single statement may combine population, disease stage, biomarker, treatment, timing, outcome, feasibility, and cost. Panelists can agree with some parts and disagree with others, leaving the rating impossible to interpret.
| Problem | Better design |
|---|---|
| “The intervention is effective, affordable, and feasible.” | Rate effectiveness, affordability, and feasibility separately. |
| One recommendation covers two different patient groups or settings. | Create separate items or use clearly defined scenario variations. |
| One response must represent both validity and feasibility. | Use parallel rating dimensions with a predefined rule for how both results affect the decision. |
A practical pre-launch review
Review the questionnaire at two levels. First, inspect the study pathway: who contributes, who decides what advances, and whether every necessary perspective has a defined opportunity to influence the content. Second, inspect every item: one proposition, one clear population or scenario, one meaningful rating task, and a response scale that matches the question.
- Pilot the instructions, items, scales, branching, and feedback with people similar to the intended panel.
- Separate distinct constructs such as agreement, importance, clarity, validity, feasibility, and appropriateness.
- Use controlled scenario variations when context must change systematically.
- Predefine how new, revised, merged, split, and removed items will be documented.
- Ask whether a panelist could reasonably agree with one part of an item and disagree with another.
Software can structure rounds and rating dimensions, but it cannot decide whether an item is scientifically valid. The study team remains responsible for content expertise, methodological review, and pilot testing.
Report the path, not only the final percentage
A reader should be able to see who contributed during each stage, how the initial list was developed, who could add or remove items, what changed between rounds, and why. For compound items that were split or rewritten, retain the wording history so the final result can be traced back to the original concern.
Report subgroup participation by round when stakeholder roles matter. If some groups entered later, explain the reason and the consequences for interpretation. If a steering group made decisions between rounds, document the criteria and decision record.
Surveylet can support separate stakeholder groups, multi-round questionnaires, distinct rating dimensions, controlled feedback, response tracking, and exports. These functions help preserve the study record, while the research team defines and defends the methodology.
What these errors have in common
Both errors hide the source of the result. Restricted early input can hide which perspectives never entered the questionnaire. Compound statements can hide which component caused agreement or rejection. A stronger Delphi design makes both the contributors and the judgment task explicit.
The goal is not to make every Delphi study identical. It is to ensure that the final classification means what the report claims it means.
Design each round so the result remains interpretable
Calibrum can help configure a Surveylet workflow around your approved questionnaire, stakeholder structure, rating dimensions, feedback plan, and round decisions.
