Delphi consensus is not enough: clarity, importance, and feasibility matter too
A statement can attract agreement and still be unclear, impractical, or insufficiently important. When several judgments matter to the final decision, a defensible Delphi study evaluates them separately and defines in advance how they will be combined.
One percentage cannot answer every question
“Consensus” often describes agreement on a single rating dimension. But many Delphi projects need to know more: Is the item understandable? Is it important? Is it supported by evidence or expert judgment? Can it be implemented? Is it appropriate for the stated population and setting?
Common rating dimensions
Clarity
Is the wording specific, understandable, and interpreted consistently enough to support a reliable judgment?
Importance or priority
Does the panel consider the item sufficiently important to retain, recommend, or prioritize?
Validity
Does the item represent a sound, relevant, and defensible component of the proposed definition, framework, or recommendation?
Feasibility
Can the item be implemented, measured, or delivered in the settings for which it is intended?
Appropriateness
Do expected benefits outweigh harms or disadvantages for the defined scenario, population, or indication?
Agreement
To what extent does the panel endorse the statement, recommendation, definition, or proposed action?
How to predefine a multidimensional decision rule
| Planning decision | What to specify before analysis | Example |
|---|---|---|
| Required dimensions | Which judgments determine acceptance and which are descriptive only? | An item must meet both importance and clarity criteria. |
| Threshold for each dimension | The qualifying response range, required percentage, median, disagreement rule, or other statistic. | At least 80% rate importance 7–9, with no prespecified disagreement. |
| Combination rule | Whether every required dimension must pass or whether another documented rule applies. | Retain only if both importance and clarity pass; revise if importance passes but clarity fails. |
| Item action | What happens after each possible combination of results? | Accept, exclude, revise, carry forward, or refer for steering-committee review. |
| Stakeholder rule | Whether overall consensus is sufficient or important groups must also meet a criterion. | Report overall acceptance but flag when patient and professional panels differ. |
What the final report should show
- Every rating dimension, response scale, threshold, and combination rule defined in the protocol.
- Results for each dimension instead of one blended pass/fail label.
- The specific criterion that caused an item to be accepted, revised, excluded, or carried forward.
- Full distributions, disagreement, and relevant stakeholder-group patterns with clear denominators.
- How qualitative comments explain unclear wording, implementation barriers, or group-specific concerns.
- Any post-launch change to the rules, together with the reason and its effect on interpretation.
A 2026 modified Delphi study illustrates why this matters: items judged strategically important were excluded when their definitions remained ambiguous. The broader lesson is that agreement on value does not cure a failure of clarity.
How Surveylet and Calibrum can help
Surveylet can collect separate structured ratings for each required dimension, provide controlled feedback, organize stakeholder groups, and export the results for criterion-specific analysis. Calibrum can help define the decision rules, configure the workflow, interpret quantitative and qualitative findings, and document why each item passed, failed, or required revision.
Design the criteria before the answers arrive
Tell us what your panel must decide and which judgments matter to that decision. We can help structure the rating dimensions, thresholds, item actions, stakeholder comparisons, and reporting plan before launch.
