Can a Delphi panel look like it agrees while remaining divided?
Yes. A result can satisfy a predefined consensus rule while still containing a consequential minority view. Report the classification and inspect the pattern behind it; they answer different questions.
The same median and percentage can tell different stories
Consider two hypothetical panels of 20 people rating one item on a nine-point scale. Suppose the protocol requires at least 80% of valid ratings in the 7–9 band. These examples use that single percentage rule for illustration; they are not RAND/UCLA classifications.
| Panel | Ratings | Median | In the 7–9 band | Pattern to inspect |
|---|---|---|---|---|
| A | 16 ratings of 8; 4 ratings of 6 | 8 | 16/20 = 80% | High ratings with a less-positive minority. |
| B | 16 ratings of 8; 4 ratings of 1 | 8 | 16/20 = 80% | High ratings alongside a strongly opposing minority. |
Both pass this one rule, but the remaining four judgments are very different. A separate rule limiting low-end ratings could change the classification; its effect should be determined in advance, not added after seeing Panel B.
New methodological research asks what the data can distinguish
Looten and Saguin’s 2026 paper, The information limit of consensus detection on bounded ordinal scales, examines model-based discrimination between consensus and latent polarization. It describes information limits and shows how mean-and-variance summaries can be less informative than fuller use of the ratings; the authors provide companion verification code. Read the paper.
This is theoretical methodological work, not a validation of a new universal Delphi classification system. Its sample-size implications depend on the competing rating models, their separation, and the tolerated error. They do not establish that every Delphi needs a particular minimum number of experts.
Our practical interpretation is to preserve the response pattern, disclose the limits of small panels, and avoid treating a compact summary as proof that a meaningful split has been ruled out.
Inspect distributions, rationales, and context together
- Response distribution: Show the count at each scale point, including opposing ratings that a central summary may obscure.
- Per-item denominator: Report valid ratings, missing responses, and “unable to judge” selections separately under the protocol’s rules.
- Stakeholder pattern: Examine relevant groups without assuming every split follows a role, country, or profession.
- Comments: Look for different interpretations, settings, assumptions, or priorities behind the ratings.
- Round history: Check whether apparent convergence reflects reconsideration or the loss of dissenting participants.
A separated cluster is a reason to investigate, not automatic proof of two stable camps. Small counts, ambiguous wording, different use of the scale, or missing scenario information can produce a similar-looking pattern.
Stability also does not establish consensus: a group can remain consistently divided. See when Delphi should preserve two answers for the separate question of what to do with persistent, interpretable disagreement.
Keep protocol results separate from interpretive cautions
Maintain the predefined consensus classification in the primary analysis. If a distribution, subgroup, or sensitivity check reveals an important concern, report it alongside that result rather than silently changing the rule.
Use “indeterminate” only where the protocol or a clearly identified supplementary analysis defines what it means. It is not interchangeable with “no consensus” or RAM’s uncertain/equivocal category. A RAM result can be uncertain because of its median band or disagreement rule; an information limitation is a different explanation.
Exploratory flags can guide discussion or future research. Identify them as exploratory and do not let them become post hoc exclusions presented as planned decisions.
Report the limits without inventing a new sample-size rule
Give the panel-size rationale, stakeholder composition, per-item counts, exact consensus rules, distributions, and attrition by round. Explain which conclusions are descriptive judgments of the recruited panel and which would need broader validation.
If a sensitivity check is used, state its purpose, assumptions, decision rule, and whether it was planned. For example, inspect whether a conclusion changes when a necessary stakeholder group is considered separately; do not describe that check as a universally validated polarization test.
For very small panels, one rating can materially alter a percentage or median. That does not make small-panel work useless, but it limits the confidence with which absence of disagreement can be claimed.
A suitable interpretation might be: “The item met the predefined criterion, but a small opposing cluster remained. Its comments raised a setting-specific concern that warrants further evaluation.”
Use Surveylet’s study record, with explicit analytical boundaries
Surveylet supports rating distributions, comments, stakeholder grouping, rounds, and analysis-ready exports. These provide material for reviewing whether a summary conceals an important pattern; interpretation remains the research team’s responsibility.
In continuous Real-Time Delphi, an administrator can create a snapshot by creating a survey round at any time. Independently, each question can use a minimum panelist response count before its feedback is displayed. The response-count setting does not create a snapshot and is not a statistical assurance that hidden polarization is detectable.
For help presenting the findings, discuss distribution tables, stakeholder comparisons, and limitations through Calibrum’s Delphi Analysis & Final Report service.
Sources and scope
- Looten V, Saguin E. The information limit of consensus detection on bounded ordinal scales. Statistics & Probability Letters; 2026. This is older research newly surfaced in our October review.
- Looten V, Saguin E. Companion verification code. Zenodo; 2026.
- Fitch K, et al. The RAND/UCLA Appropriateness Method User’s Manual. RAND; 2001.
Use the recent paper to examine the limits of inference, not as a reason to discard established protocol rules. New diagnostics should be assessed for their intended setting before they influence clinical or policy recommendations.
Report agreement without erasing the remaining questions
We can help organize the study record and reporting requirements so readers see both the predefined result and its limitations.
