Skip to content
Look beyond one summary statistic

Can a Delphi panel look like it agrees while remaining divided?

Yes. A result can satisfy a predefined consensus rule while still containing a consequential minority view. Report the classification and inspect the pattern behind it; they answer different questions.

The same median and percentage can tell different stories

Consider two hypothetical panels of 20 people rating one item on a nine-point scale. Suppose the protocol requires at least 80% of valid ratings in the 7–9 band. These examples use that single percentage rule for illustration; they are not RAND/UCLA classifications.

Illustrative ratings, not data from a published study
PanelRatingsMedianIn the 7–9 bandPattern to inspect
A16 ratings of 8; 4 ratings of 6816/20 = 80%High ratings with a less-positive minority.
B16 ratings of 8; 4 ratings of 1816/20 = 80%High ratings alongside a strongly opposing minority.

Both pass this one rule, but the remaining four judgments are very different. A separate rule limiting low-end ratings could change the classification; its effect should be determined in advance, not added after seeing Panel B.

Do not claim more than the rule establishes. “Met the predefined criterion” does not mean unanimity, absence of polarization, or proof that the panel’s conclusion is correct.

New methodological research asks what the data can distinguish

Looten and Saguin’s 2026 paper, The information limit of consensus detection on bounded ordinal scales, examines model-based discrimination between consensus and latent polarization. It describes information limits and shows how mean-and-variance summaries can be less informative than fuller use of the ratings; the authors provide companion verification code. Read the paper.

This is theoretical methodological work, not a validation of a new universal Delphi classification system. Its sample-size implications depend on the competing rating models, their separation, and the tolerated error. They do not establish that every Delphi needs a particular minimum number of experts.

Our practical interpretation is to preserve the response pattern, disclose the limits of small panels, and avoid treating a compact summary as proof that a meaningful split has been ruled out.

Inspect distributions, rationales, and context together

  • Response distribution: Show the count at each scale point, including opposing ratings that a central summary may obscure.
  • Per-item denominator: Report valid ratings, missing responses, and “unable to judge” selections separately under the protocol’s rules.
  • Stakeholder pattern: Examine relevant groups without assuming every split follows a role, country, or profession.
  • Comments: Look for different interpretations, settings, assumptions, or priorities behind the ratings.
  • Round history: Check whether apparent convergence reflects reconsideration or the loss of dissenting participants.

A separated cluster is a reason to investigate, not automatic proof of two stable camps. Small counts, ambiguous wording, different use of the scale, or missing scenario information can produce a similar-looking pattern.

Stability also does not establish consensus: a group can remain consistently divided. See when Delphi should preserve two answers for the separate question of what to do with persistent, interpretable disagreement.

Keep protocol results separate from interpretive cautions

Maintain the predefined consensus classification in the primary analysis. If a distribution, subgroup, or sensitivity check reveals an important concern, report it alongside that result rather than silently changing the rule.

Use “indeterminate” only where the protocol or a clearly identified supplementary analysis defines what it means. It is not interchangeable with “no consensus” or RAM’s uncertain/equivocal category. A RAM result can be uncertain because of its median band or disagreement rule; an information limitation is a different explanation.

Different labels, different reasons. An item that fails a percentage threshold, an item that meets a RAM disagreement rule, and an item whose small sample cannot resolve a suspected pattern should not be assigned one unexplained catch-all label.

Exploratory flags can guide discussion or future research. Identify them as exploratory and do not let them become post hoc exclusions presented as planned decisions.

Report the limits without inventing a new sample-size rule

Give the panel-size rationale, stakeholder composition, per-item counts, exact consensus rules, distributions, and attrition by round. Explain which conclusions are descriptive judgments of the recruited panel and which would need broader validation.

If a sensitivity check is used, state its purpose, assumptions, decision rule, and whether it was planned. For example, inspect whether a conclusion changes when a necessary stakeholder group is considered separately; do not describe that check as a universally validated polarization test.

For very small panels, one rating can materially alter a percentage or median. That does not make small-panel work useless, but it limits the confidence with which absence of disagreement can be claimed.

A suitable interpretation might be: “The item met the predefined criterion, but a small opposing cluster remained. Its comments raised a setting-specific concern that warrants further evaluation.”

Use Surveylet’s study record, with explicit analytical boundaries

Surveylet supports rating distributions, comments, stakeholder grouping, rounds, and analysis-ready exports. These provide material for reviewing whether a summary conceals an important pattern; interpretation remains the research team’s responsibility.

In continuous Real-Time Delphi, an administrator can create a snapshot by creating a survey round at any time. Independently, each question can use a minimum panelist response count before its feedback is displayed. The response-count setting does not create a snapshot and is not a statistical assurance that hidden polarization is detectable.

Diagnostic ideas are not built-in promises. A validated polarization test, model-based panel-size calculation, or automatic indeterminate flag would require separate methodological and product work. This article does not claim that Surveylet already supplies those functions.

For help presenting the findings, discuss distribution tables, stakeholder comparisons, and limitations through Calibrum’s Delphi Analysis & Final Report service.

Sources and scope

Use the recent paper to examine the limits of inference, not as a reason to discard established protocol rules. New diagnostics should be assessed for their intended setting before they influence clinical or policy recommendations.

Report agreement without erasing the remaining questions

We can help organize the study record and reporting requirements so readers see both the predefined result and its limitations.