Skip to content
CalibrumConsensus analysis
Delphi Consensus

Consensus Thresholds in Delphi Studies

How study teams can define, measure, interpret, and report consensus clearly in Delphi and Real-Time Delphi studies.

What is a consensus threshold?

In a Delphi study, a consensus threshold is the rule used to decide whether an item has reached sufficient agreement. It helps the study team determine which recommendations, outcomes, definitions, priorities, or criteria should be retained, revised, excluded, or reviewed further.

Consensus thresholds are important because they make the interpretation of results more transparent. Instead of saying that “most experts agreed,” the study can explain exactly how agreement was measured and what level of agreement was required.

A consensus threshold should usually be defined before the study begins, not after the results are known. This helps protect the credibility of the study and makes the final conclusions easier to defend.

Simple definition: A consensus threshold is the predefined rule that tells the study team when an item has enough agreement to be considered consensus.

Why consensus thresholds matter

Delphi studies are often used when decisions need to be transparent, evidence-informed, and supported by expert judgment. Consensus thresholds help turn expert responses into interpretable study outputs.

They reduce ambiguity

Clear thresholds help avoid vague conclusions such as “the panel generally agreed.”

They improve transparency

Readers can see how the study team classified items as included, excluded, uncertain, or requiring revision.

They support reporting

Predefined rules make it easier to prepare manuscripts, reports, tables, figures, and supplemental materials.

They protect credibility

Defining thresholds before analysis reduces concern that rules were chosen to fit desired results.

They guide round decisions

Thresholds can help determine whether items move forward, are revised, or are dropped between rounds.

They clarify disagreement

Consensus rules can also help identify items where the panel remains divided or uncertain.

Consensus does not always mean unanimity

In Delphi studies, consensus rarely means that every participant agrees completely. Complex topics often involve some disagreement, even among experts. Instead, consensus usually means that responses meet a predefined level of agreement, importance, appropriateness, or stability.

For example, a study might classify an item as reaching consensus if a certain percentage of participants rate it as important, or if the median rating is high and responses are tightly clustered.

Unanimity

Everyone agrees. This is rare and usually not required in Delphi studies.

Consensus

The panel reaches a predefined level of agreement or stability that is strong enough for the study purpose.

Important distinction: A Delphi study can have meaningful consensus even when a minority of participants disagree.

Common ways to define consensus

There is no single consensus threshold that fits every Delphi study. The right approach depends on the study objective, field, rating scale, number of participants, stakeholder groups, and planned use of the results.

Percentage agreement

A common approach is to define consensus based on the percentage of participants who rate an item above or below a certain level.

Median rating

The study may require a median rating above a certain value, especially when using numerical rating scales.

Interquartile range

The interquartile range can show whether responses are tightly clustered or widely dispersed.

Ranking position

For prioritization studies, consensus may be based on whether an item falls within a top-ranked group.

Stability across rounds

Some studies examine whether responses change very little from one round to the next.

Stakeholder-group agreement

Some studies require agreement within multiple groups, such as clinicians, patients, researchers, or policy leaders.

Examples of consensus rules

The following examples show how a Delphi team might define consensus. These are not universal rules; they are examples of how thresholds can be structured.

Study goal Example threshold Possible interpretation
Identify important outcomes 70% or more rate the outcome as highly important Outcome may be included or advanced to the next stage
Evaluate recommendation statements Median agreement rating is high and disagreement is low Statement may be accepted as consensus
Prioritize research questions Item appears in the top-ranked group across the panel Question may be classified as a high priority
Assess appropriateness High rating plus narrow response spread Scenario may be classified as appropriate
Compare stakeholder groups Threshold met overall and within key groups Item may be considered broadly supported
Practical point: The threshold should match the purpose of the study. A clinical recommendation, a core outcome set, and a forecasting study may need different rules.

Choosing the right threshold

Choosing a consensus threshold is partly methodological and partly practical. A stricter threshold may produce a shorter list of items with stronger agreement. A more flexible threshold may retain more items but require careful interpretation.

Stricter thresholds may be useful when...

  • The final recommendations may influence clinical, policy, or high-stakes decisions.
  • The study needs a focused final list.
  • The panel is large enough to support stronger agreement criteria.
  • The team wants to minimize inclusion of borderline items.

More flexible thresholds may be useful when...

  • The study is exploratory or early-stage.
  • The goal is to identify possible priorities rather than final recommendations.
  • The panel is small or includes highly diverse stakeholder perspectives.
  • The topic is emerging and uncertainty is expected.
Balance matters: A threshold that is too strict may exclude useful items. A threshold that is too loose may make the final consensus less meaningful.

Consensus in vs. consensus out

Some Delphi studies define rules not only for including items, but also for excluding items. This can help the team manage long item lists and make round-by-round decisions more consistent.

Consensus in

The item meets the predefined threshold for inclusion, importance, agreement, appropriateness, or priority.

Consensus out

The item meets a predefined threshold for exclusion, low importance, disagreement, or lack of relevance.

Items that do not meet either threshold may be classified as uncertain, retained for another round, revised based on comments, or discussed by a steering committee depending on the study protocol.

Handling disagreement

Disagreement is not a failure. In many Delphi studies, disagreement is useful because it shows where experts, disciplines, regions, or stakeholder groups view the topic differently.

The study team should decide how disagreement will be identified and reported. This may include response distributions, comment summaries, stakeholder-group comparisons, or a category for items that did not reach consensus.

Wide distribution

Ratings are spread across the scale, suggesting that the panel is divided.

Stakeholder differences

One group supports an item while another group does not.

Comment themes

Open-text responses explain why participants interpreted an item differently.

Reporting tip: Items that do not reach consensus can be just as informative as items that do, especially when they reveal important uncertainty or stakeholder differences.

Consensus across stakeholder groups

Many Delphi studies include more than one type of participant. In medical research, this may include clinicians, researchers, patients, caregivers, and policy leaders. In other fields, it may include technical experts, educators, administrators, community representatives, or industry stakeholders.

When stakeholder groups are included, the study team should decide whether consensus will be assessed across the whole panel, within each stakeholder group, or both.

Approach What it shows Possible concern
Overall panel consensus Whether the full group meets the threshold Large groups may mask disagreement from smaller stakeholder groups
Within-group consensus Whether each stakeholder group agrees independently Small group sizes may make thresholds harder to interpret
Combined approach Whether an item is supported overall and across key groups Requires clear reporting and careful interpretation

When should thresholds be defined?

Consensus thresholds should generally be defined before the Delphi study begins. They may be documented in the protocol, study plan, statistical analysis plan, ethics submission, steering-committee materials, or manuscript methods.

Before launch Define rating scales, consensus thresholds, stakeholder-group rules, and item-retention logic.
During the study Apply the rules consistently when deciding what moves forward, is revised, or is removed.
After completion Report the thresholds clearly so readers understand how final conclusions were reached.
Best practice: Avoid choosing thresholds only after seeing the results. Predefined rules make the process more transparent and defensible.

How thresholds affect round-to-round decisions

Consensus thresholds can guide what happens after each round. They can help the team decide which items are accepted, removed, revised, or carried forward.

Accept

Items that meet the inclusion threshold may be accepted into the final list or no longer need re-rating.

Remove

Items that meet exclusion rules may be removed to reduce participant burden in later rounds.

Revise

Items with mixed responses or important comments may be reworded and presented again.

Retain

Items that are close to threshold may be carried forward for another round.

Separate

Items with different concepts may be split into clearer, more specific items.

Discuss

Some items may need steering-committee review, especially if stakeholder groups disagree.

Common threshold mistakes

Consensus thresholds are simple in concept, but they can create problems when they are unclear, poorly matched to the study objective, or inconsistently applied.

Defining thresholds too late

Rules chosen after seeing the data may appear less objective.

Using one rule for every purpose

Importance, agreement, feasibility, and appropriateness may require different interpretation.

Ignoring disagreement

A high percentage agreement may hide meaningful disagreement in a subgroup.

Over-relying on averages

A mean score alone may hide a split panel or polarized responses.

Not explaining uncertain items

Items that do not meet consensus should still be reported or classified clearly.

Changing rules mid-study

Any protocol changes should be documented and justified carefully.

Reporting consensus thresholds

The final report should make the consensus process easy to understand. Readers should be able to see how thresholds were defined, how they were applied, and how the final items were classified.

Methods section

Describe rating scales, thresholds, stakeholder-group rules, number of rounds, and item-retention logic.

Results tables

Show ratings, agreement percentages, medians, distributions, consensus status, and changes across rounds.

Narrative summary

Explain which items reached consensus, which did not, and what disagreement or uncertainty remained.

Helpful output: Clear tables and consensus-status labels can make the final report easier for steering committees, journal reviewers, and decision-makers to interpret.

How Surveylet supports consensus-oriented workflows

Surveylet was designed to support Delphi and Real-Time Delphi studies, including expert panels, multi-round workflows, structured feedback, response tracking, and consensus-oriented exports.

This can help research teams manage rating scales, round-by-round feedback, stakeholder groups, response summaries, and reporting workflows more efficiently than a general-purpose survey tool.

Structured ratings

Support rating scales, comments, item review, and organized Delphi questionnaires.

Feedback workflows

Help participants review group feedback and reconsider responses across rounds or in Real-Time Delphi designs.

Reporting exports

Export data for consensus analysis, steering-committee review, reports, manuscripts, and supplemental materials.

Need help defining consensus criteria?

Surveylet helps researchers and organizations manage Delphi and Real-Time Delphi studies online, including expert panels, rating scales, structured feedback, multiple rounds, and consensus-oriented reporting.