Consensus Thresholds in Delphi Studies
How study teams can define, measure, interpret, and report consensus clearly in Delphi and Real-Time Delphi studies.
What is a consensus threshold?
In a Delphi study, a consensus threshold is the rule used to decide whether an item has reached sufficient agreement. It helps the study team determine which recommendations, outcomes, definitions, priorities, or criteria should be retained, revised, excluded, or reviewed further.
Consensus thresholds are important because they make the interpretation of results more transparent. Instead of saying that “most experts agreed,” the study can explain exactly how agreement was measured and what level of agreement was required.
A consensus threshold should usually be defined before the study begins, not after the results are known. This helps protect the credibility of the study and makes the final conclusions easier to defend.
Why consensus thresholds matter
Delphi studies are often used when decisions need to be transparent, evidence-informed, and supported by expert judgment. Consensus thresholds help turn expert responses into interpretable study outputs.
They reduce ambiguity
Clear thresholds help avoid vague conclusions such as “the panel generally agreed.”
They improve transparency
Readers can see how the study team classified items as included, excluded, uncertain, or requiring revision.
They support reporting
Predefined rules make it easier to prepare manuscripts, reports, tables, figures, and supplemental materials.
They protect credibility
Defining thresholds before analysis reduces concern that rules were chosen to fit desired results.
They guide round decisions
Thresholds can help determine whether items move forward, are revised, or are dropped between rounds.
They clarify disagreement
Consensus rules can also help identify items where the panel remains divided or uncertain.
Consensus does not always mean unanimity
In Delphi studies, consensus rarely means that every participant agrees completely. Complex topics often involve some disagreement, even among experts. Instead, consensus usually means that responses meet a predefined level of agreement, importance, appropriateness, or stability.
For example, a study might classify an item as reaching consensus if a certain percentage of participants rate it as important, or if the median rating is high and responses are tightly clustered.
Unanimity
Everyone agrees. This is rare and usually not required in Delphi studies.
Consensus
The panel reaches a predefined level of agreement or stability that is strong enough for the study purpose.
Common ways to define consensus
There is no single consensus threshold that fits every Delphi study. The right approach depends on the study objective, field, rating scale, number of participants, stakeholder groups, and planned use of the results.
Percentage agreement
A common approach is to define consensus based on the percentage of participants who rate an item above or below a certain level.
Median rating
The study may require a median rating above a certain value, especially when using numerical rating scales.
Interquartile range
The interquartile range can show whether responses are tightly clustered or widely dispersed.
Ranking position
For prioritization studies, consensus may be based on whether an item falls within a top-ranked group.
Stability across rounds
Some studies examine whether responses change very little from one round to the next.
Stakeholder-group agreement
Some studies require agreement within multiple groups, such as clinicians, patients, researchers, or policy leaders.
Examples of consensus rules
The following examples show how a Delphi team might define consensus. These are not universal rules; they are examples of how thresholds can be structured.
| Study goal | Example threshold | Possible interpretation |
|---|---|---|
| Identify important outcomes | 70% or more rate the outcome as highly important | Outcome may be included or advanced to the next stage |
| Evaluate recommendation statements | Median agreement rating is high and disagreement is low | Statement may be accepted as consensus |
| Prioritize research questions | Item appears in the top-ranked group across the panel | Question may be classified as a high priority |
| Assess appropriateness | High rating plus narrow response spread | Scenario may be classified as appropriate |
| Compare stakeholder groups | Threshold met overall and within key groups | Item may be considered broadly supported |
Choosing the right threshold
Choosing a consensus threshold is partly methodological and partly practical. A stricter threshold may produce a shorter list of items with stronger agreement. A more flexible threshold may retain more items but require careful interpretation.
Stricter thresholds may be useful when...
- The final recommendations may influence clinical, policy, or high-stakes decisions.
- The study needs a focused final list.
- The panel is large enough to support stronger agreement criteria.
- The team wants to minimize inclusion of borderline items.
More flexible thresholds may be useful when...
- The study is exploratory or early-stage.
- The goal is to identify possible priorities rather than final recommendations.
- The panel is small or includes highly diverse stakeholder perspectives.
- The topic is emerging and uncertainty is expected.
Consensus in vs. consensus out
Some Delphi studies define rules not only for including items, but also for excluding items. This can help the team manage long item lists and make round-by-round decisions more consistent.
Consensus in
The item meets the predefined threshold for inclusion, importance, agreement, appropriateness, or priority.
Consensus out
The item meets a predefined threshold for exclusion, low importance, disagreement, or lack of relevance.
Items that do not meet either threshold may be classified as uncertain, retained for another round, revised based on comments, or discussed by a steering committee depending on the study protocol.
Handling disagreement
Disagreement is not a failure. In many Delphi studies, disagreement is useful because it shows where experts, disciplines, regions, or stakeholder groups view the topic differently.
The study team should decide how disagreement will be identified and reported. This may include response distributions, comment summaries, stakeholder-group comparisons, or a category for items that did not reach consensus.
Wide distribution
Ratings are spread across the scale, suggesting that the panel is divided.
Stakeholder differences
One group supports an item while another group does not.
Comment themes
Open-text responses explain why participants interpreted an item differently.
Consensus across stakeholder groups
Many Delphi studies include more than one type of participant. In medical research, this may include clinicians, researchers, patients, caregivers, and policy leaders. In other fields, it may include technical experts, educators, administrators, community representatives, or industry stakeholders.
When stakeholder groups are included, the study team should decide whether consensus will be assessed across the whole panel, within each stakeholder group, or both.
| Approach | What it shows | Possible concern |
|---|---|---|
| Overall panel consensus | Whether the full group meets the threshold | Large groups may mask disagreement from smaller stakeholder groups |
| Within-group consensus | Whether each stakeholder group agrees independently | Small group sizes may make thresholds harder to interpret |
| Combined approach | Whether an item is supported overall and across key groups | Requires clear reporting and careful interpretation |
When should thresholds be defined?
Consensus thresholds should generally be defined before the Delphi study begins. They may be documented in the protocol, study plan, statistical analysis plan, ethics submission, steering-committee materials, or manuscript methods.
How thresholds affect round-to-round decisions
Consensus thresholds can guide what happens after each round. They can help the team decide which items are accepted, removed, revised, or carried forward.
Accept
Items that meet the inclusion threshold may be accepted into the final list or no longer need re-rating.
Remove
Items that meet exclusion rules may be removed to reduce participant burden in later rounds.
Revise
Items with mixed responses or important comments may be reworded and presented again.
Retain
Items that are close to threshold may be carried forward for another round.
Separate
Items with different concepts may be split into clearer, more specific items.
Discuss
Some items may need steering-committee review, especially if stakeholder groups disagree.
Common threshold mistakes
Consensus thresholds are simple in concept, but they can create problems when they are unclear, poorly matched to the study objective, or inconsistently applied.
Defining thresholds too late
Rules chosen after seeing the data may appear less objective.
Using one rule for every purpose
Importance, agreement, feasibility, and appropriateness may require different interpretation.
Ignoring disagreement
A high percentage agreement may hide meaningful disagreement in a subgroup.
Over-relying on averages
A mean score alone may hide a split panel or polarized responses.
Not explaining uncertain items
Items that do not meet consensus should still be reported or classified clearly.
Changing rules mid-study
Any protocol changes should be documented and justified carefully.
Reporting consensus thresholds
The final report should make the consensus process easy to understand. Readers should be able to see how thresholds were defined, how they were applied, and how the final items were classified.
Methods section
Describe rating scales, thresholds, stakeholder-group rules, number of rounds, and item-retention logic.
Results tables
Show ratings, agreement percentages, medians, distributions, consensus status, and changes across rounds.
Narrative summary
Explain which items reached consensus, which did not, and what disagreement or uncertainty remained.
How Surveylet supports consensus-oriented workflows
Surveylet was designed to support Delphi and Real-Time Delphi studies, including expert panels, multi-round workflows, structured feedback, response tracking, and consensus-oriented exports.
This can help research teams manage rating scales, round-by-round feedback, stakeholder groups, response summaries, and reporting workflows more efficiently than a general-purpose survey tool.
Structured ratings
Support rating scales, comments, item review, and organized Delphi questionnaires.
Feedback workflows
Help participants review group feedback and reconsider responses across rounds or in Real-Time Delphi designs.
Reporting exports
Export data for consensus analysis, steering-committee review, reports, manuscripts, and supplemental materials.
Need help defining consensus criteria?
Surveylet helps researchers and organizations manage Delphi and Real-Time Delphi studies online, including expert panels, rating scales, structured feedback, multiple rounds, and consensus-oriented reporting.
