Consult Evaluation: Scottish Government's Non-surgical Cosmetic Procedures Consultation
Source: UK Incubator for AI (i.AI) · 2025-03-20
Consult is a tool for analysing public consultations, which uses AI to generate themes for each question and map responses to those themes. It alleviates the burden of running consultations by facilitating quicker analysis for a fraction of the cost.
➔ Consult was generally good at mapping responses to themes and produced near-identical rankings of themes compared to expert reviewers.
The overall performance (F1) score was 0.76 out of 1 (generally considered ‘good’) and reviewers made no changes to the Consult themes for 60% of responses.
Differences between Consult and reviewer-identified themes had negligible impact on the overall theme rankings – the key driver of policy recommendations.
➔ Reviewing Consult themes is quick, with a median time of 23s per response. By reducing analysis time, Consult freed up time to focus on the implications.
➔ Reviewers were impressed with the initial themes generated and felt that Consult helped reduce reviewer bias. However, they still wanted an opportunity to influence the themes, and the process for this was cognitively demanding and time intensive.
➔ Consult struggled to identify missing themes. The theme generation step may require more iteration, and working with the theme longlist is likely to help.
➔ Reviewers were keen to move to a world where only a share of responses need reviewing, but needed more confidence in the mapping before doing that.
Introduction. i.AI rapidly design, test and deliver AI products for government. We take pride in being a highly technical team, composed of experts with deep knowledge and experience from both the public and private sector. This allows us to tackle complex challenges and innovate continuously towards our overall mission: to harness the opportunity of AI for public good.
What is Consult?
Each year the government runs around 600 public consultations, with some receiving more than 100,000 responses. Analysing these consultations takes hundreds of thousands of hours of civil service time, or are contracted out at substantial cost. The time-intensive process also delays the findings and, ultimately, the development of policy.
Consult, an AI-powered tool developed by i.AI, aims to alleviate this burden by facilitating quicker analysis, of comparable quality, for a fraction of the cost. It does this using a two stage process:
1. ‘Theme Generation’: Identify common themes in the responses to generate a list of themes for each question.
2. ‘Theme Mapping’: Classify the full set of responses using one or more of these generated themes.
Consult has the potential to save substantial time and money for Government, but it is critical that this is not at the expense of high quality analysis that captures the public’s views. We are therefore carrying out several evaluations to ensure that Consult is suitable for wider use, and continually improving its delivery. We have conducted several retrospective evaluations using historic consultation data. This evaluation builds on our previous evaluations in several ways:
1. This is the first evaluation we have done on a live consultation, so allows us to better understand how Consult performs in practice.
2. It is the first evaluation we have done with the current version of Consult. This includes a new approach to reviewing the initial theme list (generation stage), and a new interface for reviewing mapped themes (mapping stage).
3. It addresses a new topic. In order to be confident in Consult’s use for consultation analysis in general, we need to evaluate it across a range of different topics.
Related work. How teams currently analyse consultation responses Conventional analysis of public consultations is a multi-stage process. Although the approach varies by consultation (for example, based on their size) it generally involves an initial review of a subset of responses to identify common themes. Once these are identified, a team is trained on the theme framework (to ensure they have a common understanding of each theme) before reading through all of the responses and mapping each response to the list of possible themes. At the end, a summary shows how frequently each theme appeared.
Method. Consult broadly follows two steps: theme generation and theme mapping. First, it uses a topic modelling approach to identify common themes for each question (using the whole dataset, rather than a sample). It then maps each response to the available themes using LLMs. We are developing a dashboard to summarise the findings, allowing policymakers to see how often each theme appeared, as well as more nuanced metrics such as the overlap between themes.
This two stage process allows us to include a ‘human in the loop’ at each stage. In the 1 first stage (theme generation), expert reviewers working on the policy area can review and approve the themes before they are used in the mapping process. This theme sign-off process is a new development that we were keen to test. For this evaluation, this stage involved four expert reviewers from the Scottish Government.
For the second stage (theme mapping), reviewers verified (and, if necessary, amended) the theme mappings that Consult provided for each response. Using the Consult application, reviewers were presented with an individual free text response to an individual question and the themes that Consult had mapped to it. If they deemed it necessary, reviewers made changes to the themes which Consult had mapped to the response using the tickboxes. Once they were content with the set of themes mapped to the response, they saved the mapping and continued to the next response. You can see a hypothetical example below. For this evaluation, this stage involved six expert reviewers, who split the responses between them.
Research questions This evaluation was focussed on answering three key questions:
1. How similarly do Consult and reviewers map themes to responses?
2. To what extent do any differences between Consult and the reviewers impact the headline findings of the consultation?
3. How could the process be improved?
Our approach combined quantitative analysis of the Consult and reviewer theme mappings, with user research covering both stages of the Consult process.
Once the assigned themes had been reviewed by our subject experts, we had a dataset that included (for each response and question number) the themes that were assigned by Consult, and the themes assigned by our reviewers.
Our main analysis aimed to answer the following questions:
1. How aligned were the Consult-mapped themes with the reviewer-mapped themes?
2. To what extent do differences between Consult and the reviewer themes impact the implied headline findings?
Discussion. Consult correctly identified all themes for three-fifths of responses and had generally good overall performance, with an average F1 score of 0.76 (out of 1).
For 60% of responses, the themes identified by Consult were unchanged by the reviewers – they were perfectly ‘correct’. Calculating performance across all responses (including those with a partial match) the mean multi-label F1 score was 0.76 out of 1, with a 95% confidence interval of (0.75, 0.77).
This is particularly promising given that there is no objective truth for what a correct theme is – this was based on the assessment of a single reviewer. The potential for subjectivity and human error both make it harder for Consult to be accurate (and mean that even reviewers would not achieve a perfect score). In a previous evaluation of Consult where we had two reviewers, the equivalent F1 score between the reviewers was 0.81 (95% CI: 0.78, 0.83) – only slightly higher than the performance of Consult. Similarly, the reviewers only mapped the exact same themes as each other 62% of the time (roughly the same as Consult’s performance on this evaluation).
Differences between Consult and the expert reviewers had minimal effects on the overall theme rankings.
From a policy perspective, identifying the top themes in a consultation is usually more important than knowing the exact number of times the theme appeared. Therefore we are more concerned if the differences in theme mapping lead to the top theme shifting to third place, than if the most common theme is consistent but counted 220 times rather than 300 times.
Reviewers were more likely to add themes than remove them, with changes concentrated on a few key themes.
Inspecting the specific changes made by the expert reviewer, themes were added 1,671 times, but only removed 763 times. Looking specifically at Question 2, removals were heavily concentrated on theme A, which accounted for more than half of all removals.
Additions were less heavily concentrated, but theme B was strongly over-represented and accounted for nearly a third of additions. Other questions showed similar patterns.
The fact that differences were driven by a small number of themes may explain the outlier themes that moved substantially in the overall ranking. It also suggests that certain themes are more ‘problematic’ in terms of how Consult and reviewers interpret them.
Future work may enable us to detect and correct these problem themes earlier in the process.
For 14% of responses reviewers added an ‘Other’ theme, suggesting our original theme list was not exhaustive and Consult struggled to make adjustments.
Both Consult and the reviewers were able to label responses as ‘No Reason Given’ or add ‘Other’ to indicate a relevant response where the reason was not captured in the current theme list. However, there were substantial differences in how these were used. Both Consult and the reviewers used the ‘No Reason Given’ label, but this was used 1.5 times more by Consult. In contrast, Consult very rarely assigned the ‘Other’ label (47 occurrences) while this was used nearly 17 times more by reviewers (nearly 800 occurrences). In one-third (32%) of cases where ‘Other’ was added, ‘No Reason Given’ Reviewers said they selected the ‘Other’ theme for one of three reasons:
1. They had identified a new theme they thought was important and had not been captured.
2. They needed a way to categorise responses with a negative sentiment.
3. Human error, i.e. creating an ‘Other’ theme because they had forgotten or missed a pre-existing theme.
Conclusion. Consult is an effective tool for analysing consultations:
● It identifies good themes in the data.
● Variation between Consult and reviewers was similar to the expected level of person-to-person difference.
● Differences between Consult and reviewer-identified themes had negligible impact on the overall theme rankings, implying similar qualitative interpretations.
● Consult speeds up analysis and allows the reviewers to focus more on the ‘so what’.
However, there are areas for improvement:
● There were discrepancies between Consult and the reviewer-mapped themes.
While this is partly due to the subjective nature of the task, it can impact reviewers’ trust in the process. The ability to review and alter theme mappings through the Consult app is important for maintaining that trust.
● It is hard to benchmark Consult’s performance, given theme mapping is inherently subjective. This would be easier if responses were double reviewed, which we will test in future evaluations.
● Consult rarely flags new themes (using the ‘Other’ theme) during the mapping stage. It is therefore important that the correct themes end up in the framework.
● The current theme sign-off process is cumbersome and would not scale well.
Focuses for future evaluations:
In our future rounds of evaluation, we will prioritise the following:
➔ Refining the theme sign-off process, focusing on usability and ensuring the correct themes end up in the framework. This could include mapping a small sample as part of the sign-off process.
➔ Testing whether amending specific themes can reduce the discrepancy between Consult and reviewers.
➔ Using two reviewers per response to benchmark Consult against person-to-person alignment.
Lines of inquiry this paper opens 19
Research framings built by reading the notes related to this paper — the questions it feeds into.
Do AI coding tools measurably improve developer productivity and code quality?- How do non-experts evaluate AI-generated outputs when they lack implementation expertise?
- Why haven't AI agents replaced human code review workflows?
- Can rubric-graded response quality predict real-world clinical workflow success?
- Do patients show the same bias toward expert-labeled medical advice?
- Would clinicians' ratings change if authorship was visible from the start?
- Why did clinicians guess authorship at chance level despite strong preferences?
- Do physicians follow incorrect advice more when they trust its source?
- How much diagnostic accuracy is gained when physicians receive expert advice?
- Why does labeling advice as AI from a doctor change how people trust it?
- Can people tell which medical advice is accurate based only on how it reads?
- Can taste and judgment become the scarce resource in AI-assisted work?
- Can AI systems generate policy themes as well as humans can map them?