AI and Elections: How Well Do AI Platforms Answer Voter Questions?

Paper · Source
Knowledge After the Web

Source: States United Democracy Center · 2026-08-04

A States United study before the 2026 midterms finds that ChatGPT and Google AI offer incomplete and inconsistent responses to common voter questions, even as the accuracy of AI platforms improves.

Across two rounds of empirical testing in late 2025 and early 2026, our research found that while AI platforms have demonstrated improved accuracy over time, they are incomplete substitutes for official election information; inconsistent in directing voters to authoritative state sources; and vulnerable to format changes driven, in part, by commercial pressure.

Improved accuracy over time. In the first preliminary round, the error rate for responses from Google AI and ChatGPT was 6.9% and 8.2%, respectively. In the second primary round, the rate of verifiable factual errors fell to 0% across both platforms. The accuracy of responses from the AI platforms improved. However, the chatbot responses did not provide everything a voter needs to know. Our finding of improved accuracy comes with a caveat: technically accurate responses were often incomplete, failing to direct voters to official sources that would further enable their participation. Responses with an improved 0% error rate still came with information gaps, such that the improvements were not always complete, current, or usable.

Inconsistent direction of voters to official state websites. The single most important measure of voter utility examined in this study was whether AI platforms directed voters to their state’s official election website. AI platforms guided voters to state websites less than 50% of the time, and for some questions, they almost never did. ChatGPT mentioned a state election site in 39.4% of its responses, and Google AI did so in 55.6% of its responses.

Incomplete answers about candidate listings. When asked who is running for a particular office, AI platforms produced incomplete answers at strikingly high rates, the most of any question type. ChatGPT returned incomplete candidate lists for 88.9% of the gubernatorial queries. Arizona, with its large and actively changing candidate field, drove most of the incompleteness. But in this study, AI systems trained on historical data and reliant on periodic retrieval were structurally ill-suited to track real-time candidate filings.

Overreliance on Wikipedia for election information sourcing. Wikipedia accounted for 12.3% of all 3,481 cited links across the primary study and dominated the sourcing of candidate questions. This finding is a meaningful concern given that Wikipedia is openly editable and not authoritative for time-sensitive or jurisdiction-specific election information.

Vulnerability to format changes. Google AI changed its output format in the course of the study, replacing written summaries with lists of links. The links-only responses, beginning on or around Feb. 2, 2026, offered no prose, guidance, or consistent recommendations to visit official state sources. This change was applied unevenly across query topics and appeared more pronounced for election-related queries than others. When asked why its output changed, Google AI attributed the change to efforts to improve accuracy and reduce hallucinations, while also describing the change as part of making AI search more monetizable.

The bottom line—based on two rounds of empirical research and every platform condition tested—is that voters seeking election information on leading AI platforms are inconsistently directed to the authoritative information on their state’s official election website. Although AI tools are advancing, they should be seen as a complement to, rather than a replacement for, official voter information.

Across both ChatGPT and Google AI, the factual accuracy improved meaningfully between rounds. In the preliminary investigation, 6.9% and 8.2% of the responses from Google AI and ChatGPT, respectively, contained verifiable factual errors. In the primary investigation, conducted several months after the preliminary investigation, no responses across any platform conditions were coded as inaccurate. This is real progress and should be acknowledged.

It is worth being precise about what a 0% error rate does and does not establish. It means that for the questions we tested, we found no verifiable factual claim to be wrong. It does not mean, however, that the responses were complete: ChatGPT returned incomplete candidate lists in 88.9% of gubernatorial queries. A 0% error rate also does not mean voters were reliably pointed to the office that could confirm the answer: ChatGPT mentioned a state election site in 39.4% of its responses, and after Google changed its output format, Google AI did so in none of its responses. And it does not mean the results hold beyond the conditions we tested, since we asked a defined set of questions in round two about three states over a period of weeks. The improvement in accuracy is real; however, it does not make them ready to serve as a voter’s primary source for election information.

Whether AI platforms direct voters to their state’s official election website is the single most important measure of voter utility examined in this study. State election websites are maintained by the election officials who administer the process, and they are updated as rules, deadlines, and candidate filings change. They are the source best positioned to give voters the deadlines, polling places, and candidate filings that apply to them, because the office publishing the information is responsible for it. According to our data, AI platforms guide voters to these sites less than 50% of the time, and for some questions, they almost never do so.

When voters asked who was running for a particular office, AI platforms produced incomplete answers at strikingly high rates. ChatGPT returned an incomplete list of declared candidates in 88.9% of the gubernatorial candidate queries, 35.7% of the attorney general queries, and 28.6% of the secretary of state queries. Google AI (Incognito) summary responses showed a 44.4% incompleteness rate on governor and 42.9% on both attorney general and secretary of state queries. These are not failures of accuracy in the traditional sense because the AI identified real candidates. However, the problem was that the platforms did not name all of them.

This is a structural limitation, not a temporary one. Candidate fields are unique across states and change with every filing deadline, withdrawal, and announcement. AI systems that retrieve information periodically rather than continuously will continue to lag the actual ballot. Compounding the problem: When AI answered candidate questions, it relied heavily on Wikipedia (see Finding 4) rather than on official state candidate filing records. Direct state election site links were provided in 0% of candidate responses across nearly all conditions tested. This means that if voters ask AI who was running, they will rarely be pointed to the place that would tell them definitively.

Across the primary study, AI platforms cited 3,481 links in total. State and local government sources accounted for the largest share (1,445 combined links, or 41.5%) and trustworthy non-profits for the next largest (822 links, or 23.6%). The remaining 35% of links are where the analytical concern lies. Wikipedia alone accounted for 429 of those sources, representing 12.3% of all the links and was the most heavily relied-upon source for candidate questions specifically. The Mixed and Unreliable category includes mixed-content platforms like Reddit, YouTube, and other third-party sites that are not vetted by election officials, accounting for another 264 links, or 7.6%. Taken together, Wikipedia and Mixed and Unreliable sources made up roughly one in five links that AI platforms surfaced when voters asked election questions.

What these two categories share is that voters cannot reliably evaluate the content on the other end of the click. Wikipedia is broadly accurate on many topics, but it is openly editable and the qualifications of contributors cannot be verified, and entries on time-sensitive subjects such as filing deadlines, candidate lists, polling procedures, ballot-counting rules may lag official sources or contain errors volunteer editors have not yet caught. Mixed-content platforms have a different problem: State election officials post verified informational videos to YouTube and answer voter questions on Reddit, but so do individuals spreading election mis- and disinformation, and an AI response that points a voter to either platform gives no indication which is which. When a voter sees a YouTube link in an AI-generated answer, they have no way of knowing whether the destination is a secretary of state’s verified channel or an unverifiable creator.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

How can AI systems reliably guide voters without introducing political bias? Why do confident AI outputs mislead human trust calibration? What gaps exist between benchmark performance and real deployment outcomes? How do AI hiring systems affect authenticity, fairness, and candidate preferences?