SYNTHESIS NOTE
Topics›Expertise in the Age of AI Content›this note

Does disclosing AI assistance make readers trust articles less?

When articles carry a label saying they used AI tools, do human and AI raters downgrade their quality assessments? This matters because writers worry disclosure could harm how their work is received.

Synthesis note · 2026-10-06 · sourced from Expertise in the Age of AI Content

The central finding is that an AI disclosure costs an article a little with both kinds of rater. The authors ran a pre-registered survey with 1,970 human participants and collected 2,520 LLM ratings, all of one human-written news article. Both groups penalized the AI-disclosed version. The authors call the penalty an "AI disclosure discount" and say it "carries epistemic stigmatization," but they size it at "less than 0.15 on a 7-point scale across all experiments," which they describe as "a perceptible yet not overwhelming penalty."

The design is a 2×3×3 between-subjects factorial: disclosure present or absent, author race (Asian, Black, White), and author gender (man, woman, non-binary), giving eighteen conditions. The article was identical across conditions. The control line was a statistical update ("Statistical information updated as of Oct. 11, 2024."), and the treatment added "This article was created with assistance from Artificial Intelligence (AI) tools." Human raters scored information trustworthiness, comprehensiveness, writing quality, and likelihood of sharing on 7-point Likert scales. The mechanism is interpretive. The introduction frames readers as wanting to "calibrate judgments to discern the boundary between human insight and synthetic fluency," and the authors read the penalty as stigma attached to the disclosure itself. The excerpt does not test that reading against an alternative.

Against the nearest notes, this is the same shape as Does telling people an AI wrote something actually stop them from believing it?: disclosure changes the judgment without blocking it. The outcomes differ, though. That note measures sway in an argument, while this study measures ratings of one news article on four scales, so the figures are not comparable. The authors say their result "aligns with prior work showing that AI disclosure has a statistically significant but relatively modest impact on perception." The introduction also cites Baek et al. for labeling that "can reduce perceptions of credibility, creativity, and shareability." The excerpt does not say why the two findings differ in size. The writers' side of the trade-off is in Do writers want to see each other's AI prompts in shared editors?: collaborators want to see when AI was used, and the introduction notes that writers "may hesitate to disclose" for fear of how their work will be perceived. This study measures the audience cost of that disclosure, and finds it small.

The excerpt does not establish how the LLM raters were prompted, which models were used beyond the two named later in the discussion, how the human sample was recruited, or the statistical tests behind "consistent" and "pronounced." It gives no per-dimension results. It also does not show whether the penalty depends on genre, which the authors list as open. The Method says only the biography and disclosure language varied, while Phase 1 also lists a photo. The implication is narrow. The number measures a rating cost for one disclosure sentence on one news article. It is not a general discount on AI-assisted writing, and it should not be read as the size of the effect in publishing or hiring.

Inquiring lines that read this note 38

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Are AI-generated articles systematically disadvantaged in search ranking and user engagement? Can readers reliably distinguish AI-written text from human writing? How do writers navigate authorship and delegation with AI? Does disclosing AI authorship change how audiences evaluate the writing? How reliably can humans and AI detectors identify machine-generated text? How do educators verify student capability when AI can produce indistinguishable work? Why do confident AI outputs mislead human trust calibration? Do restrictions on reviewer LLM use actually shape peer review behavior? How do clinicians calibrate trust in AI medical recommendations?

Related concepts in this collection 6

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
18 direct connections · 117 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

disclosed AI assistance lowers ratings from human and LLM raters alike, but by less than 0.15 on a 7-point scale