What tasks do expert data storytellers trust to LLMs?
Expert visual data storytellers make strategic choices about which narrative work to delegate to LLMs and which to protect. Understanding these boundaries reveals how human judgment and automation can coexist in knowledge work.
The paper interviewed 12 expert visual data storytellers to ask which activities authors entrust to LLMs and which they "protect." Its central finding is that participants "rarely treated LLMs as autonomous storytellers." They "selectively delegate execution-oriented tasks" while "retaining control over activities that shape narrative intent and story meaning." Two further findings sit alongside this: LLM assistance is "most productive after human seeding and constraint-setting," and it "shifts labor from production to verification."
The discussion gives the reasoning behind the boundary. Delegation is described as "a nuanced negotiation" rather than a binary choice, with levels that vary by the specific activity, "perceived potential risks," and "desired authorial control." Authors are advised to decide up front which story components stay under human control, which the paper frames as preserving agency "without rejecting automation." Verification is the other half. Participants "treated model outputs as hypotheses to be checked rather than facts to be accepted," especially for claims, causal explanations, or audience-facing content. The paper therefore treats checking that statements and visualizations stay grounded in data, context, and intent as part of authorship, since "maintaining ownership requires" it.
This gives practitioner-side grounding to a pattern already in the library. The point that Can LLMs generate more novel ideas than human experts? argues that generation is cheap and evaluation stays human, and that human validation becomes the bottleneck. These interviews describe the same division of labor from the author's side, but the paper attributes the boundary to authors' own risk and control judgments, not to a claim about what LLMs cannot do. The verification stance also acts as a working counterweight to Does polished AI output trick audiences into trusting it?: authors who check outputs against data are declining to let a finished-looking result stand in for their judgment. The finding that help is most productive after seeding and constraint-setting is compatible with Can designers shape LLM behavior without deep technical knowledge?, where designer judgment is put in through structured prompt authoring before the model produces anything.
The excerpt does not say which tasks counted as "execution-oriented" or which as "narrative intent," how participants were recruited, or what tools they used. It reports no measure of time saved, output quality, or verification cost, so "shifts labor" is an interview-based description and not a measured trade-off. It also does not say whether the protected boundaries track real risk or only perceived risk. The supportable implication is narrow: for expert authors in this domain, tools that assume the LLM will run the story end to end do not match reported practice. Tools that help authors mark what they retain and that make grounding checkable match it better, which is the direction the paper's own design implications take.
Related concepts in this collection 3
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Can LLMs generate more novel ideas than human experts?
Research shows LLM-generated ideas score higher for novelty than expert-generated ones, yet LLMs avoid the evaluative reasoning that characterizes expert thinking. What explains this apparent contradiction?
the same production-to-evaluation split seen from practitioners' delegation choices rather than from LLM capability asymmetry
-
Does polished AI output trick audiences into trusting it?
When AI generates professional-looking graphs, diagrams, and presentations, do audiences mistake visual polish for analytical depth? This matters because appearance might substitute for actual expertise.
treating outputs as hypotheses to check is a practitioner counter to polished output standing in for judgment
-
Can designers shape LLM behavior without deep technical knowledge?
Explores whether LLMs can be treated as adaptable design materials that designers can tinker with directly, rather than fixed components handed over by engineers. Matters because it determines whether user-centered judgment reaches model adaptation early.
human seeding and constraint-setting parallels designers putting judgment in through system-prompt authoring
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Harnessing LLMs Without Surrendering Control: Delegation Boundaries in Visual Data Storytelling Authoring
- DATATALES: Investigating the use of Large Language Models for Authoring Data-Driven Articles
- Data Journalist Agent: Transforming Data into Verifiable Multimodal Stories
- Bridging the gulf of envisioning: Cognitive design challenges in llm interfaces.
- DiaSynth: Synthetic Dialogue Generation Framework for Low Resource Dialogue Applications
- Rethinking Interpretability in the Era of Large Language Models
- The LLM Fallacy: Misattribution in AI-Assisted Cognitive Workflows
- Demystifying Agent Skills: Why They Work-Until They Don't
Original note title
expert visual data storytellers delegate execution-oriented tasks to LLMs and keep narrative intent under human control — labor shifts to verification