Line of inquiry
Inquiring lines›What determines reliable reasoning…›What determines LLM output consist…›this line of inquiry
How can we detect and account for LLM involvement in academic writing?
A broader line of inquiry — a family of 47 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 47
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Can readers reliably distinguish LLM-generated research writing from human writing?
- Can LLMs reliably assess the quality of ideas they generate?
- Can text-based algorithms reliably detect LLM assistance in scientific abstracts?
- How does LLM-modified writing narrow linguistic diversity in peer review?
- Can researchers detect individual papers modified by LLMs reliably?
- Does verification become the real bottleneck in LLM-assisted authorship?
- How does stylistic matching contribute to LLM self-preference in evaluations?
- Why do LLMs systematically prefer text from their own family?
- Did LLM use in scientific writing plateau after the initial surge?
- What heuristics do readers use to detect or fail to detect LLM writing?
- Do LLM adopters actually cite more diverse and younger research?
- Does rhetorical robustness across multiple LLM models predict stable scientific review?
- Do humans and LLMs agree on novelty assessment in research?
- How much does rhetorical framing shift LLM reviewer scores independent of content?
- Why do readers rate LLM-edited text more favorably?
- What structural barriers prevent LLMs from making evaluative judgments about writing?
- How much do LLM reviewers shift scores based on rhetorical framing alone?
- Can reviewer-author matching by LLM use amplify biases in acceptance decisions?
- Can researchers prevent their expectations from shaping LLM outputs?
- Do readers who prefer LLM-edited abstracts check the substance or just clarity?
- What methods can reliably detect LLM-generated academic papers at scale?
- Does exposure to LLM answers actually change how people think critically?
- Why does disclosure of LLM authorship change reader trust and preference?
- Can you detect LLM arguments by measuring convergence with the original post?
- How does framing critical topics shape LLM review scores?
- How do years of A/B testing compare to one-shot LLM content generation?
- Can forensic features reliably distinguish LLM arguments from human arguments?
- Why do computer science papers show more LLM modification than other fields?
- How do LLM reviewer scores respond when rewriting is applied recursively or jointly?
- How does the absence of evaluative stance appear in LLM academic writing?
- What quality differences exist between flagged and unflagged peer reviews?
- Why do some LLM clusters cite broader psychology than others?
- What linguistic features most strongly signal LLM authorship in counter-arguments?
- Can LLM persuasion be fairly evaluated without stratifying by reader background?
- Why did LLM-assisted writing plateau instead of continuing to grow?
- Can LLM-generated reference reviews detect machine-written peer review submissions?
- How do newer LLM generations differ from human writing patterns in detectable ways?
- What mechanisms drive rating compression in fully LLM-generated peer reviews?
- How much do LLM persuasiveness claims hide heterogeneous effects across different reader ideologies?
- Did reviewers successfully circumvent ICML's hidden-instruction watermark detection method?
- Does LLM use reduce writing costs differently across linguistic backgrounds?
- Do LLMs match top human creative writers in literary quality?
- What concerns does widespread LLM use raise for scientific independence?
- Can knowledge density explain why LLM writing feels coherent but fatiguing?
- Does personalized rubric training in one writer's case actually generalize?
- Do LLMs trained on Wikipedia content count as indirect readership?
- Did GPT-4 see the entire paper or only a portion of it?