Theme of inquiry
How do prompts and framing affect LLM outputs?
A question within its area, explored through 6 lines of inquiry below — each a family of specific questions the research asks.
35 specific questions
- Do LLMs achieve similar persuasive outcomes through different rhetorical mechanisms than humans?
- What role does stylistic convergence play in LLM persuasion effectiveness?
- How susceptible are language models to rhetorical pressure during debates?
- How does rhetorical familiarity bias models toward their own arguments?
- Do LLMs mirror the style of text they are prompted to respond to?
- Do LLMs address the prompter but persuade the public differently?
- Can lightweight linguistic features reliably detect LLM generated arguments?
91 specific questions
- Should LLM reasoning be studied as latent state trajectories rather than surface text?
- Does LLM reasoning always match the outputs it generates?
- Why does LLM knowledge fail to influence their actual outputs?
- How do knowing and doing diverge in LLM decision-making?
- How faithful are natural language explanations from LLMs really?
- When should an LLM engage extended reasoning versus responding directly?
- Can LLMs improve at simple deduction through different training approaches?
16 specific questions
- Why do research ideation systems suffer from diversity collapse despite high novelty metrics?
- Why do LLM-generated ideas score higher novelty yet lower feasibility than expert ideas?
- Why does diversity collapse occur in multi-agent research ideation despite high novelty?
- Can LLMs generate more novel research ideas than human experts?
- Why do LLM research ideas lack diversity despite high average novelty?
- Why does LLM research ideation collapse into low diversity despite high novelty?
- Can LLM diversity collapse in research ideation be reversed or mitigated?
17 specific questions
- Why do LLMs generate novel ideas but struggle to evaluate them?
- Do LLMs generate more novel ideas than they can evaluate?
- How can LLMs evaluate their own creative outputs for utility and novelty?
- Why do LLMs excel at generation but struggle with evaluation?
- What makes novelty assessment harder to automate than idea generation?
- Why do models generate creative ideas but fail to evaluate their legitimacy?
- What structural barriers prevent LLMs from making evaluative judgments about writing?
43 specific questions
- Why do language models presume common ground rather than build it?
- Why do language models presume common ground instead of building it?
- Does social grounding in language improve through iterative human integration?
- Can language models develop genuine social grounding through human interaction?
- Can LLMs use implicit background knowledge the way humans do in ordinary conversation?
- How do LLMs differ from humans in their grounding mechanisms?
- Can convention formation improve communicative grounding beyond word sharing?
32 specific questions
- Can LLMs reliably assess the quality of ideas they generate?
- Can language models accurately evaluate the quality of their own ideas?
- Does LLM judge preference for LLM arguments amplify errors in contested factual domains?
- Why do LLM outputs match researcher priors without solving tasks correctly?
- Can parallel evaluation reduce position and length bias in LLM judging?
- What other evaluation biases exist in LLM judge systems?
- Can researchers prevent their expectations from shaping LLM outputs?