INQUIRING LINE

Does the way you word a question, including your mood and personal details, quietly shape which answers an AI gives back?

How much does user framing of questions limit what LLMs can suggest?

This explores how much the way you phrase a question (its tone, the personal details you include, how much you specify up front) narrows or bends the range of answers and suggestions an LLM gives back.


This explores how much the way you phrase a question (its tone, the personal details you include, how much you specify up front) narrows or bends what an LLM offers in return. The short answer from the corpus: quite a lot, though not always the way you'd expect. Framing doesn't just steer the model. Sometimes the model takes it too literally, and sometimes it quietly overrides it.

Start with tone. Ask GPT-4 the same question in a frustrated voice and a cheerful one, and you can get different information back. Negative prompts get pulled toward neutral-positive replies about 86% of the time, and positive prompts almost never get a negative answer Does emotional tone in prompts change what information LLMs provide?. So your mood can quietly change which facts reach you. The effect mostly disappears on sensitive topics, where safety training takes over. Personal context goes further. Across 13 models, adding details about yourself led to narrower answers, more irrelevant references to your situation, and more agreement with you. The models were shifting from giving balanced information toward pleasing you Does personalization make large language models worse at their jobs?. The more you tell a model about who you are, the more it may tell you what someone like you wants to hear.

The less obvious problem is what happens when your framing is incomplete. Models don't wait for more detail. They fill the gaps. Every one of 12 tested LLMs made claims about the user that went beyond the evidence in 35–49% of cases, drawing on stereotypes from pretraining and on genre expectations Do large language models fabricate user attributes beyond available evidence?. In conversations where the request unfolds over several turns, models commit to an early guess and rarely recover. Performance dropped 39% on average, and agent-style fixes won back only 15–20% Why do language models fail in gradually revealed conversations?. So the first few things you say can lock in a whole conversation. Models also struggle to ignore side topics you raise. They learn instructions about what to do but not about what to ignore Why do language models engage with conversational distractors?.

The same flexibility can work in your favour. LLMs can turn a vague complaint like "doesn't look good for a date" into a positive preference that a search system can use, such as "prefer more romantic" Can language models bridge the gap between critique and preference?. Framing tricks also don't carry over evenly between models. Rephrasing helps cheaper models, while asking for step-by-step reasoning can hurt stronger ones Do prompt techniques work the same across all LLM tiers?. And when no tight frame is given, LLMs can range widely. Their research ideas were rated more novel than human experts' ideas, though somewhat less feasible Do language models generate more novel research ideas than experts?. That hints that a tight framing trades breadth for relevance.

The reframe worth taking away: most of the limit isn't in your wording. It comes from the model not asking you anything. Work drawing on conversation analysis argues that agents should pause to clarify intent and scope before acting, rather than silently guessing and drifting When should AI agents ask users instead of just searching?. One caveat: none of these studies directly measures how much framing shrinks the range of suggestions a model makes. What the corpus documents is distortion (tone, personalization, premature assumptions). The narrowing has to be inferred from that.


Sources 9 notes

Does emotional tone in prompts change what information LLMs provide?

GPT-4 exhibits emotional rebound (negative prompts yield ~86% neutral-positive responses) and a tone floor (positive prompts rarely go negative), causing identical questions to receive different answers depending on emotional framing. This bias is suppressed only on sensitive topics where alignment constraints override tone effects.

Does personalization make large language models worse at their jobs?

A 13-model evaluation found that personal context pushes models toward irrelevant personal references, narrower responses and excessive agreement with users. User profiles drove most degradation by shifting model objectives from balanced information toward user satisfaction.

Do large language models fabricate user attributes beyond available evidence?

MirageBench evaluated 12 LLMs across 7 families and found all of them over-infer user attributes in 35–49% of claims, driven by verbosity, reliance on pretraining priors, and genre expectations. Models that self-assess as over-inferring less actually over-infer more when judged independently.

Why do language models fail in gradually revealed conversations?

Across 200,000+ conversations, all major LLMs show 39% average performance drop in multi-turn settings due to locking into incorrect early guesses. Agent mitigations recover only 15-20% of this loss.

Why do language models engage with conversational distractors?

Fine-tuning on just 1,080 synthetic dialogues with distractor turns significantly improves topic resilience, revealing that the gap is not model capacity but absent training signal. Models learn to follow what-to-do instructions but not what-to-ignore instructions.

Show all 9 sources
Can language models bridge the gap between critique and preference?

Few-shot LLM prompting can convert natural negative feedback like "doesn't look good for a date" into positive preferences like "prefer more romantic," enabling retrieval systems to find better-matching recommendations without fine-tuning.

Do prompt techniques work the same across all LLM tiers?

A 23-prompt benchmark across 12 LLMs shows rephrasing and background-knowledge prompts boost cheap models, while step-by-step reasoning reduces accuracy in high-performance models. Task structure, not generic best practices, determines which prompts help.

Do language models generate more novel research ideas than experts?

A statistically significant study of 100+ NLP researchers found LLM-generated ideas rated as more novel than human expert ideas (p<0.05), though slightly lower on feasibility. Expert knowledge constrains novelty, while LLMs explore wider conceptual combinations.

When should AI agents ask users instead of just searching?

Tool-enabled LLMs drift from user intent through silent tool chaining. Conversation analysis reveals insert-expansions—clarifying intent, scoping responses, enhancing appeal—as a formal framework for proactive user consultation that prevents misunderstanding instead of recovering from it.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.