INQUIRING LINE

Tweak a prompt until the answer looks right, and the researcher's own judgment quietly becomes the measuring stick.

Can ad-hoc prompt engineering be treated as a standard research practice?

This explores whether tweaking prompts until they work, the way most people use LLMs, holds up as a legitimate scientific method, or whether it quietly undermines the results it produces.


This explores whether tweaking prompts until they work can count as a legitimate research method, or whether it quietly undermines the results it produces. The collection's short answer is no, at least not in its usual form. The main problem isn't a lack of rigor in some general sense. Iterating on a prompt changes what is being measured. When one researcher keeps revising a prompt until the output looks right, their own judgment becomes part of the instrument. The evaluation criteria also start to drift toward what the model happens to do well rather than what the task actually needs. That sets up a self-fulfilling loop: the prompt gets 'better' at producing what the researcher already expected to see Does iterative prompt engineering undermine scientific validity?. A separate line of work describes the same thing from the user's side. Refining a prompt works like a steering process that pulls the model's output toward the person's own expectations, so the result is co-written by the model and the person prompting it How much does the user shape what a model generates?.

The evidence for prompting techniques themselves is also weaker than it looks. When five well-known techniques were retested across six models and five benchmarks with proper statistical controls, none produced a significant improvement. The authors trace this to the same problems behind psychology's replication crisis: small samples, loose experimental design, and selective reporting Do popular prompting techniques actually improve model performance?. Part of the reason may be that prompt effects rarely carry over from one setting to another. In one benchmark of 23 prompts across 12 models, rephrasing helped cheap models, while step-by-step reasoning actually lowered accuracy in the strongest ones Do prompt techniques work the same across all LLM tiers?. Whether chain-of-thought helps can depend on the individual question, not just the type of task Why do some questions perform better without step-by-step reasoning?. A prompt's effect also depends on how answers are sampled afterward. Prompts tuned without accounting for techniques like majority voting or best-of-N underperform by up to 50% compared with tuning both together Does prompt optimization without inference strategy fail?. So a prompt 'that works' is really a finding about one model, one kind of question, and one way of sampling answers.

The collection also points to what a more standard practice could look like. One option is to fix the rules before looking at any outputs: set the criteria in advance, have several people code the results, and check how well they agree Does iterative prompt engineering undermine scientific validity?. A research-automation system called Spark-to-Paper builds the same idea into its design. It requires the evidence to be specified before results are seen, and it keeps the model's judgment calls separate from steps that can be checked mechanically Can separating judgment from verification improve research paper reliability?. Another option is to judge the prompt itself rather than its outputs. One framework scores prompts on six dimensions drawn from communication theory and instructional design, so a prompt can be assessed without a 'good result' to work backward from Can we measure prompt quality independent of model outputs?. Prompts can also borrow structure from established methods. Turning a classic framework for analyzing arguments into explicit steps forces the model to check the assumptions linking evidence to claims, which ad-hoc chain-of-thought tends to skip Can structured argument prompts make LLM reasoning more rigorous?.

The takeaway you might not expect is that the risk in ad-hoc prompting isn't only noise. Over time, the researcher can end up redefining the task to suit the model. Exploring by trial and error is fine for discovery. Once it becomes the method behind a published claim, it needs the same safeguards as any other measurement tool: criteria set in advance, more than one judge, and testing across models and settings.


Sources 9 notes

Does iterative prompt engineering undermine scientific validity?

Iterative prompt revision by single researchers introduces individual bias, shifts evaluation criteria to match LLM capabilities rather than task requirements, and creates self-fulfilling feedback loops. A validated pipeline with inter-coder reliability and pre-specified criteria is required instead.

How much does the user shape what a model generates?

Foundation Priors research shows prompt engineering as divergence minimization between synthetic output and user priors. The refinement process systematically steers generation toward what users already expect, making outputs co-productions of model and user subjectivity.

Do popular prompting techniques actually improve model performance?

Systematic testing of five prominent prompting techniques across six models and five benchmarks found no statistically significant improvements. The field faces methodological weaknesses identical to psychology's replication crisis: small samples, poor experimental design, publication bias, and selective reporting.

Do prompt techniques work the same across all LLM tiers?

A 23-prompt benchmark across 12 LLMs shows rephrasing and background-knowledge prompts boost cheap models, while step-by-step reasoning reduces accuracy in high-performance models. Task structure, not generic best practices, determines which prompts help.

Why do some questions perform better without step-by-step reasoning?

Saliency analysis reveals that CoT prompting fails when question information doesn't aggregate into the prompt structure before reasoning begins. For simple questions, direct question-to-answer flow outperforms step-by-step reasoning, showing the optimal prompt depends on question type, not just task category.

Show all 9 sources
Does prompt optimization without inference strategy fail?

Prompts optimized without knowledge of the inference strategy (best-of-N, majority voting) systematically underperform. Joint optimization of both prompt and inference strategy yields up to 50% improvement across reasoning and generation tasks.

Can separating judgment from verification improve research paper reliability?

Spark-to-Paper architects paper generation as composable skills that isolate model judgment from executable, verifiable operations and require evidence specification before results are observed, reducing dependence on model correctness for consistency.

Can we measure prompt quality independent of model outputs?

Research identifies six evaluable dimensions—Communication, Cognition, Instruction, Logic, Hallucination, and Responsibility—with 20 sub-criteria based on Grice, cognitive load theory, and instructional design. Improvements in one dimension cascade to others, revealing prompt quality as a structured space rather than a flat checklist.

Can structured argument prompts make LLM reasoning more rigorous?

Applying Toulmin's argument model as explicit prompting steps (CQoT) improves LLM reasoning by forcing models to identify warrants and backing rather than skipping implicit premises. The method catches failures that standard chain-of-thought prompting allows.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.