Theme of inquiry

Why do models produce unreliable reasoning outputs?

A question within its area, explored through 7 lines of inquiry below — each a family of specific questions the research asks.


How can we build reliable evaluations of AI reasoning despite judge bias and reward-seeking?

40 specific questions

See all 40 questions in this line of inquiry
Does model confidence reliably signal actual accuracy in practice?

72 specific questions

See all 72 questions in this line of inquiry
How can we prevent synthetic data from contaminating statistical inference and corpora?

32 specific questions

See all 32 questions in this line of inquiry
How does self-revision in reasoning models affect accuracy and confidence?

55 specific questions

See all 55 questions in this line of inquiry
How does improved reasoning affect models' ability to acknowledge uncertainty?

49 specific questions

See all 49 questions in this line of inquiry
How do false presuppositions and sycophancy drive persistent false beliefs in models?

44 specific questions

See all 44 questions in this line of inquiry
How can we distinguish genuine model deception from honest errors?

44 specific questions

See all 44 questions in this line of inquiry