Further Explorations on the Use of Large Language Models for Thematic Analysis. Open-Ended Prompts, Better Terminologies and Thematic Maps

Paper · Source
Reading and SummarizationDomain Specialization in LLMs

There is a nascent area, where scholars are approaching thematic analysis (TA) using LLMs, following the six phases developed by BRAUN and CLARKE (2006). TA is a qualitative method of analysis where the researcher labels (codes) portions of data with relevant meaning and then organises these codes/labels into patterns (the themes). BRAUN and CLARKE stipulated that TA encompasses the following phases: 1. familiarisation with the data; 2. initial coding; 3. identification of themes; 4. revision of themes; 5. renaming and summarising of themes; and 6. write-up of the results.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

How do evaluation biases undermine LLM quality assessment systems? How should dialogue systems best leverage conversation history for retrieval? What makes specific clarifying questions more effective than generic ones? How do formal dialogue structures reveal conversation coherence mechanisms? How faithfully do LLMs reflect their actual reasoning in outputs and explanations? Do language models learn genuine linguistic structure or just surface patterns? Do reasoning traces faithfully represent or merely mimic actual model reasoning? How should retrieval systems optimize for multi-step reasoning during inference? How do LLMs distinguish causal reasoning from temporal and semantic associations? How do training data properties shape reasoning capability development? How do neural networks separate factual knowledge from reasoning abilities? How should human oversight be integrated with autonomous AI systems? Why does verification consistently lag behind AI generation? Does domain specialization cause models to lose capabilities elsewhere? Does fine-tuning modify underlying model capabilities or only behavioral outputs? What prevents language models from reliably adopting diverse personas? Do harness improvements transfer across model scales or memorize shortcuts? How should models express uncertainty rather than forced confident answers?