Generative AI Uses and Risks for Knowledge Workers in a Science Organization
Generative AI could enhance scientific discovery by supporting knowledge workers in science organizations. However, the realworld applications and perceived concerns of generative AI use in these organizations are uncertain. In this paper, we report on a collaborative study with a US national laboratory with employees spanning Science and Operations about their use of generative AI tools. We surveyed 66 employees, interviewed a subset (N=22), and measured early adoption of an internal generative AI interface called Argo lab-wide. We have four findings: (1) Argo usage data shows small but increasing use by Science and Operations employees; Common current and envisioned use cases for generative AI in this context conceptually fall into either a (2) copilot or (3) workflow agent modality; and (4) Concerns include sensitive data security, academic publishing, and job impacts. Based on our findings, we make recommendations for generative AI use in science and other organizations.
Introduction. Generative AI, as a new core technology, has demonstrated significant potential to enhance scientific discovery by supporting knowledge workers in science organizations and institutions (e.g., government and industry research labs, universities, think tanks) [1, 49, 50]. However, the real-world applications and perceived risks of generative AI use in these organizations are uncertain. If generative AI could make science organizations with science- and operations-focused employees more efficient and speed up time to discovery on topics such as drug development and climate solutions, it would have important consequences for facilitating scientific gains to help society at large. Prior literature has looked at generative AI as a tool for scientific research, including the development of science-specific large language models [50, 55] and driving complex scientific tasks [10, 26, 39, 45]. This research, however, focuses on science research tasks and does not study the science workplace more broadly. Another area of literature looks at generative AI in the workplace, with a particular focus on professional knowledge workers [3, 12, 16, 21, 43, 48, 59]. Very little research, however, studies generative AI opportunities for scientists as knowledge workers [30], and to the best of our knowledge prior research has not studied generative AI for both science- and operations-focused workers in a science organization. In addition, prior work has investigated generative AI risks and concerns (e.g., [12, 30, 34, 37]), however, there has not been a focus on concerns that are specific to the novel context of a science organization. In this paper, we add to the literature by reporting on an investigation of the practical, real-world applications and perceived risks of generative AI use (focusing predominately on large language models as opposed to image-generating models) across a multidisciplinary science and engineering research center with the primary mission to deliver scientific progress. To conduct our study, we collaborated with an Information Technology (IT) group at a US national laboratory.1 The national lab, Argonne, we worked with employs several thousand people and is tasked with basic and applied science and engineering research. The lab includes Science divisions working to publish academic papers in research areas such as climate science, materials science, high performance computing, and more. Science teams also include software engineers, data scientists, and engineers who specialize in building and operating scientific experiments. In addition, the lab has multiple Operations divisions that employ people in areas such as IT, Human Resources (HR), Finance, Communications, Facilities (e.g., grounds crew and building maintenance), the Fire Department, and Security. We studied generative AI perceptions and uses for employees spanning Science and Operations roles and measured early adopter usage for a recently released internal generative AI interface called Argo based on a private instance2 of OpenAI’s GPT-3.5 Turbo large language model (LLM). We study all usage during the first several months Argo was available, and use the term early adopter to highlight that we study these initial users of the system.
Two features of the national lab make it a useful case study for science organizations and non-science organizations interested in using generative AI. First, numerous knowledge work organizations have a similar dichotomy between knowledge specialists (e.g., scientists, lawyers, investors) and operations workers, and in this paper we consider that they might have different needs but similar levels of interest with respect to generative AI. Second, the national lab regularly deals with sensitive data, such as classified, national security, or pre-published scientific data, and thus must take privacy and security risks seriously, similar to organizations such as banks and other government institutions. Our research questions are as follows:
RQ1: How are Science and Operations workers at a national lab currently using, and how do they envision using, generative AI to support their work? RQ2: What risks (including privacy, security, and ethics) exist for using generative AI at a national lab?
To answer our research questions, we conducted a survey (N= 66) and in-depth semi-structured interviews (N= 22) with Argonne employees, and also analyzed Argo usage data from the first eight months of deployment (which overlapped with the study time frame). We have four main findings: (1) Following its initial launch, we found that Argo was being used by a growing number of early adopters (more often in Science) and that most survey respondents were familiar with, and experimenting using, generative AI. However, few had made it a consistent part of their work; We identified common generative AI use cases that conceptually binned into either a (2) copilot or (3) workflow agent modality.
Related work. In recent years, scientists have been exploring how deep learning and generative AI might be useful for advancing scientific research [1, 49, 54]. Some have called for building science-specific large language models (LLMs) trained only on science literature [50, 55]. Another area of interest has been in the use of LLMs to generate science-focused programming code, including calculation kernels executed on HPC systems and parallel code snippets to train LLMs [8, 15, 31]. Generative AI is also being leveraged to create synthetic user research data for Human-Computer Interaction (HCI) research [19, 33] and drive automated process workflows for scientific tasks and experimental equipment [10, 26, 39, 45]. Little research exists, however, on how generative AI is impacting the day-to-day aspects of scientific knowledge work. One exception is a study by Morris [30] who interviewed 20 scientists, at a variety of institutions, and identified generative AI opportunities as well as concerns, ranging from literature reviews and data analysis to applications in higher education given many participants’ roles as university professors. Our paper differs from Morris’ by focusing on scientists at an organization that has a science mission (as opposed to university professors and scientists working in the tech industry). Notably, we include the perspective of Operations workers, investigate organization-level trends, and specifically seek perspectives on privacy and security. Another closely related study did not exclusively focus on generative AI. In this work, Crosby et al. [7] use a human-centered methodology to inform the design of a suite of tools for ocean scientists that leverage ML to process image data. Finally, some work has looked at the ability for LLMs to summarize academic papers for scholars [14, 57] or create data visualizations [29]. We expand on these studies by contributing a comprehensive look at how both Science and Operations staff in a science organization might use generative AI in their work and what concerns they have.
Researchers have also begun investigating the use of generative AI in knowledge work, a classification of labor involving the production of information-driven products and services as the key economic output [43, 59]. Some professions considered knowledge work include data science, law, marketing, and finance [43, 59]. Given generative AI’s ability to process text-based information, there has been growing interest in how this technology will impact the jobs of professional knowledge workers [3, 12, 16, 21, 27, 28, 43, 48, 59]. One thread of research looks to measure the productivity gains by knowledge workers who have access to generative AI tools [20].
Method. 3.1 Research Context In January 2024, the IT group at Argonne National Lab broadly deployed a generative AI chatbot named Argo, powered by a private instance of OpenAI’s GPT-3.5 Turbo,3 for use by all members of the Argonne National Lab community. Argo was designed for exclusive internal lab use so it does not save query and LLM response data and it does not share such information with OpenAI or other third-party services. Employees could use Argo as they would a service like ChatGPT: it included a browser-based interface that a user could type a prompt into and get a reply. Argo could also be accessed by employees via an API. Argo was intended to be used only for work purposes and required a lab login and VPN to access if not on-site at the lab. Beyond providing a secure generative AI assistant for employees, Argo was not released with a specific purpose. Employees could find out about Argo through official announcements, emails, or events targeted at raising awareness. Argonne National Lab is organized into Science and Operations divisions. The goal of Science divisions is to produce scientific research, while the goal of the Operations divisions is to keep the organization running smoothly. Science workers may be operating instruments, running experiments, managing data analysis, writing grants, and publishing papers. Operations workers may be producing science communication, working in administrative roles, ensuring lab safety (e.g. radiation exposure, cybersecurity), and building software and hardware infrastructure. Importantly, while the lab is divided into distinct divisions there is overlap in tasks between them. For example, both divisions employ technical workers such as software engineers. In addition, both divisions require some workers with a scientific background. In this paper, we distinguish between the two divisions to tease out if and when the need for generative AI differs between Science and Operations roles since this is a common division of labor in knowledge work organizations. We conducted a survey (April–June 2024) and interviews (April– July 2024) during the first eight months of Argo’s deployment to purposefully capture generative AI perceptions and use levels during this early stage. Our research was approved by the university 3.2 Data Collection 3.2.1 Survey. We first wanted to broadly capture national lab employee perceptions of generative AI uses and concerns in a survey. The survey, created in Qualtrics, included questions such as How familiar are you with large language models (LLMs)?, How often do you use LLMs as part of your work?, and Please describe ethical concerns you have about using LLMs in your work, if any. We also asked How often do you use LLMs for the following work tasks? and provided a list of 15 tasks such as writing code, feedback on experimental design, and editing human-written text. This list of tasks was drawn from a study of 20 scientists about their use of generative AI [30]. We also asked demographics questions. At the end of the survey, we gave participants the option to share their email if they would be willing to sign up for an interview. See the full survey in Appendix A.1.4 3.2.2 Interviews. As survey responses were completed, we contacted respondents who had provided email addresses for follow-up interviews. In total, we contacted 40 respondents, 22 of whom agreed to participate in semi-structured interviews. Each interview lasted 30 minutes and was conducted over Zoom and recorded. The interview protocol was designed to elicit more depth on generative AI applications, risks, and concerns than the survey. We asked participants to recall current or envisioned scenarios for generative AI, and also prompted them based on their survey responses. We then went in-depth on each scenario, asking questions such as To what extent have LLMs been helpful for this task? and When using an LLM to do this task doesn’t work, what goes wrong? In addition, we dug into participant views on privacy, security, and ethics concerns. See our full interview protocol in Appendix A.2.
Discussion. We categorize the first set of generative AI use cases described by participants in both Science and Operations roles as copilot-style interactions. This means they entail conversational interactions between the user and the AI where the user gets real-time responses to questions posed to the AI. We use the conceptual framing of a copilot in order to arrange our findings around features and affordances this copilot would need to be most useful in a science organization. We note that participants themselves rarely used the term copilot, rather we impose it for conceptual organization of the findings. At the time of data collection, participants said they were most often using LLMs such as ChatGPT to get help writing structured text that they can easily verify is correct, such as emails and reports. Participants envisioned goal, however, was to use a LLM to extract insights from unstructured text data such as scientific literature or survey results. We group Science and Operations employee responses together in this section since we did not find a large difference in use for copilot-style interactions.
4.2.1 Current Uses: Writing Structured Code/Text. Survey and interview participants described numerous examples of structured 4.2.2 Envisioned Uses: Extracting Insights from Large Unstructured Text Data. As science-focused knowledge workers, national lab employees must process significant quantities of unstructured textual data, such as scientific literature or organization regulations. At the time of the study, employees were hesitant to trust generative AI with extracting insights due to fears of hallucinations and reliability, which we return to in the Section 4.4.1. If these issues were resolved, however, we found a strong desire among participants to be able to get help from generative AI with managing, organizing, and learning from designated data sources. In the survey, 60% of respondents said they had at least tried summarizing literature (Figure 5) and 20% of survey free responses mentioned use cases related to querying unstructured text data (Table 3).
We categorize the second set of generative AI use cases described by participants as workflow agents. As opposed to a copilot, an agent navigates a complex task autonomously or semi-autonomously and returns the output to the user. In the context of a science organization, we found agents were emerging as a way of driving workflows in both Science and Operations. Scientific workflows included steps such as downloading data from an instrument or database, running multiple data analysis steps, and producing graphs or other visualizations. Operations workflows included tracking if work is progressing on-time, managing procurement processes, and automating common database interactions. Tasks related to workflow agents that survey respondents had tried included: analyzing data (43%); merging, cleaning, formatting data (35%); figure, graph generation (27%); creating synthetic data (21%); and labeling data (21%). In the survey free responses, 27% of answers (equally split between Science and Operations respondents) mentioned workflow automation as a use case for generative AI.
4.3.1 Current Uses: Initial Steps Towards Workflow Agents in Science and Operations. Participants in both Science and Operations described cases where they were testing using LLMs to automate some of their workflow. Participants reported LLMs are already able to automate some of these workflows in a scientific environment to make them more efficient, but generative AI workflow agents are in early stages of development and use. Science workflows: Multiple participants described how they were beginning to use LLMs to automate their customized scientific workflows.
Conclusion. In this paper, we study the practical, real-world applications and perceived risks of generative AI use across Science and Operations teams in a multidisciplinary science and engineering research center, Argonne National Lab. To understand current and envisioned generative AI use cases and privacy, security, ethics, and other concerns surrounding generative AI in a science organization, we report on usage statistics for the first release of a private instance of GPT-3.5 called Argo at the lab, a survey (N= 66), and in-depth interviews (N= 22). We find that there is an upward trend of Argo users, split between both Science and Operations employees, although use is largely experimental at this time. Uses cases fall into either a copilot or workflow agent generative AI modality. Risks include reliability, overreliance, privacy and security, the impact on academic publishing, and concerns around generative AI taking jobs. We end with recommendations for organizations interested in implementing generative AI systems and for HCI researchers working on crafting these systems.
Limitations. This paper provides a case study of a single science organization. Given how little is currently known about organizational use of generative AI assistants, findings from this case study are applicable to a broad range of knowledge work institutions such as those with knowledge specialists (e.g. scientists, lawyers, etc.) and operations workers; in addition, future research should study a variety of science and other knowledge work organizations. When studying usage in the organization, we did not have access to the number of employees accessing commercial LLMs and so our usage data for Argo under-counts total LLM usage.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
Does AI-assisted work increase total productivity or just shift time?- How does working with AI shift where knowledge workers spend their time?
- How do workers signal effort and voice when using AI tools?
- Does receiving AI output shift workers' time away from their own productive tasks?
- Does AI-assisted work reduce time spent on coordination and communication?
- Does generative AI push knowledge workers toward different types of tasks?
- Does generative AI substitute for labor or complement worker productivity?
- Does generative AI narrow performance gaps between different professional backgrounds?
- Do gains from AI assistance disappear when workers complete tasks alone?
- Does benefit from AI partnership depend on the individual worker?
- Does generative AI adoption shift work away from coordination tasks?
- What does selective and critical GenAI use look like in daily practice?
- Does AI adoption push knowledge work away from communication toward solo tool use?
- Why does AI adoption shift knowledge work toward individual documentation focus?
- Does AGI focus distract firms from developing task-creating AI innovations?
- What specific training approaches help managers integrate AI into team workflows?
- What makes colleagues willing to share how they actually use GenAI at work?
- What directions does AI-generated workslop flow within organizations most often?
- Does AI shift knowledge work away from communication toward solo documentation tasks?