Pairit: A Platform for Live Experiments on Human-AI Collaboration

Paper · arXiv 2609.09789 · Published September 9, 2026
Design Frameworks

Organizational design in the era of artificial intelligence requires experimental methods that can test how human–AI groups coordinate, delegate, and make decisions. Programmable platforms coordinate live human-to-human sessions or real-time human–AI chat, but researchers cannot easily declare experiment protocols in which AI participants both communicate and act on shared work within one auditable configuration. Here we introduce Pairit, an online platform that facilitates the design, testing, and deployment of experiments that test human–AI organizational designs and interventions. Through a single YAML configuration file, researchers declare an executable experiment graph—pages, routing, randomization, matchmaking, chat, shared workspaces, server-hosted agents, surveys, timers, and custom HTML components—and combine any number of humans and AI agents in live sessions. We have validated the feasibility of the platform through multiple live deployments, including peer-reviewed published studies, capturing high-resolution process traces of communication, negotiation, and collaborative work in live human–AI dyads.

Introduction. Researchers run experiments to causally test design choices in human–AI teams. Programmable platforms such as oTree [3] and Empirica [1] coordinate live human-to-human sessions, and newer systems such as Deliberate Lab [12] support real-time human–AI group chat. Yet researchers still cannot easily declare experiment protocols in which AI agents converse, co-edit shared documents, and take protocol-defined actions within one auditable configuration that also specifies team assignment, communication channels, and routing. Without that capability, researchers cannot easily and systematically vary team composition, live communication, and how agents collaborate during tasks. Because each new design requires custom software, live experiment protocols rarely become shareable research objects other scholars can inspect or run. Here we introduce Pairit, which lets researchers declare live experiment protocols with any mix of human and AI participants in a single auditable configuration, run them as live sessions, and share them with other labs.

Discussion / Conclusion. Pairit enables new research on coordination, communication, and delegation in live human–AI teams. Until now, studying these dynamics meant building live infrastructure from scratch: synchronizing participants, routing conversation, managing shared work, and embedding agents that both speak and act. That work remains slow and design-intensive even with AI-assisted coding, which is why many live human–AI designs were rarely tested in experiments. Pairit turns that build work into a declarative experiment graph, so researchers can systematically vary who participates, what role AI plays, and when it intervenes. Researchers can now test facilitation, delegation, and bargaining assistance in repeatable, customizable, and shareable live experiments rather than one-off custom builds. Pairit also makes experiment protocols auditable, shareable, and reproducible. Standard prose methods sections cannot fully specify live experiment protocols.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

How should dialogue recommender systems manage conversation history and state? How can AI alignment serve diverse human preferences at scale? How can recommendation systems balance personalization with stability and coverage? How can AI systems learn from failures without cascading errors? How do knowledge graphs enable efficient multi-hop reasoning over alternatives? How should personalization be implemented to improve AI assistant effectiveness? How should conversational agents balance goal-driven initiative with user control? How should dialogue systems best leverage conversation history for retrieval? How do social dynamics and selection effects compound in rating aggregates? How can persona representations reduce language model variance and improve task accuracy? How do language models inherit human biases from training data? What dimensions of recommendation quality do standard metrics miss? How do formal dialogue structures reveal conversation coherence mechanisms? How can we distinguish genuine user preferences from measurement artifacts?