Does AI Assistance Leave a Temporal Fingerprint? Detecting Overreliance in AI-Assisted Writing and Programming
The rapid adoption of generative AI has made final artifacts unreliable evidence of student learning, and AI detectors that examine only the finished product are inaccurate and ethically contentious. Process data offers an alternative, but prior work covers only English essay writing. We ask whether AI assistance carries a temporal signature, whether it generalizes from writing to programming, and whether it distinguishes ordinary collaboration from wholesale delegation. We analyze three public corpora: CoAuthor (1,447 keystrokelevel co-writing sessions), RealHumanEval (editor telemetry from 243 programmer records), and a pre-LLM CS1 corpus (5.1 million keystrokes) as a human-only baseline, comparing minimal-AI work, collaborative AI use, and simulated wholesale delegation. Three findings emerge. First, the signature generalizes: AI contributions arrive in bursts far outside the author's own baseline in both mediums (paired dz = 1.13 and 3.54). Second, engagement diverges by medium: 93% of AI-inserted characters survived to writers' final documents, while only 14% of accepted code suggestions survived intact.
Introduction. When students can produce polished essays and code with conversational agents, the finished artifact no longer reveals how learning occurred. Institutional responses have leaned on AI text detectors, which operate on the final product alone, show unreliable accuracy, and produce contested accusations that erode trust (Deep et al., 2025). The Learning Visibility Framework reframes this as a measurement problem: what is lost when AI enters the assessment loop is visibility into the learning process, and what assessment requires is process-based evidence alongside outcomes (Davalos & Zhang, 2026). Process data is a promising basis for such evidence: keystroke features distinguish authentic composition from transcription of AI-generated text with high accuracy in English essays (Crossley et al., 2024; Deane et al., 2026; Kundu et al., 2024; Roh et al., 2025). But this evidence covers a single medium. No peerreviewed study has tested whether AI-inserted code is detectable from temporal process features, and none has compared signatures across mediums.
Discussion / Conclusion. Process evidence flags delegation, not assistance. Ordinary AI collaboration is only weakly separable from minimally assisted work, and that is the desirable outcome: students using AI as permitted should not be flagged, and in our results, they essentially never are (at most 0.5% of real work classed as delegation). Wholesale delegation, by contrast, produces a categorical signature that survives without provenance labels or keystroke resolution. The separability is partly by construction, since concentrated production defines the simulated behavior; what the result establishes is sensitivity and observability, not prevalence or field accuracy. The other half of the asymmetry, that collaboration resists detection, is corroborated by a deployed keystroke screening tool that reports missing hybrid What this method cannot detect, and what a flag can never mean.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
How does AI-generated content transformation affect public discourse quality?- Does positive sentiment bias in AI content harm information quality?
- What makes AI-generated punditry different from human expert commentary online?
- What changes when published text was never written for its readers?
- How does the author-function itself change when AI replaces human authorship?
- Why do users prefer AI text versions even when they misrepresent their own views?
- Does AI make writers appear more politically extreme to readers?
- Why does AI writing seem more competent and informative than human writing?
- Does AI writing make authors appear more privileged or educated?
- How does AI assistance affect perceived emotional tone in writing?
- How does perceived writer confidence shift with AI-assisted composition?
- Does homogenization at the text level cause homogenization of perceived authors?
- Does AI writing erase markers of non-native English speaker identity?
- Can demographic distortion in AI writing affect who appears credible in public discourse?
- Why are education and language fluency more affected than race perception?
- When do readers defer to AI text without genuine processing?
- What makes readers treat AI-generated text as authoritative?
- What specific distortions does AI writing assistance introduce into text?
- How do writer preferences for AI output affect their willingness to edit it?