Behavior in complex human settings

The common problem is not simply whether an AI system is accurate. It is how models reason, persuade, interpret, coordinate, and fail when language, judgment, incentives, culture, and institutions shape the task.

Model behavior & human judgment

Reasoning, ethical judgment, decision-making, persuasion, deception, manipulation, calibration, and linguistic failure modes. This work tests how behavior changes under syntactic reframing, ambiguity, social pressure, and shifts in context—and whether a model’s explanation tracks what it actually does.

Model behavior and evaluation research →

Agents & multi-agent systems

Coordination, debate, explanation, adversarial interaction, emergent behavior, and evaluation of agents. Current mentored and collaborative work examines deception in multi-agent negotiation, emotional stability among interacting agents, and the use of agent debate in consequential decisions.

Agentic and multi-agent research →

Narrative, cognition & emotion

Narrative is a cognitive and emotional tool: a structure for memory and interpretation, a mechanism of persuasion, and a factor in how models and people organize judgment. Current work connects literary study, cognitive science, and philosophy of information to model responses, behavioral effects, human-AI interaction, and vulnerabilities to persuasion.

Narrative and emotion research →

Archival Intelligence

Rescuing endangered cultural archives.

Elkins is Co-PI of Archival Intelligence, a Schmidt Sciences Humanities and AI Virtual Institute project building open AI tools for endangered archives, beginning in New Orleans.

The research asks how AI systems can retrieve, represent, and interpret cultural and historical materials without stripping away provenance, context, ambiguity, or expert knowledge. Generic retrieval can find a record while missing why it matters, who is authorized to interpret it, or how place and community shape its meaning. Those are system-design problems.

The project joins archival retrieval, knowledge representation, interpretation, domain expertise, and technical infrastructure to make AI systems more culturally and historically accurate. Related work addresses cultural data governance and public knowledge infrastructure, including participation in UNESCO’s AI, IP & Culture Repository co-design process.

Explore Archival Intelligence →

Evaluation, robustness, and institutions

Evaluation & robustness

Behavioral evaluation, benchmark design, model comparison, confidence and calibration, red-teaming, ethics-based auditing, and failure-mode analysis. Elkins and Chun conducted the first ethics-based audit of moral reasoning in deployed large language models; current work continues to examine vulnerabilities conventional performance metrics can miss.

AI evaluation and robustness →

AI, institutions & governance

Standards, comparative regulation, open-model governance, public-interest AI, and institutional evaluation grounded in evidence about model behavior. Elkins co-leads the MLA team in the NIST CAISI AI Consortium and coauthored the first systematic EU–China–US regulatory comparison after passage of the EU AI Act.

Governance and public AI →

Making difficult phenomena testable

How can context, ambiguity, interpretation, and historical meaning remain present in computational research?

The program began with literature and philosophy; its methods now inform research on models, agents, and institutions.

Narrative sentiment methodology

Elkins and Chun introduced the first rigorous methodology for narrative sentiment analysis, establishing a reproducible framework for selecting, comparing, validating, and interpreting sentiment models across narrative texts. Earlier applications had not supplied a systematic procedure for model selection, robustness, smoothing and scale choices, or interpretation across methods.

The Shapes of Stories

SentimentArcs & Middle Reading

Developed by Jon Chun, SentimentArcs is a novel ensemble method for comparing narrative sentiment trajectories across texts. The ensemble reduces dependence on any single sentiment system; Middle Reading connects those computational comparisons to interpretation at the scale of passages, scenes, and whole works.

Methods for narrative and emotion →

Explainability, translation & multimodality

The program includes the first application of explainable AI to narrative analysis, computational comparison of translations, and the first multimodal method for measuring long-form film sentiment-arc coherence across modalities. Each asks what a representation preserves, what it discards, and how those choices affect interpretation.

Narrative, translation, and transmission →

From methods to models

The same concerns now shape behavioral evaluation, persuasion research, agent analysis, and culturally grounded AI. Can GPT-3 Pass a Writer’s Turing Test?, coauthored by Elkins and Chun, was the first writer’s Turing test of a large language model and made literary imitation and reasoning experimentally comparable.

AI, creativity, and authorship → Philosophy of mind →

Publications and projects

Model behavior & evaluation · 2024

Informed AI Regulation

Ethics-based auditing of moral reasoning across deployed language models.

Narrative method · Cambridge University Press, 2022

The Shapes of Stories

Rigorous methodology and ensemble comparison for narrative sentiment analysis.

Cultural and historical AI · Current project

Archival Intelligence

Open AI systems for endangered cultural archives, provenance, retrieval, and interpretation.

Google Scholar → Books and major works → AI in higher education →

Scholarly reception

The citation record now spans 15 fields and 55 subfields, including AI, digital humanities, literary studies, education, medicine, law, and related fields.