01
Model behavior & human judgment
Reasoning, ethical judgment, persuasion, deception, manipulation, calibration, and failure modes produced by framing or linguistic variation.
AI researcher
Katherine Elkins studies how increasingly capable AI systems reason, persuade, interpret, coordinate, and interact with people—and how to build systems that are more culturally and historically accurate.
Her research spans model behavior, human judgment, narrative, agentic and multi-agent systems, behavioral evaluation, cultural memory, and governance. She combines computational methods with expertise in language, history, culture, and human reasoning to design, test, and interpret AI research.
Current research
Many consequential AI questions are simultaneously technical and human: what a model does under pressure, how wording or narrative changes a judgment, when an explanation tracks behavior, and how agents coordinate or deceive. Elkins works directly on those questions through computational experiments, behavioral testing, and interdisciplinary interpretation.
01
Reasoning, ethical judgment, persuasion, deception, manipulation, calibration, and failure modes produced by framing or linguistic variation.
02
Agentic behavior, coordination, debate, emotional stability, and deception in systems where models interact with one another and make consequential decisions.
03
Narrative as a cognitive and emotional tool: a structure for memory and interpretation, a mechanism of persuasion, and a factor in how models and people organize judgment.
04
Behavioral evaluation, model comparison, ethics-based auditing, red-teaming, benchmark design, and testing that reveals vulnerabilities conventional metrics can miss.
05
AI systems that work with archives, cultural memory, provenance, retrieval, context, and expert interpretation without treating knowledge as context-free data.
06
Standards, regulation, public-interest AI, and the institutional consequences of deploying increasingly capable models in high-stakes environments.
Current project
Schmidt Sciences
Humanities and AI Virtual Institute
Building open AI tools for rescuing endangered cultural archives—and AI systems that are more culturally and historically accurate.
Elkins is Co-PI of a collaborative project beginning in New Orleans that develops systems for preserving, organizing, retrieving, and interpreting cultural records. The technical and intellectual goal is to work with cultural and historical materials without stripping away provenance, context, ambiguity, or expert knowledge. The project treats archives as a demanding AI research environment rather than a humanities application added after system design.
Research in practice
Elkins co-leads the five-member team representing the Modern Language Association in the federal consortium. The team connects language, ambiguity, framing, persuasion, narrative, and human judgment to national work on AI measurement and evaluation.
Kenyon describes the program Elkins co-founded as the world’s first Human-Centered AI curriculum and lab: a research model giving humanities and social-science students the technical fluency to conduct original computational and AI research while keeping disciplinary knowledge central.
Current work compares human, statistical, and language-model judgments in high-stakes prediction and examines how syntax, framing, and reported reasoning can diverge from model behavior.
Methodological foundations
Elkins’s computational research developed rigorous ways to operationalize narrative, emotion, interpretation, translation, cultural meaning, and judgment without reducing away context or ambiguity. She now applies and extends those methods to increasingly capable models, agents, and AI systems; narrative remains an active object of research as a tool of cognition, emotion, memory, and persuasion.
Cambridge University Press · 2022
Introduced the first rigorous methodology for narrative sentiment analysis: a reproducible framework for selecting, comparing, validating, and interpreting sentiment models across narrative texts.
Computational interpretation
Middle Reading joins close interpretation to computational scale. SentimentArcs is a novel ensemble method for comparing narrative sentiment trajectories across texts.
Frontiers in Computer Science · 2024
Computational comparison makes visible what changes across translations—and what model outputs can miss when language carries cultural and interpretive context. Follow the broader transmission research →
Selected firsts
Kenyon describes the program as the world’s first Human-Centered AI curriculum and lab.
First writer’s Turing test of a large language model.
First rigorous methodology for narrative sentiment analysis.
First application of explainable AI to narrative analysis.
First ethics-based audit of moral reasoning in deployed large language models.
First systematic comparison of AI regulation across the European Union, China, and the United States after passage of the EU AI Act.
Scholarly reception
The ethics-audit confidence-scoring method has been adopted in later LLM value evaluation. The open-model benefit–risk framework has shaped FAccT, TMLR, and governance research. Narrative methods have been extended in NLP, translation, behavioral science, persuasion research, and large-scale studies of stories.
Speaking & exchange
Talks and conversations address model behavior, AI governance, narrative and persuasion, cultural memory, human-centered research, higher education, and the design of AI systems grounded in cultural and institutional contexts.
Contact
Research collaborations, institutes, standards work, speaking, and media inquiries →