Katherine Elkins, AI researcher

Katherine Elkins

AI researcher

Katherine Elkins studies how increasingly capable AI systems reason, persuade, interpret, coordinate, and interact with people—and how to build systems that are more culturally and historically accurate.

Her research spans model behavior, human judgment, narrative, agentic and multi-agent systems, behavioral evaluation, cultural memory, and governance. She combines computational methods with expertise in language, history, culture, and human reasoning to design, test, and interpret AI research.

AI behavior in complex human environments

Many consequential AI questions are simultaneously technical and human: what a model does under pressure, how wording or narrative changes a judgment, when an explanation tracks behavior, and how agents coordinate or deceive. Elkins works directly on those questions through computational experiments, behavioral testing, and interdisciplinary interpretation.

01

Model behavior & human judgment

Reasoning, ethical judgment, persuasion, deception, manipulation, calibration, and failure modes produced by framing or linguistic variation.

02

Agents & multi-agent systems

Agentic behavior, coordination, debate, emotional stability, and deception in systems where models interact with one another and make consequential decisions.

03

Narrative, cognition & emotion

Narrative as a cognitive and emotional tool: a structure for memory and interpretation, a mechanism of persuasion, and a factor in how models and people organize judgment.

04

Evaluation & robustness

Behavioral evaluation, model comparison, ethics-based auditing, red-teaming, benchmark design, and testing that reveals vulnerabilities conventional metrics can miss.

05

Culturally & historically grounded AI

AI systems that work with archives, cultural memory, provenance, retrieval, context, and expert interpretation without treating knowledge as context-free data.

06

AI, institutions & governance

Standards, regulation, public-interest AI, and the institutional consequences of deploying increasingly capable models in high-stakes environments.

Schmidt Sciences
Humanities and AI Virtual Institute

Archival Intelligence

Building open AI tools for rescuing endangered cultural archives—and AI systems that are more culturally and historically accurate.

Elkins is Co-PI of a collaborative project beginning in New Orleans that develops systems for preserving, organizing, retrieving, and interpreting cultural records. The technical and intellectual goal is to work with cultural and historical materials without stripping away provenance, context, ambiguity, or expert knowledge. The project treats archives as a demanding AI research environment rather than a humanities application added after system design.

Current institutional work

NIST CAISI AI Consortium · 2024–present

Measurement, evaluation, and trustworthy AI

Elkins co-leads the five-member team representing the Modern Language Association in the federal consortium. The team connects language, ambiguity, framing, persuasion, narrative, and human judgment to national work on AI measurement and evaluation.

Research context →

Human-Centered AI Lab

Technical research with domain expertise

Kenyon describes the program Elkins co-founded as the world’s first Human-Centered AI curriculum and lab: a research model giving humanities and social-science students the technical fluency to conduct original computational and AI research while keeping disciplinary knowledge central.

Mentored research →

Model evaluation

Behavioral vulnerabilities

Current work compares human, statistical, and language-model judgments in high-stakes prediction and examines how syntax, framing, and reported reasoning can diverge from model behavior.

Research overview →

Making difficult human phenomena empirically testable

Elkins’s computational research developed rigorous ways to operationalize narrative, emotion, interpretation, translation, cultural meaning, and judgment without reducing away context or ambiguity. She now applies and extends those methods to increasingly capable models, agents, and AI systems; narrative remains an active object of research as a tool of cognition, emotion, memory, and persuasion.

Cambridge University Press · 2022

The Shapes of Stories

Introduced the first rigorous methodology for narrative sentiment analysis: a reproducible framework for selecting, comparing, validating, and interpreting sentiment models across narrative texts.

Computational interpretation

Middle Reading & SentimentArcs

Middle Reading joins close interpretation to computational scale. SentimentArcs is a novel ensemble method for comparing narrative sentiment trajectories across texts.

Original methods and research designs

  1. 2016

    Human-Centered AI

    Kenyon describes the program as the world’s first Human-Centered AI curriculum and lab.

  2. 2020

    Writer’s Turing test

    First writer’s Turing test of a large language model.

  3. 2022

    Narrative sentiment methodology

    First rigorous methodology for narrative sentiment analysis.

  4. 2023

    Explainable AI for narrative

    First application of explainable AI to narrative analysis.

  5. 2024

    Ethics-based model audit

    First ethics-based audit of moral reasoning in deployed large language models.

  6. 2024

    Comparative AI regulation

    First systematic comparison of AI regulation across the European Union, China, and the United States after passage of the EU AI Act.

Methods taken up across fields

The ethics-audit confidence-scoring method has been adopted in later LLM value evaluation. The open-model benefit–risk framework has shaped FAccT, TMLR, and governance research. Narrative methods have been extended in NLP, translation, behavioral science, persuasion research, and large-scale studies of stories.

Research in public

Talks and conversations address model behavior, AI governance, narrative and persuasion, cultural memory, human-centered research, higher education, and the design of AI systems grounded in cultural and institutional contexts.

Research collaborations, institutes, standards work, speaking, and media inquiries