Katherine Elkins
Current research
AI behavior in complex human environments
Many consequential AI questions are simultaneously technical and human: what a model does under pressure, how wording or narrative changes a judgment, when an explanation tracks behavior, and how agents coordinate or deceive.
Elkins works directly on those questions through computational experiments, behavioral testing, and interdisciplinary interpretation.
Her research spans model behavior, human judgment, narrative, agentic and multi-agent systems, behavioral evaluation, cultural memory, and governance.
Research constellation
Research constellation
Move across the map, or tap a star, to read about it.
The constellation, as text
Every star on the map
The graph above is drawn from this list. Each entry names what it connects to, so the full structure reads without the animation.
Research areas
-
Model behavior & human judgment
Reasoning, ethical judgment, persuasion, deception, manipulation, calibration, and failure modes produced by framing or linguistic variation. This work tests whether a model's explanation tracks what it actually does.
Connected to: Agents & multi-agent systems, Narrative, cognition & emotion, Evaluation & robustness, NIST CAISI AI Consortium, How Well Can GenAI Predict Human Behavior?, Can GPT-3 Pass a Writer’s Turing Test?, Informed AI Regulation, Syntactic Framing Fragility
-
Agents & multi-agent systems
Agentic behavior, coordination, debate, emotional stability, and deception in systems where models interact with one another and make consequential decisions.
Connected to: Model behavior & human judgment, Human-Centered AI curriculum & lab, Deception, stability & debate among agents
-
Narrative, cognition & emotion
Narrative as a cognitive and emotional tool: a structure for memory and interpretation, a mechanism of persuasion, and a factor in how models and people organize judgment.
Connected to: Model behavior & human judgment, Culturally & historically grounded AI, AI creativity, authorship & co-creation, Philosophy of mind, literature & information, The Shapes of Stories, SentimentArcs & Middle Reading, In Search of a Translator, The Shapes of Cinderella
-
Evaluation & robustness
Behavioral evaluation, model comparison, ethics-based auditing, red-teaming, benchmark design, and testing that reveals vulnerabilities conventional metrics can miss.
Connected to: Model behavior & human judgment, AI, institutions & governance, NIST CAISI AI Consortium, How Well Can GenAI Predict Human Behavior?, Informed AI Regulation, Syntactic Framing Fragility
-
Culturally & historically grounded AI
AI systems that work with archives, cultural memory, provenance, retrieval, context, and expert interpretation without treating knowledge as context-free data.
Connected to: Narrative, cognition & emotion, Archival Intelligence, In Search of a Translator, The Shapes of Cinderella, Provenance Infrastructure as a Safeguard for Cultural Commons
-
AI, institutions & governance
Standards, regulation, public-interest AI, and the institutional consequences of deploying increasingly capable models in high-stakes environments.
Connected to: Evaluation & robustness, Archival Intelligence, NIST CAISI AI Consortium, Informed AI Regulation, Near to Mid-term Risks and Opportunities of Open-Source Generative AI, Comparative AI regulation: EU, China, US, Provenance Infrastructure as a Safeguard for Cultural Commons
Related topics
-
AI creativity, authorship & co-creation
What language models can generate, imitate, translate, and co-create, and what their outputs reveal about authorship, intention, literary value, and creative practice.
Connected to: Narrative, cognition & emotion, In Search of a Translator, Can GPT-3 Pass a Writer’s Turing Test?
-
Philosophy of mind, literature & information
How knowing happens: how consciousness organizes experience, how memory reconstructs the past, how perception is mediated, and how narrative shapes the self.
Connected to: Narrative, cognition & emotion, Proust’s In Search of Lost Time: Philosophical Perspectives
-
AI in higher education
What universities should teach when students can delegate traditional academic tasks to machines. The answer moves students upstream: from users of AI systems to researchers who evaluate, build, interpret, and question them.
Connected to: Human-Centered AI curriculum & lab
Projects & roles
-
Archival Intelligence
Open AI tools for rescuing endangered cultural archives, beginning in New Orleans. The research asks how AI systems can retrieve, represent, and interpret cultural and historical materials without stripping away provenance, context, ambiguity, or expert knowledge.
Connected to: Culturally & historically grounded AI, AI, institutions & governance, Provenance Infrastructure as a Safeguard for Cultural Commons
-
NIST CAISI AI Consortium
Elkins co-leads the five-member team representing the Modern Language Association in the NIST CAISI AI Consortium, bringing research on language, ambiguity, framing, persuasion, narrative, and human judgment into federal work on AI measurement and evaluation.
Connected to: Model behavior & human judgment, Evaluation & robustness, AI, institutions & governance, Informed AI Regulation
-
Human-Centered AI curriculum & lab
An interdisciplinary research model that combines technical training with domain expertise and original student research.
Connected to: Agents & multi-agent systems, AI in higher education, Deception, stability & debate among agents
-
How Well Can GenAI Predict Human Behavior?
Audits how LLMs reason in high-stakes decisions over humans, using recidivism prediction as a test case and comparing human experts, statistical models, and LLMs through the FATE framework: Fairness, Accuracy, Transparency, and Explainability.
Connected to: Model behavior & human judgment, Evaluation & robustness
-
Deception, stability & debate among agents
Current mentored and collaborative work examines deception in multi-agent negotiation, emotional stability among interacting agents, and the use of agent debate in consequential decisions.
Connected to: Agents & multi-agent systems, Human-Centered AI curriculum & lab
Books & selected work
-
The Shapes of Stories
Introduced the first rigorous methodology for narrative sentiment analysis: a reproducible framework for selecting, comparing, validating, and interpreting sentiment models across narrative texts.
Connected to: Narrative, cognition & emotion, SentimentArcs & Middle Reading
-
Proust’s In Search of Lost Time: Philosophical Perspectives
Brings Proust into philosophy as a thinker of consciousness, perception, time, selfhood, ethics, and aesthetic experience. Elkins wrote the introduction and the chapter “Proust’s Consciousness.”
Connected to: Philosophy of mind, literature & information, In Search of a Translator
-
SentimentArcs & Middle Reading
Middle Reading joins close interpretation to computational scale. SentimentArcs, developed by Jon Chun, is a novel ensemble method for comparing narrative sentiment trajectories across texts.
Connected to: Narrative, cognition & emotion, The Shapes of Stories
-
In Search of a Translator
Computational comparison of Proust across literary translations makes visible what changes, and what model outputs can miss when language carries cultural and interpretive context.
Connected to: Narrative, cognition & emotion, Culturally & historically grounded AI, AI creativity, authorship & co-creation, Proust’s In Search of Lost Time: Philosophical Perspectives
-
The Shapes of Cinderella
Compares original-language emotional trajectories of Ye Xian, Perrault’s Cendrillon, and the Grimm versions. What travels is a recognition scaffold, while each culture rebuilds the tale as a distinct emotional-moral architecture.
Connected to: Narrative, cognition & emotion, Culturally & historically grounded AI
-
Can GPT-3 Pass a Writer’s Turing Test?
The first writer’s Turing test of a large language model, which made literary imitation and reasoning experimentally comparable.
Connected to: Model behavior & human judgment, AI creativity, authorship & co-creation
-
Informed AI Regulation
The first ethics-based audit of moral reasoning in deployed large language models. Its regulatory claim is empirical: informed regulation should begin from observed model behavior.
Connected to: Model behavior & human judgment, Evaluation & robustness, AI, institutions & governance, NIST CAISI AI Consortium, Syntactic Framing Fragility
-
Syntactic Framing Fragility
Documents systematic failures in how frontier models respond to syntactic framing changes in safety-critical moral prompts, extending ethics-based auditing from value comparison to linguistic robustness.
Connected to: Model behavior & human judgment, Evaluation & robustness, Informed AI Regulation
-
Near to Mid-term Risks and Opportunities of Open-Source Generative AI
A benefit–risk framework for open-source generative AI.
Connected to: AI, institutions & governance
-
Comparative AI regulation: EU, China, US
The first systematic comparison of AI regulation across the European Union, China, and the United States after passage of the EU AI Act.
Connected to: AI, institutions & governance
-
Provenance Infrastructure as a Safeguard for Cultural Commons
Technical mechanisms for tracking cultural-data provenance and supporting rights-aware AI training practices, included among the report’s operational proposals.
Connected to: Culturally & historically grounded AI, AI, institutions & governance, Archival Intelligence
Six research areas
AI behavior in complex human environments
-
Model behavior & human judgment
Reasoning, ethical judgment, persuasion, deception, manipulation, calibration, and failure modes produced by framing or linguistic variation.
Model behavior and human judgment -
Agents & multi-agent systems
Agentic behavior, coordination, debate, emotional stability, and deception in systems where models interact with one another and make consequential decisions.
Agents and multi-agent systems -
Narrative, cognition & emotion
Narrative as a cognitive and emotional tool: a structure for memory and interpretation, a mechanism of persuasion, and a factor in how models and people organize judgment.
Narrative, cognition and emotion -
Evaluation & robustness
Behavioral evaluation, model comparison, ethics-based auditing, red-teaming, benchmark design, and testing that reveals vulnerabilities conventional metrics can miss.
Evaluation and robustness -
Culturally & historically grounded AI
AI systems that work with archives, cultural memory, provenance, retrieval, context, and expert interpretation, keeping knowledge tied to its context.
Culturally and historically grounded AI -
AI, institutions & governance
Standards, regulation, public-interest AI, and the institutional consequences of deploying increasingly capable models in high-stakes environments.
AI, institutions and governance
Current project · Schmidt Sciences Humanities and AI Virtual Institute
Archival Intelligence
Building open AI tools for rescuing endangered cultural archives, and AI systems that are more culturally and historically accurate.
Elkins is Co-PI of a collaborative project beginning in New Orleans that develops systems for preserving, organizing, retrieving, and interpreting cultural records. The goal is to work with cultural and historical materials without stripping away provenance, context, ambiguity, or expert knowledge.
Explore Archival IntelligenceResearch in practice
Research leadership
-
NIST CAISI AI Consortium · 2024–present
Measurement, evaluation, and trustworthy AI
Elkins co-leads the five-member team representing the Modern Language Association in the NIST CAISI AI Consortium, bringing research on language, ambiguity, framing, persuasion, narrative, and human judgment into federal work on AI measurement and evaluation.
-
Schmidt Sciences
Archival Intelligence
Co-Principal Investigator of a Schmidt Sciences Humanities and AI Virtual Institute project developing open AI systems for endangered cultural archives, bringing together technical research, archival expertise, cultural institutions, and research infrastructure across multiple institutions.
-
Kenyon College · since 2016
Human-Centered AI
Elkins and Jon Chun founded the Human-Centered AI curriculum and lab at Kenyon in 2016, creating an interdisciplinary research model that combines technical training with domain expertise and original student research.
Research, education & curriculum
AI in Higher Education
Katherine Elkins’s work asks what universities should teach when students can increasingly delegate traditional academic tasks to machines. Her answer moves students upstream: from users of AI systems to researchers who evaluate, build, interpret, and question them.
Elkins and Jon Chun founded the Human-Centered AI curriculum and lab at Kenyon in 2016. The program connects technical training, disciplinary expertise, original student research, and a larger argument about what universities are for.
Explore the curriculum, lab, and research archiveMethodological foundations
Making difficult human phenomena empirically testable
Elkins’s computational research developed rigorous ways to operationalize narrative, emotion, interpretation, translation, cultural meaning, and judgment while keeping context and ambiguity in view. She now applies and extends those methods to increasingly capable models, agents, and AI systems.
The Shapes of Stories
Introduced the first rigorous methodology for narrative sentiment analysis: a reproducible framework for selecting, comparing, validating, and interpreting sentiment models across narrative texts.
Middle Reading & SentimentArcs
Middle Reading joins close interpretation to computational scale. SentimentArcs is a novel ensemble method for comparing narrative sentiment trajectories across texts.
Translation, language, and interpretation
Computational comparison makes visible what changes across translations, and what model outputs can miss when language carries cultural and interpretive context.
Selected firsts
Original methods and research designs
- 2016
Human-Centered AI
Elkins and Jon Chun founded the Human-Centered AI curriculum and lab at Kenyon.
- 2020
Writer’s Turing test
First writer’s Turing test of a large language model.
- 2022
Narrative sentiment methodology
First rigorous methodology for narrative sentiment analysis.
- 2023
Explainable AI for narrative
First application of explainable AI to narrative analysis.
- 2024
Ethics-based model audit
First ethics-based audit of moral reasoning in deployed large language models.
- 2024
Comparative AI regulation
First systematic comparison of AI regulation across the European Union, China, and the United States after passage of the EU AI Act.
Scholarly reception
Methods taken up across fields
The ethics-audit confidence-scoring method has been adopted in later LLM value evaluation. The open-model benefit–risk framework has shaped FAccT, TMLR, and governance research. Narrative methods have been extended in NLP, translation, behavioral science, persuasion research, and large-scale studies of stories.
Read the evidenceThe citation record spans 15 fields and 55 subfields.
The citation record spans AI, digital humanities, literary studies, education, medicine, law, and related fields.
Speaking & exchange
Research in public
Talks and conversations address model behavior, AI governance, narrative and persuasion, cultural memory, human-centered research, higher education, and the design of AI systems grounded in cultural and institutional contexts.
Selected speakingContact