Associate Professor at the University of Southern Denmark
I am an Associate Professor at the Department of Mathematics and Computer Science at the University of Southern Denmark (SDU), where I lead the AI Safety & Interpretability Lab and co-lead OdenseNLP. I am also an affiliated member of the Pioneer Centre for Artificial Intelligence (P1), a member of the Digital Democracy Centre, and a member of the International Association for Safe and Ethical AI.
My research advances AI safety by understanding how and why AI models behave the way they do – and by developing methods to keep that behavior legible and controllable. Specifically, my research areas comprise interpretability, emergent communication, and multi-agent LLM systems. One line of my work studies model internals to identify the functional roots of model behavior, using mechanistic interpretability and experimental paradigms from psycholinguistics (calibration of activation oracles, culture neurons). In another line of work, I study the interaction between models – how communication emerges in populations of agents, and why some communication protocols stay readable while others drift away from human language (compositionality in language learning, emergent languages in LLM agent populations). These strands converge on the key question of how to retain legibility in machine reasoning and communication, and at what cost (auditability vs. accuracy).
I am the principal investigator of the MIST project on scalable mechanistic interpretability for safe and trustworthy LLM agents, funded by the Novo Nordisk Foundation, and I lead the alignment workstream in the Danish Foundation Models project, where we train open multilingual models that excel at Danish using only permissible data (DFM Mimir v1) – contributing to European AI sovereignty.
Contact: lukas 'at' lpag.de
Design: Adapted from Diane Mounter.
Privacy: No personal data, no cookies.