Lukas Galke Poech | Advancing AI safety through interpretability-informed control

Lukas Galke Poech

Associate Professor at the University of Southern Denmark

I am an Associate Professor at the Department of Mathematics and Computer Science at the University of Southern Denmark (SDU), where I lead the AI Safety & Interpretability Lab and co-lead OdenseNLP. I am also an affiliated member of the Pioneer Centre for Artificial Intelligence (P1), a member of the Digital Democracy Centre, and a member of the International Association for Safe and Ethical AI.

My research advances AI safety by understanding how and why AI models behave the way they do – and by developing methods to keep that behavior legible and controllable. Specifically, my research areas comprise interpretability, emergent communication, and multi-agent LLM systems. One line of my work studies model internals to identify the functional roots of model behavior, using mechanistic interpretability and experimental paradigms from psycholinguistics (calibration of activation oracles, culture neurons). In another line of work, I study the interaction between models – how communication emerges in populations of agents, and why some communication protocols stay readable while others drift away from human language (compositionality in language learning, emergent languages in LLM agent populations). These strands converge on the key question of how to retain legibility in machine reasoning and communication, and at what cost (auditability vs. accuracy).

I am the principal investigator of the MIST project on scalable mechanistic interpretability for safe and trustworthy LLM agents, funded by the Novo Nordisk Foundation, and I lead the alignment workstream in the Danish Foundation Models project, where we train open multilingual models that excel at Danish using only permissible data (DFM Mimir v1) – contributing to European AI sovereignty.

News

  • August 2026: Pivot programming languages in LLMs accepted at EMNLP 2026.
  • August 2026: Opening talk on mechanistic interpretability at the Danish AI Safety Conference 2026 slides.
  • August 2026: Model release: Mimir V1.
  • April 2026: Prolog-as-a-tool accepted at ACL 2026 Findings. arXiv preprint.
  • February 2026: Three papers accepted to LREC2026.
  • November 2025: The MIST project on mechanistic interpretability for safe and trustworthy LLM agents has been funded by the Novo Nordisk Foundation.
  • November 2025: FlexDeMo has been accepted to AAAI 2026.
  • October 2025: Culture Neurons have been accepted to AACL-IJCNLP 2025 Findings.

Contact: lukas 'at' lpag.de
Design: Adapted from Diane Mounter.
Privacy: No personal data, no cookies.