Data Science & AI Lecture Series

Controlling LLM's via Activation Geometry

Amirali Abdullah

<<< All talks

Controlling LLM's via Activation Geometry

When Wednesday, April 15, 2026, 10:30 AM – 11:00 AM (MT)
WhereDinosaur Room (CSC 206)

Abstract

Controlling the behavior of large language models at inference time is an increasingly important problem. In this talk, I present a simple and unified approach to steering model behavior based on activation geometry. By learning a single classifier over hidden representations, we can derive directions that control multiple attributes such as helpfulness, style, or safety, and compose them dynamically without retraining. This framework enables flexible, low cost control of model outputs and highlights a geometric view of representation space beyond fixed linear directions. I will discuss empirical results showing how this approach supports multi attribute control in practice, and briefly outline how such steering mechanisms can be useful in scientific settings where reliable and interpretable model behavior is critical. Our recent followup work suggests that similar activation level interventions can extend across modalities, enabling systematic analysis and control in text to image models through composable operations.

Speaker

Amirali Abdullah

Amirali Abdullah is a Lead AI Researcher at Thoughtworks Inc and a Research Advisor at Martian Learning. His research focuses on the interpretability and control of large language models, with particular emphasis on activation-level steering, representation geometry, and the structure of learned features.

Tags: large language models natural language processing


Part of the Data Science & AI Lecture Series. Something wrong on this page? Edit _data/talks/2026-04-15-amirali-abdullah.toml.