Data Science Seminar

Challenges and Progress Towards AI Alignment via Reinforcement Learning from Human Feedback

Daniel Brown

<<< All talks

Challenges and Progress Towards AI Alignment via Reinforcement Learning from Human Feedback

When Wednesday, February 7, 2024, 1:30 PM – 2:30 PM (MT)
WhereGC 2560 (Gardner Commons)

Abstract

In this talk I will discuss recent progress and challenges towards using human feedback to develop AI systems that behave in ways that are aligned with human preferences. One problem that arises when robots and other AI systems learn from human input is that there is often a large amount of uncertainty over the human’s true intent and the corresponding desired AI behavior. To address this problem, I will discuss prior and ongoing research along three main topics: (1) Enabling robots and other AI systems to learn models of human preferences in ways that are sample efficient and safe, (2) Reward misidentification and causal confusion when learning from human feedback, and (3) Active learning methods for learning better aligned features and reward functions.

Speaker

Daniel Brown

Utah

Daniel Brown is an assistant professor in the School of Computing and Robotics Center at the University of Utah. He completed his postdoc at UC Berkeley in 2023 and he received his Ph.D. in Computer Science from UT Austin in 2020. Daniel’s research focuses on helping robots and other AI systems to safely and efficiently interact with and learn from humans. His research spans the areas of human-robot interaction, reward learning, human-in-the-loop machine learning, and AI safety, with applications in robot manipulation, bio-inspired swarms,autonomous driving,and medical and assistive robotics.

Tags: machine learning robotics


Part of the Data Science Seminar. Something wrong on this page? Edit _data/talks/2024-02-07-daniel-brown.toml.