Data Science Seminar
Challenges and Progress Towards AI Alignment via Reinforcement Learning from Human Feedback
Daniel Brown
Challenges and Progress Towards AI Alignment via Reinforcement Learning from Human Feedback
| When | Wednesday, February 7, 2024, 1:30 PM – 2:30 PM (MT) |
|---|---|
| Where | GC 2560 (Gardner Commons) |
Abstract
In this talk I will discuss recent progress and challenges towards using human feedback to develop AI systems that behave in ways that are aligned with human preferences. One problem that arises when robots and other AI systems learn from human input is that there is often a large amount of uncertainty over the human’s true intent and the corresponding desired AI behavior. To address this problem, I will discuss prior and ongoing research along three main topics: (1) Enabling robots and other AI systems to learn models of human preferences in ways that are sample efficient and safe, (2) Reward misidentification and causal confusion when learning from human feedback, and (3) Active learning methods for learning better aligned features and reward functions.
Speaker
Daniel Brown
Utah
Daniel Brown is an assistant professor in the School of Computing and Robotics Center at the University of Utah. He completed his postdoc at UC Berkeley in 2023 and he received his Ph.D. in Computer Science from UT Austin in 2020. Daniel’s research focuses on helping robots and other AI systems to safely and efficiently interact with and learn from humans. His research spans the areas of human-robot interaction, reward learning, human-in-the-loop machine learning, and AI safety, with applications in robot manipulation, bio-inspired swarms,autonomous driving,and medical and assistive robotics.
Tags: machine learning robotics
Part of the Data Science Seminar. Something wrong on this page? Edit _data/talks/2024-02-07-daniel-brown.toml.