로보트
Risk-Aware Preference Learning for Stochastic Outcomes
2026. 7. 20.
arXiv:2607.15483v1 Announce Type: new
Abstract: Learning reward functions from human preferences is a widely used approach for aligning robot behavior with user expectations in human-robot interaction. Most existing approaches assume that humans evaluate uncertain outcomes using expected utility (EU), aggregating outcome utilities linearly with their probabilities. However, behavioral evidence shows that humans are systematically risk-sensitive, overweighting rare negative events and exhibiting
출처
- [GL] arXiv cs.RO · 2026-07-20T04:00:00.000Z
