Catchflow
목록으로
로보트

Risk-Aware Preference Learning for Stochastic Outcomes

2026. 7. 20.
arXiv:2607.15483v1 Announce Type: new Abstract: Learning reward functions from human preferences is a widely used approach for aligning robot behavior with user expectations in human-robot interaction. Most existing approaches assume that humans evaluate uncertain outcomes using expected utility (EU), aggregating outcome utilities linearly with their probabilities. However, behavioral evidence shows that humans are systematically risk-sensitive, overweighting rare negative events and exhibiting
출처