reinforcement learning from human feedback

training method using human feedback to rank responses and train a reward model that improves model outputs

Discussed in 3 episodes Wikidata (Q115570683) →

Episodes