Reinforcement Learning from Human Feedback (RLHF)
Reinforcement learning from human feedback (RLHF) aligns a model's behaviour with human preferences: people compare model outputs, a reward model learns from those judgements, and the model is then optimised to produce outputs the reward model scores highly.
A typical pipeline starts with a pre-trained language model that has been through supervised fine-tuning on example responses. Human labellers compare pairs of model outputs and indicate which is better. A reward model is trained to predict these preferences, and the language model is optimised against it with a reinforcement learning algorithm such as PPO, with a penalty that keeps it close to the original model.
RLHF was central to making instruction-following assistants more helpful and less harmful, as described in OpenAI's InstructGPT work, and preference-based methods are now widely used in training chat models. Alternatives include direct preference optimisation (DPO), which learns from preference pairs without a separate reward model, and reinforcement learning from AI feedback, in which a model generates the preference judgements.
Preference data is expensive and reflects the views and inconsistencies of the labellers. Optimising too hard against a reward model can lead to reward hacking, where outputs score well without being better, and to overly verbose or sycophantic responses. RLHF shapes behaviour but does not guarantee factual accuracy.
Key points
- Aligns models with human preferences using compared outputs
- A reward model learns preferences and the LLM is optimised against it
- Central to instruction-following chat assistants
- DPO is an alternative that skips the separate reward model
Where AiVibe comes in
AiVibe delivers AI and machine learning services, chatbots and virtual assistants with RAG, MCP tools and voice, AI quality management including bias detection and model validation, and the AIMURUGA AI agent, and builds Intel-based edge AI devices using the Intel Distribution of OpenVINO toolkit.