Listen: Alignment

The effort to make a model behave as intended (helpful, honest, harmless) via RLHF and related methods; it sets a model's defaults about refusals, trust, and which sources it treats as credible.

Listen

Transcript

A raw AI model has incredible language skills, but it does not know how it is supposed to behave. That is where alignment comes in. Alignment is the process of training a model to be helpful, honest, and harmless. It ensures the AI follows human instructions and values, instead of just predicting the next word.

Developers achieve this through methods like Reinforcement Learning from Human Feedback, where the model learns from human preferences, or even by following a written constitution of principles.

This alignment process sets the model's default behaviors. It decides how the AI handles uncertainty, how it refuses harmful prompts, and which sources it trusts. This has a massive impact on what information actually reaches the public. An AI's ideas about which brands are reputable and which sources are credible are baked in during alignment. In other words, to be recommended by an AI, you first have to fit the model's learned definition of what is safe and trustworthy.