How is a chatbot trained?

Knowledge Videos

Short animated explainers, organised by subject.

How is a chatbot trained?

Pretraining, fine-tuning and human feedback turn a text predictor into an assistant.

Transcript

First, pretraining. The model reads a huge amount of text, and learns to predict the next word.

The result is a base model. It is great at continuing text, but it is not yet a helpful assistant.

Next, fine tuning on example conversations teaches it to follow instructions.

Now the same question gets a direct answer.

Then comes human feedback. People compare pairs of answers, and a reward model learns their preferences.

The model is nudged toward answers people rate higher. This is called reinforcement learning from human feedback.

Pretraining teaches language, fine tuning teaches the format, and feedback teaches what people prefer. Together, they turn a text predictor into an assistant.

More in this series