How is a chatbot trained?
Pretraining, fine-tuning and human feedback turn a text predictor into an assistant.
Transcript
First, pretraining. The model reads a huge amount of text, and learns to predict the next word.
The result is a base model. It is great at continuing text, but it is not yet a helpful assistant.
Next, fine tuning on example conversations teaches it to follow instructions.
Now the same question gets a direct answer.
Then comes human feedback. People compare pairs of answers, and a reward model learns their preferences.
The model is nudged toward answers people rate higher. This is called reinforcement learning from human feedback.
Pretraining teaches language, fine tuning teaches the format, and feedback teaches what people prefer. Together, they turn a text predictor into an assistant.
More in this series
1:14How does AI work?
Neurons and weights, learning from mistakes, and predicting the next word.
1:13How does AI read a sentence?
Tokens, embeddings and attention, stacked in layers to score the next token.
1:17What are billions of parameters?
What a parameter is, how big a billion is, and why big models need racks of GPUs.
1:26How do networks learn from errors?
Loss as a landscape, gradient descent and backpropagation.
1:23What does temperature do?
Scores become probabilities, and temperature sharpens or flattens them.
1:19How does AI draw pictures?
Diffusion models add noise to learn, then remove it to create, steered by a prompt.
1:15How can AI use your own documents?
Retrieval-augmented generation: embed, retrieve, augment the prompt, generate.
1:14Why does AI make things up?
Likely is not the same as true: gaps, snowballing errors, and what helps.
1:16How much can an AI remember?
The context window, forgetting, the cost of long inputs, and workarounds.
1:09What is an AI agent?
A model in a loop with tools and guardrails.
1:12Can AI be biased?
Skewed data, where bias comes from, proxies, and how to audit and fix it.
1:10What is overfitting?
Underfit, good fit and overfit curves, train vs test error, and the fixes.