A full language model training pipeline built on Mistral-7B, taken through every stage rather than stopping at fine-tuning:
- Continued pretraining on Indonesian Wikipedia, to move the base model’s distribution toward Indonesian.
- Supervised fine-tuning on translated Alpaca and OASST, for instruction-following.
- DPO alignment on Wikipedia-derived preference data.
All three checkpoints are published (base, instruct, and aligned), so the effect of each stage can be inspected separately instead of taken on trust.
What I got out of it
This is the project where I learned what end-to-end LLM training actually costs you, and the answer is mostly patience and disk space. The training code is the small part. The real work is data: translating it, cleaning it, discovering that your preference pairs are subtly degenerate, and doing it again.
It is also where the thinking behind Anak Baik started. Building an Indonesian model and then asking what it would refuse to do turned out to be a short path to a research question.