All Topics

RLHF

1 post with this topic

Read DotLM-165M: How I trained a 165M parameter language model from scratch

DotLM-165M: How I trained a 165M parameter language model from scratch

DotLM is a 165M parameter reasoning-capable SLM trained for all four stages of language modeling: Pretraining, Instruction Tuning, Alignment, and Reasoning using synthetically generated STE dataset.

39 min read
FeaturedNLPSLMAgenticAILLM TrainingSFTRLHFReasoningInference OptimizationChatUI