← Writing library

DotLM-165M: How I trained a 165M parameter language model from scratch

DotLM is a 165M parameter reasoning-capable SLM trained for all four stages of language modeling: Pretraining, Instruction Tuning, Alignment, and Reasoning using synthetically generated STE dataset.

39 min readShanmukha Sainath

DotChat Demo

This blog assumes you have a basic understanding of LLMs and how they're trained. If not, I would recommend you to first go through this llm-course by Maxime Labonne.

Why DotLM?

Last time I wrote a blog post (about the STE Dataset), OpenClaw was trending. Now, as I write this, Andrej Karpathy's "LLM Knowledge Bases" concept is the latest focus. People are finding innovative ways to integrate LLMs into their daily workflows; I personally use Claude Code for many of my projects, both personal and professional. It’s remarkable to see what these models are capable of, and they only continue to improve. Anthropic's latest internal model, Mythos, is particularly mind-blowing, while not yet public, its reported capabilities are incredible and dangerous.

As I mentioned in my previous blog post, I have always been fascinated by how a set of matrices can be trained to almost replace an entry-level software engineer. As an ML researcher, I understand the mathematics behind them, still it amazes me how they work. Few years back I wouldn't have imagined that a Language model would be able to do so much and has capabilities to even solve un-solved Maths and Physics problems (GPT-5.4 cracked a 20 year old Maths problem). This led me to build a miniature version of it myself from the ground up. This led me to develop a tiny, reasoning-capable language model—often referred to as an SLM (Small Language Model).

If you're also interested in training an LLM from scratch, I highly recommend the Smol Training Playbook from Hugging Face.

Dataset

There are plenty of open-source datasets like FineWeb, Wikipedia, and RedPajama that are perfect for pretraining. Since my objective was to train a reasoning-capable model (a characteristic usually reserved for much larger systems) using fewer parameters, I needed a high information density dataset. Extensive research on existing datasets and their composition led me to create my own STE dataset, covering all four stages of LLM training. The synthetic data focuses on core concepts across various domains like Physics, Chemistry, Biology, Mathematics, Computer Science, and General Knowledge. I also added few samples from FineWeb, Wikipedia to the pretraining dataset (only because I haven't generated enough pretraining tokens).

Samples from STE Dataset

{
    "text": "Suspensions occur when tiny particles are mixed into a liquid but do not dissolve, remaining suspended in the liquid, creating a cloudy or opaque solution. This phenomenon can be vividly observed in various everyday situations, such as seeing muddy water or orange juice with pulp. Both are excellent examples of suspensions; the particles are distributed throughout the liquid but can settle out over time if left undisturbed. At an airport, envision the hustle and bustle of travelers rushing through the terminal, bags in tow, while staff members work diligently at check-in counters and security stations. As you stand near a café, you notice a freshly squeezed orange juice sitting on the counter. Unlike clear, bottled juice that appears as a smooth, uniform liquid, this fresh juice has bits of pulp suspended within it. The pulp consists of tiny pieces of orange flesh and membranes that do not dissolve in the juice. Instead, they remain dispersed throughout, creating that familiar cloudy appearance. If you were to let the juice sit undisturbed for a while, you'd see the pulp gradually sinking to the bottom of the glass, demonstrating the principle of settling, which is characteristic of suspensions. Consider another scenario involving a muddy puddle formed after a rainstorm at the airport. When rainwater mixes with dirt, the resulting muddy water is a classic suspension. The fine particles of soil are temporarily held within the water but do not dissolve, giving the water a brownish hue. If left alone, the heavier soil particles settle to the bottom, creating a distinct separation between the muddy water above and the clear water below. This illustrates how suspensions differ from true solutions, where the solute fully dissolves and cannot be separated through simple settling. The key distinction between suspensions and other mixtures, like solutions or colloids, lies in the particle size and behavior. In a solution, such as sugar dissolved in water, the sugar molecules break apart into individual units so small that they cannot be seen and remain evenly distributed, with no tendency to settle out. On the other hand, in a colloid, such as whipped cream, the particles are larger than those in a solution but smaller than those in a suspension, preventing them from settling quickly. Knowing this, one can appreciate why suspensions are utilized in specific applications where separation or settling is acceptable. An example of this is the preparation of certain medications in liquid form. Some drugs, especially those containing active ingredients like antibiotics in suspension, are formulated to ensure they remain evenly mixed for a certain period. Patients must shake the bottle before use to ensure the particles are properly distributed. If the particles were in a solution, they would not be visible or would be rapidly absorbed, potentially leading to dosing inconsistencies. At the airport, travelers benefit from these principles of suspension in various products. For example, salad dressings often contain suspended oil, vinegar, and spices. If left undisturbed, the oil would separate and float above the dressing, demonstrating the settling behavior typical of suspensions. The need to shake or stir the dressing before use ensures a uniform flavor experience, much like shaking the orange juice to redistribute the pulp. Understanding suspensions not only explains everyday phenomena but also informs practical solutions in industries ranging from food and beverage to pharmaceuticals. The careful design of these mixtures highlights the importance of particle interactions, stability, and the ability to utilize settling properties for effective applications. Whether it's a refreshing glass of orange juice or a medicinal syrup, the science of suspensions plays a crucial role in our daily lives, all while reminding us how particles behave in a liquid medium.",
    "topic": "material_science",
    "subtopic": "suspension",
    "concept": "Particles mixed but not dissolved",
    "model": "gpt-4o-mini",
    "provider": "openai",
    "token_count": 718,
    "timestamp": "2026-01-20T13:05:33.911030"
  },
 {
  "messages": [
      {
          "role": "user",
          "content": "If I'm hiking and I notice a small stream flowing down from a mountain, what will happen to that stream as I continue to hike higher up the mountain?"
      },
      {
          "role": "assistant",
          "content": "The relevant principle here is gravity, which pulls water downhill. As you hike higher up the mountain, the stream will appear to be getting smaller and smaller, eventually disappearing into the ground or becoming so small that it's no longer visible. This is because you're moving upstream, towards the source of the water. Think of it like walking up an inclined plane; the higher you go, the less water you'll see, as it's all flowing downhill behind you."
      }
  ],
  "topic": "earth_science",
  "subtopic": "mountains",
  "query_type": "predictive",
  "model": "mistral-small-3.1",
  "provider": "openrouter",
  "timestamp": "2026-01-21T12:05:25.229851"
}
{
    "prompt": "What is the difference between an arc and a chord in a circle, and how do their measurements relate to the circle's overall dimensions?",
    "chosen": "In a circle, a chord is a straight line segment whose endpoints lie on the circle, effectively cutting through the interior, while an arc is a portion of the circumference of the circle, defined by two endpoints on the circle. The length of a chord can be calculated using the radius and the angle subtended at the center, while the length of an arc is proportional to the central angle it subtends, typically expressed as a fraction of the circle's circumference. The relationship between these measurements is rooted in the geometry of the circle, where the radius remains constant and defines both the length of the chord and the arc based on their respective angles.",
    "rejected": "In a circle, an arc and a chord are two distinct features with distinct goals. The arc \"wants\" to span a portion of the circle's circumference, and its length is directly proportional to the angle it subtends at the center of the circle. Conversely, a chord \"seeks\" to intersect the circle at two points, and its length is inversely related to the radius of the circle. Specifically, the ratio of an arc's length to its corresponding chord's length is constant, reflecting the circle's inherent geometry and desire for harmony.",
    "topic": "geometry_space",
    "subtopic": "arc",
    "error_type": "teleological",
    "judge_score": 0.9,
    "judge_reasoning": "Response A uses teleological reasoning by attributing desires ('wants' and 'seeks') to geometric features, which is misleading and not scientifically accurate. Response B correctly describes the definitions and relationships without such errors.",
    "model_chosen": "gpt-4o-mini",
    "model_rejected": "llama-3.1-8b",
    "timestamp": "2026-01-21T13:59:10.313421"
  }
{
    "input": "In a coordination game, why do people tend to choose the most obvious option? For example, if two people have to choose between meeting at the Eiffel Tower or a small cafe in Paris without communicating, they both go to the Eiffel Tower. Is it because they think the other will choose it?",
    "thought_trace": "We are dealing with a coordination game scenario. The user describes a situation where two people must choose between two meeting points (Eiffel Tower vs. a small cafe) without communication and notes that they tend to choose the Eiffel Tower. The user asks if this is because each thinks the other will choose it.\n\n**Understanding**: The user is asking about the reasoning behind the choice of the \"obvious\" option in a coordination game. The example is a classic one in game theory: two players must independently choose a meeting place and they succeed only if they choose the same place. The Eiffel Tower is considered the focal point.\n\n**Misconception Check**: The user's question seems to imply that the reason is that each player thinks the other will choose the Eiffel Tower. This is partially correct, but we must delve deeper. The key is not just about what one thinks the other will do, but also about what each thinks the other thinks, and so on. Moreover, the reason why they both land on the Eiffel Tower is because it is a salient or focal point. The user might not be aware of the concept of focal points.\n\n**Relevant Principles**:\n1. **Coordination Game**: A game where players benefit from choosing the same action. The payoff is higher when they coordinate, regardless of which option they choose, as long as they choose the same one.\n2. **Nash Equilibrium**: In this game, both choosing Eiffel Tower and both choosing the cafe are Nash equilibria (if they both choose Eiffel Tower, neither has an incentive to switch to the cafe, and vice versa). However, there might be multiple equilibria and the problem is how players select one without communication.\n3. **Focal Point (Schelling Point)**: A solution that people tend to choose by default because it stands out as prominent, obvious, or culturally salient. The concept was introduced by Thomas Schelling. The Eiffel Tower is a famous landmark, so it naturally serves as a focal point.\n\n**Reasoning**:\n- Step 1: The players want to coordinate. They must choose independently without communication.\n- Step 2: Each player must form a belief about what the other will choose. But since they cannot communicate, they must rely on common knowledge and salience.\n- Step 3: The Eiffel Tower is a world-famous monument, while the small cafe is not. Therefore, the Eiffel Tower is more salient. Each player knows that the Eiffel Tower is a well-known landmark and thus likely to be chosen by the other.\n- Step 4: Moreover, each player knows that the other player is also reasoning in the same way. So, each thinks: \"The Eiffel Tower is the obvious choice, so the other will probably choose it. Therefore, I should choose the Eiffel Tower too.\"\n- Step 5: This reasoning leads both to choose the Eiffel Tower, achieving coordination.\n\n**Analogy**: Imagine you and a friend are separated in a large department store and you want to meet without cell phones. You both know that the main entrance is the most prominent place. So, you both go to the main entrance. The main entrance is the focal point because it is the most obvious meeting spot.\n\n**Conclusion**: Yes, people tend to choose the most obvious option (like the Eiffel Tower) because they believe the other person will choose it. But the deeper reason is that the obvious option serves as a focal point (or Schelling point) due to its salience, which allows players to coordinate without communication. This is a well-established concept in game theory.\n\nFinal Answer:\nIn coordination games without communication, people often choose the most obvious option because it serves as a focal point (or Schelling point). The Eiffel Tower in your example is a prominent landmark, so each person expects the other to choose it. This mutual expectation leads both to select the Eiffel Tower, resulting in successful coordination. Therefore, it is indeed because they think the other will choose it, but this is driven by the salience of the option.",
    "output": "### Step-by-Step Reasoning:\n\n**Understanding**:  \nThe user is asking about behavior in a coordination game\u2014specifically, why people gravitate toward the \"most obvious\" option (like the Eiffel Tower) when they must match choices without communication. This scenario is a classic example from game theory, where success depends on both players choosing the same action independently.\n\n**Misconception Check**:  \nThe user\u2019s phrasing (\"Is it because they think the other will choose it?\") touches on the core idea but misses a deeper layer. The key isn\u2019t just individual belief about the other\u2019s choice\u2014it\u2019s about **shared expectations** and **salience**. Players don\u2019t merely guess; they rely on a culturally or logically prominent \"focal point\" that both recognize as the default. Assuming players act purely on random guesses or personal preferences would be a misconception\u2014the solution arises from common knowledge and social cues.\n\n**Relevant Principles**:  \n1. **Coordination Games**: Players earn the highest payoff only if their choices match (e.g., meeting at the same location). Multiple Nash equilibria exist (both choosing Eiffel Tower or both choosing the caf\u00e9), but no single equilibrium is inherently \"better\"\u2014coordination itself is the goal.  \n2. **Focal Points (Schelling Points)**: Introduced by Thomas Schelling, this concept explains how players coordinate without communication by choosing the most conspicuous option. Salience comes from cultural prominence, simplicity, or uniqueness (e.g., the Eiffel Tower is iconic, while a small caf\u00e9 is obscure).  \n3. **Common Knowledge**: Players don\u2019t just think about what others will do; they consider what others think they will do, creating a chain of shared reasoning. The Eiffel Tower works because both know it\u2019s famous, both know the other knows it\u2019s famous, and so on.  \n\n**Reasoning**:  \n1. **Problem Setup**: Two players must choose between two locations (Eiffel Tower or caf\u00e9) independently. Payoffs are identical if they match (e.g., +10 for meeting, 0 for mismatching).  \n2. **Why Obvious Options Win**:  \n   - **Salience Creates Predictability**: The Eiffel Tower is uniquely prominent\u2014it\u2019s a global symbol of Paris. Players use this salience to break symmetry.  \n   - **Mutual Expectations**: Player A thinks: \"The Eiffel Tower is obvious, so Player B will likely choose it. Therefore, I\u2019ll choose it too.\" Player B reasons identically. This isn\u2019t a blind guess; it\u2019s a self-reinforcing loop.  \n   - **Iterated Beliefs**: Player A also thinks: \"Player B expects me to choose the Eiffel Tower because it\u2019s obvious.\" This higher-order thinking solidifies the choice.  \n3. **Why Other Options Fail**: The caf\u00e9 lacks salience\u2014it could be any of thousands in Paris. Without a distinguishing feature, it\u2019s harder to form mutual expectations.  \n4. **Experimental Evidence**: Real-world studies (e.g., Schelling\u2019s experiments) show people default to salient options like landmarks, round numbers, or simple solutions >80% of the time, even with identical payoffs.  \n\n**Analogy**:  \nImagine two friends separated in a mall, needing to meet without phones. They\u2019ll likely choose the main entrance, not a random store\u2014because it\u2019s the most visible, \"default\" spot. Both know it\u2019s obvious, both know the other knows, so it becomes the natural choice.  \n\n**Conclusion**:  \nYes, people choose the most obvious option (like the Eiffel Tower) primarily because they believe the other person will choose it\u2014but this belief stems from the option\u2019s **salience as a focal point**. The Eiffel Tower is culturally prominent, making it a predictable choice through shared knowledge. In game theory terms, this is a **Schelling point**: a solution that stands out due to uniqueness, tradition, or simplicity, enabling coordination without communication. While personal preferences or randomness could play a role, salience is the dominant driver in real-world behavior.",
    "topic": "economics_game_theory",
    "subtopic": "coordination_game",
    "has_misconception": false,
    "model": "deepseek-r1",
    "provider": "deepinfra",
    "timestamp": "2026-01-18T17:45:26.263925",
    "thought_trace_compressed": "- Step 1: Players must coordinate their choices independently without communication.\n- Step 2: Each player forms beliefs about the other's choice based on common knowledge and salience.\n- Step 3: The Eiffel Tower is a prominent landmark, making it more salient than the small cafe.\n- Step 4: Each player reasons that the other will likely choose the Eiffel Tower due to its obviousness.\n- Step 5: This mutual reasoning leads both players to choose the Eiffel Tower, achieving coordination.",
    "output_compressed": "Yes, people choose the Eiffel Tower because they believe the other will choose it, driven by its salience as a focal point in the coordination game.",
    "compression_model": "gpt-4o-mini"
}

Pretraining Sample

1 / 4

In case of Reasoning, thought_trace and output responses were compressed using gpt-4o-mini as the original responses were too long.

DotLM Architecture

DotLM ArchitectureDotLM Architecture

Illustration inspired from Sebastian Raschka's LLM Gallery

The DotLM architecture is a decoder-only transformer inspired by Qwen 3. While key components like SwiGLU, RMSNorm, RoPE, and GQA are borrowed from Qwen 3, several architectural modifications were introduced to enable reasoning at a 165M parameter scale. My selection process focused on parameter-efficient components to minimize the model's size.

Specs

ConfigurationValue
Total Parameters165,342,720 (~165M)
Layers24
dmodeld_{model}768
dffd_{ff} (Intermediate Dim)2,048 (2.67x ratio)
Attention Heads6
KV Heads2 (Grouped Query Attention)
Head Dimension128
Context Length4,096
TokenizerBPE
Vocabulary Size16,384
Activation Function (FeedForward)SwiGLU
NormalizationRMSNorm (ϵ=106\epsilon = 10^{-6})
Positional EmbeddingRoPE (θ=10,000\theta = 10,000)
Weight TyingEnabled (Input/Output Embeddings)
Training PrecisionBF16 Mixed
DotLM Model Specifications

The optimal configuration was determined through hyperparameter tuning using Autoresearch (detailed later in this post).

Grouped Query Attention (GQA)

Grouped Query Attention (GQA) bridges the gap between Multi-Head Attention (MHA) and Multi-Query Attention (MQA) by grouping query heads to share a single key-value head. This significantly reduces the memory bandwidth required for the KV cache during inference while maintaining quality close to MHA. In MHA, we have HH query, key, and value heads. In GQA, we have HH query heads but only GG key/value heads (1<G<H1 < G < H).

Special cases of GQA are:

  • Multi-Head Attention (MHA): G=HG = H
  • Multi-Query Attention (MQA): G=1G = 1
AttentionAttention
Multi-Head AttentionMulti-Head Attention
Multi-Query AttentionMulti-Query Attention
Grouped-Query AttentionGrouped-Query Attention

Attention

1 / 4

SwiGLU

SwiGLU is an activation function that combines the SWISH (or SiLU) activation with a Gated Linear Unit (GLU). It has been empirically shown to offer better performance than standard ReLU or GeLU. It dynamically gates the flow of information using a learned linear projection passed through a Swish function.

SwiGLU(x)=Swish(xW1)(xW2)Swish(z)=zσ(z)\text{SwiGLU}(x) = \text{Swish}(x W_1) \otimes (x W_2) \\[15pt] \text{Swish}(z) = z \cdot \sigma(z)

RMSNorm

Root Mean Square Normalization (RMSNorm) is a simpler and faster alternative to standard Layer Normalization. It removes the mean-centering step, hypothesizing that the scaling invariance is the most important component of LayerNorm's success. This reduces computational overhead while maintaining training stability.

RMSNorm(x)=x1di=1dxi2+ϵγ\text{RMSNorm}(x) = \frac{x}{\sqrt{\frac{1}{d} \sum_{i=1}^{d} x_i^2 + \epsilon}} \odot \gamma

where xx is the input vector of dimension dd, ϵ\epsilon is a small constant for numerical stability, and γ\gamma is a learned scale parameter.

Tie Embeddings

Weight tying shares the parameters between the token embedding layer and the language modeling head (pre-softmax projection). This reduces the total parameter count by reusing the learned linguistic representations for both mapping tokens to vectors and vectors back to logits.

Woutput=WembeddingTW_{\text{output}} = W_{\text{embedding}}^T

Rotary Positional Embeddings (RoPE)

RoPE encodes positional information by rotating the queries and keys in the complex plane, rather than adding an absolute positional vector. This preserves relative distances and provides better extrapolation capabilities for longer context windows. For a given position mm and a 2D feature vector x=[x1,x2]Tx = [x_1, x_2]^T, it applies the rotation matrix:

RoPE(x,m)=(cos(mθ)sin(mθ)sin(mθ)cos(mθ))(x1x2)\text{RoPE}(x, m) = \begin{pmatrix} \cos(m\theta) & -\sin(m\theta) \\ \sin(m\theta) & \cos(m\theta) \end{pmatrix} \begin{pmatrix} x_1 \\ x_2 \end{pmatrix}

For higher dimensions, it pairs consecutive dimensions and rotates each pair with different frequencies θi=100002i/d\theta_i = 10000^{-2i/d}.

RoPERotary Positional Embedding Analysis

Autoresearch

As its name suggests, Autoresearch is a framework designed to automate hyperparameter tuning for neural networks using AI agents. This concept was proposed by Andrej Karpathy in a recent X post. The primary advantage of Autoresearch over traditional methods (such as Grid Search or Bayesian Optimization) is its ability to perform unstructured, autonomous code evolution rather than merely searching within a predefined parameter space. This gives researchers greater control over design criteria and the ability to discover novel architectural patterns.

The idea: give an AI agent a small but real LLM training setup and let it experiment autonomously overnight. It modifies the code, trains for 5 minutes, checks if the result improved, keeps or discards, and repeats.

Workflow

DotLM WorkflowDotLM Training Workflow using Autoresearch

We start with a base configuration file and a pretraining dataset. Autoresearch optimizes architecture settings to identify the best model, which is then used to train subsequent stages: SFT, Alignment, and Reasoning. In these later stages, Autoresearch is used to optimize training settings. While validation loss is the primary metric for parameter tuning, the agent also evaluates the quality of generated samples to ensure the model is learning meaningful patterns (a predefined set of 20 prompts is used for these quality checks).

Pretraining - Knowledge Acquisition

Pretraining is the most critical, resource-intensive, and time-consuming stage of training an LLM. During this phase, the model learns about language, facts, and the world. Once trained, the model functions as an auto-complete engine, predicting the next token based on previous sequences.

Using approximately 246M tokens from the STE pretraining dataset, the model was trained for 10 epochs (totaling ~2.46B tokens), which aligns with Chinchilla optimality for a 165M parameter model. The hyperparameters optimized were:

  • Model Architecture
    • Max Sequence Length
    • Hidden Dim
    • Attention Heads
    • KV Heads
    • RoPE theta
    • Num layers
  • Training Hyperparameters
    • Batch Size + Accumulation
    • Learning Rate
    • WSD scheduler (Warmup-Stable-Decay) settings
    • Gradient Clipping
    • Beta1, Beta2 for AdamW optimizer
HyperparameterValue
Dataset352k samples (~246M tokens)
Training Tokens~2.46B (10 Epochs)
Total Steps~27,500 Optimizer Steps
Batch Size64 (Accum=2, Effective=128)
Learning Rate1.5×1031.5 \times 10^{-3}
Weight Decay0.01
SchedulerWSD (Warmup-Stable-Decay)
Warmup Steps2,750 (~10%)
Stable Steps20,000 (~73%)
Decay Steps4,750 (~17%)
Gradient Clipping0.5
DotLM Pretraining Training Settings
Pretraining sampleSample text generation after Pretraining
Pretraining autoresearchAutoresearch progress during Pretraining
Pretraining autoresearch changesChanges in architecture and training settings during Pretraining
Pretraining metricsTraining history for the best settings

Sample text generation after Pretraining

1 / 4

The figures above illustrate the Autoresearch progress and the training history for the optimal settings. Key architectural optimizations discovered by Autoresearch include:

  • Max Sequence Length: Reduced from 1280 to 768 (aligning with the data distribution)
  • Hidden Dim: Increased from 1536 to 2048 (a 2.67x dmodeld_{model} ratio)
  • Attention Heads: Reduced from 12 to 6 (a head dimension of 128 outperformed 64 in all tests)
  • KV Heads: Standardized to 2 (GQA with 3 groups)

Addressing Misaligned Autoresearch Objectives

In the Autoresearch progress plot, you may notice an increase in validation loss after Experiment 30. This occurred because, starting at Experiment 21, the agent learned that reducing model depth lowered validation loss. Since I had not yet specified that the model required reasoning capabilities, it began sacrificing depth to minimize loss. I then intervened to update the objective, guiding the agent back in the right direction.

Be very specific about what you want the model to achieve. Otherwise, the agent will find shortcuts to reach the assigned objective, which may not align with your goals.

Supervised Fine-Tuning (SFT) - Learning to Converse

In this stage, the base model is fine-tuned on a high-quality instructional dataset to learn conversational formats and follow human instructions. We utilized 25,700 samples from the STE SFT dataset. While we limited the Pretraining Autoresearch phase to 20 minutes, we allowed the SFT, Alignment, and Reasoning Autoresearch phases to run until the completion of training. During SFT, Autoresearch focused primarily on optimizing learning rates and WSD scheduler settings. Many initial trials (conducted during the V1 phase of Autoresearch, not shown below) were discarded due to the low quality of generated samples.

HyperparameterValue
Dataset25,700 samples
Batch Size16 (Accum=4, Effective=64)
Learning Rate3.0×1043.0 \times 10^{-4}
Epochs5
Total Steps~1,988 Optimizer Steps (Best ckpt at 800 steps)
SchedulerCosine with 100 Warmup Steps
Weight Decay0.01
DotLM SFT Training Settings
SFT sampleSample text generation after SFT
SFT autoresearchAutoresearch progress during SFT
SFT autoresearch changesChanges in training settings during SFT
SFT metricsTraining history for the best settings

Sample text generation after SFT

1 / 4

Alignment - Adhering to Human Preferences

Alignment via Direct Preference Optimization (DPO) ensures the model's outputs reflect human intent and preferences. Since each training sample contains both a "chosen" and a "rejected" response, the memory requirement per sample is effectively doubled. During this stage, we observed a significant negative correlation between validation loss and the quality of generated samples: across all experiments, the quality of generated samples degraded even as validation loss continued to improve.

Standard hyperparameter tuning typically fails to address such inverse relationships; this is precisely where Autoresearch excels.

HyperparameterValue
Dataset7,172 samples
Batch Size8 (Accum=4, Effective=32)
Learning Rate5.0×1055.0 \times 10^{-5}
DPO Beta0.2 (Strong preference signal)
Epochs10
Total Steps~2,220 Optimizer Steps (Best ckpt at 600 steps)
SchedulerCosine with 111 Warmup Steps
Val Steps300 (~3 checks/epoch)
DotLM Alignment Training Settings
Alignment sampleSample text generation after Alignment
Alignment autoresearchAutoresearch progress during Alignment
Alignment autoresearch changesChanges in training settings during Alignment
Alignment metricsTraining history for the best settings

Sample text generation after Alignment

1 / 4

Reasoning - Thinking Before Speaking

The final stage involves training the model to generate Chain-of-Thought (CoT) reasoning enclosed in <think>...</think> tags. This phase utilizes a specialized reasoning dataset with compressed thought traces and outputs to fit within the 768-token context window. During this stage, Autoresearch varied the learning rate and scheduler parameters simultaneously, in contrast to the SFT phase where the learning rate was optimized independently at the start.

HyperparameterValue
Dataset6,300 samples
Batch Size16 (Accum=2, Effective=32)
Learning Rate5.0×1055.0 \times 10^{-5}
Epochs10
Total Steps~1,940 Optimizer Steps (Best ckpt at 800 steps)
SchedulerCosine with 50 Warmup Steps
Weight Decay0.01
Val Steps200 (~2 checks/epoch)
DotLM Reasoning Training Settings

How Autoresearch Refined the Training Objective

Initially, the objective was set to minimize validation loss while maximizing the quality of generated samples. However, the agent observed that certain training configurations resulted in outputs missing the required </think> closing tag. Consequently, the agent updated its objective to discard settings that failed to reliably close the reasoning block. This was one of the most significant modifications the agent introduced during the reasoning step, as it directly improved the model's structural coherence and reasoning performance.

Reasoning sampleSample text generation after Reasoning
Reasoning autoresearchAutoresearch progress during Reasoning
Reasoning autoresearch changesChanges in training settings during Reasoning
Reasoning metricsTraining history for the best settings

Sample text generation after Reasoning

1 / 4

Inference Optimization

Existing Methods

Quantization

Quantization reduces the precision of the model's weights (e.g., from 16-bit floats like FP16 or BF16 to 8-bit or 4-bit integers like INT8 or INT4). This sharply decreases the memory footprint and increases the memory bandwidth utilization, allowing for faster generation times and the ability to run on consumer hardware with less VRAM, often with negligible degradation in model quality.

Flash Attention

Flash Attention is a hardware-aware exact attention algorithm. It optimizes memory access patterns by tiling the attention computation, which drastically reduces the number of read/write operations to the GPU's High-Bandwidth Memory (HBM). This results in much faster execution and significantly lower memory consumption compared to standard attention implementations, particularly for long context lengths.

KV Cache

The KV Cache stores the computed Key (KK) and Value (VV) representations of previous tokens during autoregressive generation. Since predicting the next token requires attention over all preceding tokens, caching these values instead of recomputing them at every step transforms an O(N2)O(N^2) generative computation cost into an O(N)O(N) operation per step.

  • Static KV Cache: Pre-allocates a fixed memory buffer for the maximum supported sequence length. This approach eliminates the overhead of dynamic allocation during generation and ensures deterministic memory usage, though it can be wasteful for shorter sequences.
  • Dynamic KV Cache: Allocates memory for KV pairs on-the-fly, typically managed through paging mechanisms like PagedAttention. This significantly reduces memory fragmentation and allows for higher batch sizes and throughput by only consuming memory proportional to the actual sequence length.

Speculative Decoding

Speculative Decoding is a latency optimization technique that employs a smaller, faster "draft" model alongside the larger target LLM. The draft model rapidly generates a sequence of potential next tokens. The target model then verifies these tokens efficiently in a single forward pass. Tokens that are accepted are yielded immediately, which allows multiple tokens to be produced per step and significantly accelerates inference without altering the final output probabilities.

Paged Attention

Inspired by virtual memory paging in operating systems, Paged Attention splits the contiguous KV cache into fixed-size blocks ("pages") allocated dynamically in non-contiguous memory spaces. This largely eliminates memory fragmentation and avoids pre-allocating worst-case sequence lengths, which substantially increases the maximum batch size and throughput when serving LLMs concurrently.

CUDA Graphs

CUDA Graphs optimize the CPU overhead of launching GPU kernels. In traditional execution, the CPU must launch each kernel layer by layer, which introduces a latency bottleneck, especially for small models or small batch sizes where GPU execution outpaces the CPU launch time. CUDA Graphs record the entire sequence of GPU operations into a single topological graph and launch it with a single CPU instruction, effectively removing kernel launch latency overhead.

DotLM Inference Analysis

Speculative decoding works best when the draft model is significantly faster (e.g., 10-100x) than the target model. In the case of DotLM, using a draft model is impractical due to the model's already small size. Small models like DotLM are primarily compute-bound rather than memory-bound; the overhead of non-contiguous memory access can degrade performance, whereas simple linear memory access is much faster.

I conducted an extensive analysis of DotLM inference performance on an NVIDIA L40S GPU by combining several optimization techniques: Quantization, Flash Attention, KV Caching, and CUDA Graphs. I evaluated both FP32 and BF16 precision models, as 8-bit or 4-bit quantization might not be necessary for a model of this scale. The figure below compares the throughput and memory usage across various DotLM configurations.

DotLM Inference AnalysisDotLM Inference benchmarking

The combination of the BF16 precision model with a Static KV Cache and CUDA Graphs (using a batch size of 8) yielded the best throughput at 1460 tokens/sec with a memory footprint of 1139MB. CUDA Graphs notably improved throughput with only a marginal increase in memory usage.

Evaluation

DotLM was evaluated against multiple benchmarks to assess its performance across diverse tasks. Results are compared against the GPT-2 baseline (with Qwen 0.5B included for reference) to contextualize the improvements realized by DotLM.

1. Commonsense and Linguistic Reasoning

  • HellaSwag: This task requires the model to predict the most plausible ending to a sentence. Scoring 32% is a significant improvement over the GPT-2 baseline (~29%), demonstrating that the model understands physical and social contexts.
  • Winogrande: Focused on pronoun resolution. While DotLM slightly outperforms GPT-2 at 52%, it remains just 2% above random chance, placing it in the realm of borderline random guessing.

2. Academic Scientific Knowledge

  • SciQ: This is where DotLM-165M truly shines. Outperforming GPT-2 by 8%, the model demonstrates a strong ability to retrieve scientific facts and textbook knowledge.
  • ARC-Easy: In grade-school science questions, DotLM again leads the GPT-2 baseline. This suggests that smaller models can excel as "compact encyclopedias" when trained on high-quality, high-density data.

3. The Reasoning Frontier

  • GPQA: GPQA is designed to be "Google-proof," often requiring expert-level knowledge. A score of 22% (slightly below random) is entirely expected at this scale. Mastery of expert reasoning remains the primary differentiator between sub-billion parameter models and large-scale LLMs.
BenchmarkGPT-2 (124M)DotLM (165M)Qwen (500M)
HellaSwag29.0%32.0%48.0%
ARC-Easy40.0%42.0%60.0%
Winogrande50.2%52.0%56.2%
SciQ50.0%58.0%85.0%
GPQA24.0%22.0%26.5%
Comparison of DotLM with GPT-2 and Qwen on various benchmarks

DotChat - The ChatUI for DotLM

DotChat is a chat interface built to interact with DotLM-165M. It consists of a Next.js frontend deployed on Vercel and a Python backend deployed on Modal as a serverless GPU endpoint. The backend exposes a single streaming endpoint via FastAPI and Server-Sent Events (SSE), keeping the architecture minimal and cost-efficient.

Inference Stack

The backend inference pipeline is composed of three main modules: Tokenizer, Inference Engine, and Chat Manager. On container startup, the engine loads the model weights from a persistent Modal Volume and initializes the chat manager. The model stays warm in memory for 180 seconds between requests, eliminating repeated cold-start overhead for active user sessions.

Streaming Token Generation

The inference engine drives generation token-by-token using a manual KV cache loop. This avoids the overhead of a standard model.generate() call and enables true streaming:

  1. Prefill: The full conversational prompt is passed in one forward pass, populating the KV cache for all layers.
  2. Decode loop: Each subsequent step feeds only the single most-recent token, with the KV cache providing the full attention history.
  3. Incremental Decoding: The output tokens are dynamically decoded and mapped back to strings, allowing the server to stream chunks of text immediately via a Server-Sent Events connection.

Adding Conversational Capabilities

Maintaining a coherent conversation is one of the most significant challenges for Small Language Models (SLMs). Models with fewer parameters are highly sensitive to noise in the prompt and easily get "confused" when forced to process long, irrelevant historical context. If you simply append the entire chat history to a 165M model, it often begins to hallucinate or lose focus on the current query.

To solve this, DotChat uses a Dynamic Context Injection approach. Instead of blindly passing history, the system performs a real-time "topic shift" analysis to decide if the previous turn is actually relevant to the current user query.

The Topic Classification

When a user sends a query in conversational mode, the ChatManager executes two very fast, partial inference passes (generating only ~15 tokens) to extract Topic Vectors:

  1. Query Only: A vector representing the semantic intent of the query in isolation.
  2. Query + Previous Response: A vector representing the intent when the query is prefixed with the previous assistant's response.

By calculating the Cosine Similarity between these two vectors, we can determine if the conversation is still on the same topic.

  • Similarity ≥ 0.5: The query is semantically linked to the previous answer. A truncated version of the previous response is injected as a context prefix.
  • Similarity < 0.5: A topic shift is detected. The query is treated as a fresh start, and no context is injected.

This approach ensures that DotLM-165M only receives the context it needs to stay coherent, effectively bypassing its inherent small-scale context limitations. Though it's not as good as LLMs, it seems to work well for most of the queries that I have tested.

Conversational Pipeline Flowchart

DotChat Conversational FlowDotChat Conversational Pipeline Flowchart

ChatUI Details

The frontend is a Next.js app that consumes the Server-Sent Events (SSE) stream and updates the user interface in real time. It parses incoming text chunks and dynamically separates the <think>…</think> reasoning block from the final answer, giving users complete visibility into the model's thought process.

How much did it cost?

  • Data Creation: Data is created using APIs from OpenAI, Openrouter and Deepinfra. Total cost is approximately $170.
  • Training: All the Autoresearch runs are performed on a single NVIDIA H100 GPU from JarvisLabsAI. The total cost of training DotLM-165M is approximately $150.
Data CostCosts of Data Creation
Training CostCosts of Training of all Autoresearch runs

Costs of Data Creation

1 / 2

References

Tools

Citation

Cited as:

Shanmukha Sainath. "DotLM-165M: How I trained a 165M parameter language model from scratch". TensorWrites (Apr 2026). https://www.tensorwrites.com/posts/dotlm

BibTeX:
@article{dotlm2026,
  title   = "DotLM-165M: How I trained a 165M parameter language model from scratch",
  author  = "Shanmukha Sainath",
  journal = "TensorWrites",
  year    = "2026",
  month   = "Apr",
  url     = "https://www.tensorwrites.com/posts/dotlm"
}