<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en"><generator uri="https://jekyllrb.com/" version="4.4.1">Jekyll</generator><link href="https://alexchen.ai/feed.xml" rel="self" type="application/atom+xml"/><link href="https://alexchen.ai/" rel="alternate" type="text/html" hreflang="en"/><updated>2026-10-03T17:34:35+00:00</updated><id>https://alexchen.ai/feed.xml</id><title type="html">blank</title><subtitle>Personal site of Alex (Wei) Chen — Director at Qualcomm, founder of Nexa AI (acquired by Qualcomm), Stanford PhD. Writing on AI systems, generative AI inference, hardware-software co-design, and robotics.</subtitle><entry><title type="html">The AI Coding Paradox: When Faster Code Makes Projects Slower</title><link href="https://alexchen.ai/blog/2026/the-ai-coding-paradox/" rel="alternate" type="text/html" title="The AI Coding Paradox: When Faster Code Makes Projects Slower"/><published>2026-09-20T00:00:00+00:00</published><updated>2026-09-20T00:00:00+00:00</updated><id>https://alexchen.ai/blog/2026/the-ai-coding-paradox</id><content type="html" xml:base="https://alexchen.ai/blog/2026/the-ai-coding-paradox/"><![CDATA[ <blockquote class="block-tip"> <h5 class="no_toc" id="tldr">TL;DR</h5> <ul> <li>AI writes code fast, but teams ship <strong>working products</strong>, not code. Confusing coding velocity with delivery velocity is the first trap.</li> <li>AI helps most when <strong>implementation</strong> is the bottleneck. When the constraint is requirements, architecture, validation, or ownership, faster generation just grows the review queue.</li> <li>The data is sobering: DORA 2024 tied higher AI adoption to lower throughput and a 7.2% drop in delivery stability; METR found experienced developers were <strong>19% slower</strong> with AI while believing they were faster.</li> <li>“AI slop” is code produced faster than a team can understand it. Generation scales; understanding does not.</li> <li>The fix is not less AI but <strong>more disciplined AI</strong>: specify first, keep changes small, assign owners, protect the core, and measure outcomes instead of output.</li> </ul> </blockquote> <p>AI can write code remarkably fast. But software teams do not ship code; they ship working products. Confusing coding velocity with delivery velocity is the first trap of AI-assisted development.</p> <h2 id="where-the-bottleneck-actually-is">Where the Bottleneck Actually Is</h2> <p>AI creates clear gains when implementation is the bottleneck: generating boilerplate, test scaffolding, migrations, documentation, or well-specified features. But many projects are constrained elsewhere — uncertain product requirements, architectural decisions, customer feedback, hardware testing, regulatory validation, or cross-functional alignment. Generating code faster in these situations simply produces more work waiting for review.</p> <table> <thead> <tr> <th>Constraint</th> <th>Example</th> <th>Does faster codegen help?</th> </tr> </thead> <tbody> <tr> <td>Implementation</td> <td>Boilerplate, CRUD, migrations, test scaffolding</td> <td>Yes — large leverage</td> </tr> <tr> <td>Specification</td> <td>Unclear product requirements, shifting scope</td> <td>No — builds the wrong thing faster</td> </tr> <tr> <td>Architecture</td> <td>Choosing boundaries, ownership of state, data model</td> <td>No — locks in guesses</td> </tr> <tr> <td>External validation</td> <td>Hardware testing, regulatory sign-off, customer feedback</td> <td>No — queue grows</td> </tr> <tr> <td>Coordination</td> <td>Design reviews, cross-team alignment</td> <td>No — more to coordinate</td> </tr> </tbody> </table> <p>DoorDash CEO Tony Xu recently made this distinction: even when AI writes much of the code, engineers still spend substantial time on product reviews, design discussions, and coordination. Likewise, Google’s <a href="https://cloud.google.com/blog/products/devops-sre/announcing-the-2024-dora-report">2024 DORA research</a> found that greater AI adoption improved documentation and code-review speed, yet was associated with slightly lower delivery throughput and a 7.2% reduction in delivery stability. Local acceleration did not automatically improve the whole system.</p> <p>The effect can be even more surprising in mature codebases. In a randomized study of 16 experienced open-source developers completing 246 real issues, <a href="https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/">METR found</a> that AI tools made developers 19% slower. The developers expected a 24% speedup and still believed afterward that AI had accelerated them. Reviewing and correcting plausible-but-imperfect output consumed the apparent gain.</p> <blockquote class="block-warning"> <h5 class="no_toc" id="note">NOTE</h5> <p>The perception gap is the dangerous part. If a tool makes you slower while making you <em>feel</em> faster, no amount of self-reported productivity will catch it. Only measured outcomes will.</p> </blockquote> <h2 id="the-rise-of-ai-slop">The Rise of AI Slop</h2> <p>AI slop is not merely ugly code. It is code produced faster than a team can understand it: duplicated logic, unnecessary abstractions, inconsistent assumptions, shallow tests, and comments that describe syntax without explaining intent. Each change may look reasonable in isolation while the system gradually loses coherence.</p> <p>In the <a href="https://survey.stackoverflow.co/2025/ai">2025 Stack Overflow survey</a>, 66% of developers reported frustration with AI solutions that were “almost right,” while 45% said debugging AI-generated code took longer. Flask creator Armin Ronacher similarly described hearing from engineers who <a href="https://lucumr.pocoo.org/2026/2/13/the-final-bottleneck/">no longer know what is inside their own codebases</a>. Generation scales; understanding does not.</p> <p>The failure mode is a feedback loop, not a single bad commit:</p> <pre><code class="language-mermaid">flowchart TD
    A[Fast generation] --&gt; B[Review queue grows]
    B --&gt; C[Shallow review]
    C --&gt; D[Shared mental model erodes]
    D --&gt; E[Fixes break hidden assumptions]
    E --&gt; F[More generation to patch]
    F --&gt; B
</code></pre> <h3 id="a-subscription-system-six-months-later">A Subscription System, Six Months Later</h3> <p>Imagine a team using agents to build a subscription system. Within days, it has billing, retries, discounts, caching, and entitlement checks. Months later, an unusual renewal failure appears. The business rules are scattered across generated handlers and background jobs, and nobody knows which component owns the true state. Each attempted fix breaks another assumption. Rebuilding the workflow around an explicit state machine may become safer than continuing to patch it.</p> <p>The real failure was not one bad function; it was the loss of a shared mental model.</p> <h2 id="a-field-guide-for-disciplined-ai">A Field Guide for Disciplined AI</h2> <p>The answer is not less AI, but more disciplined AI.</p> <table> <thead> <tr> <th>Practice</th> <th>What it protects</th> </tr> </thead> <tbody> <tr> <td>Specify behavior and acceptance tests <strong>before</strong> generating code</td> <td>Building the right thing</td> </tr> <tr> <td>Keep AI-authored changes small and reviewable</td> <td>Review depth; the ability to say no</td> </tr> <tr> <td>Assign an owner who can explain every production change</td> <td>The shared mental model</td> </tr> <tr> <td>Protect core architecture, invariants, and performance-critical paths</td> <td>Coherence where it matters most</td> </tr> <tr> <td>Measure cycle time, failures, and rework — not lines of code or PR volume</td> <td>Honest signal about whether it is working</td> </tr> </tbody> </table> <p>Each of these is cheap compared to the alternative. A spec takes an hour; an unowned billing state machine takes a quarter to unwind.</p> <h2 id="closing">Closing</h2> <p>AI amplifies the system around it. When coding is the constraint, it can create enormous leverage. When clarity, validation, or ownership is the constraint, faster generation merely creates a larger queue — and eventually, a more expensive rewrite.</p> <p>The question to ask before reaching for an agent is not “can it write this?” but “is writing this the slow part?”</p>]]></content><author><name>Alex Chen</name></author><category term="engineering"/><category term="ai-coding"/><category term="software-engineering"/><category term="productivity"/><category term="agents"/><summary type="html"><![CDATA[Why AI-generated code accelerates some teams and slows others — coding velocity vs. delivery velocity, the rise of AI slop, and a short field guide for using AI without losing the plot.]]></summary></entry><entry><title type="html">Understanding Kernel Design from a Mathematical Perspective</title><link href="https://alexchen.ai/blog/2026/kernel-math/" rel="alternate" type="text/html" title="Understanding Kernel Design from a Mathematical Perspective"/><published>2026-05-11T00:00:00+00:00</published><updated>2026-05-11T00:00:00+00:00</updated><id>https://alexchen.ai/blog/2026/kernel-math</id><content type="html" xml:base="https://alexchen.ai/blog/2026/kernel-math/"><![CDATA[ <blockquote class="block-tip"> <h5 class="no_toc" id="tldr">TL;DR</h5> <ul> <li>Two mathematical tools explain most GPU kernel design decisions for ML inference: <strong>index notation</strong> (free indices are distributed across threads; dummy indices need intermediate storage) and <strong>associative mergeable summaries</strong> (a summary type with an associative merge $\oplus$ and a lift function).</li> <li><strong>Kernel fusion</strong> removes HBM round-trips: computing $O_i = (X_i + Y_i)\,\Phi(X_i + Y_i)$ in one kernel keeps the intermediate in registers because $i$ is a free index with no reduction.</li> <li><strong>FlashAttention</strong> is a mergeable summary over the KV length: per query row, keep the triple $(m, \ell, u)$ — running max, rescaled denominator, rescaled numerator — and merge partial tiles with $\ell_C = e^{m_A - m_C}\ell_A + e^{m_B - m_C}\ell_B$ (and likewise for $u$).</li> <li>The only reduction primitives needed are <strong>max</strong> and <strong>sum</strong>, both associative, so thread blocks can process tiles in any order and merge incrementally.</li> <li>General recipe: start from the final expression, split it into two disjoint parts, and ask which variables are needed to merge them — those variables become the summary state and the merge becomes $\oplus$.</li> </ul> </blockquote> <p>In this post, I discuss writing kernels for ML model inference, relating some mathematical concepts to kernel design and implementation.</p> <h2 id="mathematical-background">Mathematical Background</h2> <h3 id="index-notation">Index Notation</h3> <p>We use index notation to represent the parallelism of operators for kernels. Index notation includes the <strong>dummy index</strong> and the <strong>free index</strong>. The dummy index means we need some intermediate result to store, while the free index needs to be distributed to threads for computation.</p> <p>For example, a matmul in PyTorch can be represented as:</p> \[C_{ij} = \sum_{k} A_{ik} B_{kj}\] <p>where $i, j$ are the free indices and $k$ is the dummy index.</p> <h3 id="associative-mergeable-summary">Associative Mergeable Summary</h3> <p>An associative mergeable summary over a set $S$ consists of:</p> <ul> <li>A <strong>summary type</strong> $T$ with a merge operation $\oplus: T \times T \to T$ that is associative: $(a \oplus b) \oplus c = a \oplus (b \oplus c)$</li> <li>A <strong>lift</strong> function $f: S \to T$ that maps each element to its summary</li> <li>An <strong>identity</strong> element $e \in T$ such that $a \oplus e = e \oplus a = a$</li> </ul> <p>This generalizes the commutative monoid by not requiring commutativity — the merge order must be preserved, but partial summaries can still be computed independently and merged. For example, online softmax uses this structure: each partial summary carries a local max, a sum of exponentials, and an accumulator, and these can be merged associatively across thread blocks.</p> <h2 id="case-study">Case Study</h2> <h3 id="fused-vector-add-and-gelu">Fused Vector Add and GELU</h3> <p>Writing and reading from HBM is expensive, so a natural strategy is to fuse the vector add and GELU into a single kernel.</p> <p><strong>Unfused (two separate kernels):</strong></p> <p>First, compute the vector add and write the intermediate result $Z$ to HBM:</p> \[Z_i = X_i + Y_i\] <p>Then, read $Z$ back from HBM and apply GELU:</p> \[O_i = Z_i \cdot \Phi(Z_i)\] <p>where $\Phi$ is the standard Gaussian CDF. This requires one write and one read of the intermediate $Z$ to/from HBM.</p> <p><strong>Fused (single kernel):</strong></p> \[O_i = (X_i + Y_i) \cdot \Phi(X_i + Y_i)\] <p>The intermediate result $X_i + Y_i$ stays in registers — no HBM round-trip is needed. Since $i$ is a free index with no dummy indices, every element is independent and trivially parallelizable across threads.</p> <h3 id="flashattention">FlashAttention</h3> <p>The typical attention computation is:</p> \[O = \text{softmax}\!\left(\frac{QK^T}{\sqrt{d}}\right)V\] <p>In index notation:</p> \[O_{ij} = \text{softmax}\!\left(\frac{Q_{ik}K_{km}}{\sqrt{d}}\right)V_{mj}\] <p>where $Q\in\mathbb{R}^{S\times d}$, $K\in\mathbb{R}^{L\times d}$, $V\in\mathbb{R}^{L\times d_v}$, $d$ is the key/value dimension, $S$ is the sequence length, and $L$ is the KV cache length. Here $i, j$ are the free indices and $k, m$ are the dummy indices. We should write this as one kernel — otherwise we must materialize $QK^T$ to HBM, which is expensive when the sequence or KV cache length is long.</p> <p>The daunting part is softmax, which breaks naive parallelism over the dummy index $m$. But softmax is row-wise independent, so each query row can be handled separately. The remaining problem: for a single query row, how do we split the reduction over a long KV cache length $L$ across thread blocks?</p> <h4 id="step-1-what-do-we-actually-need">Step 1: What Do We Actually Need?</h4> <p>For a single query row $q$, the attention output is:</p> \[O = \sum_{j=1}^{L} \frac{\exp(s_j)}{\sum_{l=1}^{L} \exp(s_l)}\, v_j\] <p>where $s_j = \frac{q \cdot k_j}{\sqrt{d}}$. For numerical stability, we subtract the max $m = \max_j s_j$:</p> \[O = \frac{\sum_{j=1}^{L} \exp(s_j - m)\, v_j}{\sum_{j=1}^{L} \exp(s_j - m)}\] <p>So to compute the output, we need three things: the max $m$, the denominator $\ell = \sum_j \exp(s_j - m)$, and the numerator $u = \sum_j \exp(s_j - m)\, v_j$.</p> <h4 id="step-2-split-into-two-disjoint-subsets">Step 2: Split into Two Disjoint Subsets</h4> <p>Take two disjoint subsets $A, B$ with $A \cup B = {1, \dots, L}$. For each subset, we independently compute its own local version of these three quantities:</p> \[m_A = \max_{j \in A} s_j, \quad \ell_A = \sum_{j \in A} \exp(s_j - m_A), \quad u_A = \sum_{j \in A} \exp(s_j - m_A)\, v_j\] \[m_B = \max_{j \in B} s_j, \quad \ell_B = \sum_{j \in B} \exp(s_j - m_B), \quad u_B = \sum_{j \in B} \exp(s_j - m_B)\, v_j\] <p>Each subset uses its own local max for numerical stability. The question is: can we recover the global result from $(m_A, \ell_A, u_A)$ and $(m_B, \ell_B, u_B)$ alone?</p> <h4 id="step-3-merge">Step 3: Merge</h4> <p>Yes. Let $m_C = \max(m_A, m_B)$. We rescale everything to this common max:</p> \[\ell_C = e^{m_A - m_C}\, \ell_A + e^{m_B - m_C}\, \ell_B\] \[u_C = e^{m_A - m_C}\, u_A + e^{m_B - m_C}\, u_B\] \[O = \frac{u_C}{\ell_C}\] <p>This works because $e^{m_A - m_C}$ corrects each partial sum from its local max to the global max. Critically, the merge does not depend on the order of $A$ and $B$ — swapping them gives the same result. This is exactly what we need for GPU execution, where SMs process tiles in arbitrary order. By recursively applying this divide-and-conquer strategy, we can compute the attention output for the entire query matrix in parallel.</p> <h4 id="step-4-reduction-operators-across-threads">Step 4: Reduction Operators Across Threads</h4> <p>The merge formulas in Step 3 rely on two fundamental reduction operators: <strong>max</strong> and <strong>sum</strong>. Both are associative and commutative, which means they can be computed in any order across threads — exactly what we need for GPU parallelism.</p> <p><strong>Max reduction.</strong> Each thread block computes a local max $m_i$ over its tile of scores. To obtain the global max, we reduce across thread blocks:</p> \[m = \max(m_1, m_2, \dots, m_P)\] <p>where $P$ is the number of partitions. This is a tree reduction: pairs of threads compare their local maxes, the winners compare again, and so on in $\lceil \log_2 P \rceil$ steps. In practice this happens first within a warp via shuffle instructions, then across warps via shared memory.</p> <p><strong>Sum reduction.</strong> Once the global max $m$ is known (or as we merge incrementally), each thread block’s partial denominator must be rescaled and summed:</p> \[\ell = \sum_{i=1}^{P} e^{m_i - m}\, \ell_i\] <p>This is again a standard sum reduction — each thread contributes $e^{m_i - m}\, \ell_i$, and the partial sums are accumulated with the same tree pattern. The numerator $u$ follows identically:</p> \[u = \sum_{i=1}^{P} e^{m_i - m}\, u_i\] <p>These are exactly the merge formulas from Step 3, applied across $P$ thread blocks instead of two subsets. The key insight is that max and sum are the only two reduction primitives needed — max gives us the global shift for numerical stability, and sum (with the exponential correction factor) lets us accumulate both the denominator and the weighted value numerator. Because both reductions are associative, thread blocks can finish their local work in any order and merge results incrementally as they become available.</p> <h4 id="identifying-the-auxiliary-variables">Identifying the Auxiliary Variables</h4> <p>The critical auxiliary state is the triple $(m, \ell, u)$ — the running max, the rescaled denominator, and the rescaled numerator. These are exactly the variables FlashAttention maintains per query row. Each thread block processes a tile of keys, produces a partial $(m, \ell, u)$, and the results are merged incrementally with $\oplus$.</p> <p>This way of thinking generalizes: whenever you face a reduction that isn’t a simple sum or max, start from the final expression you need, try splitting it into two disjoint parts, and ask what variables are needed to merge them. The variables you discover become your summary state, and the merge formula becomes your $\oplus$. For softmax, the “recipe” turned out to be $(m, \ell, u)$ — the max is the non-obvious ingredient that makes everything else rescalable.</p>]]></content><author><name>Alex Chen</name></author><category term="ml"/><category term="kernel"/><category term="inference"/><category term="optimization"/><summary type="html"><![CDATA[Relating mathematical concepts — index notation, associative mergeable summaries — to GPU kernel design for ML inference, with case studies on fused operators and FlashAttention.]]></summary></entry><entry><title type="html">Gated Delta Networks: Improving Mamba2 with the Delta Rule</title><link href="https://alexchen.ai/blog/2026/gated-delta-networks/" rel="alternate" type="text/html" title="Gated Delta Networks: Improving Mamba2 with the Delta Rule"/><published>2026-04-13T00:00:00+00:00</published><updated>2026-04-13T00:00:00+00:00</updated><id>https://alexchen.ai/blog/2026/gated-delta-networks</id><content type="html" xml:base="https://alexchen.ai/blog/2026/gated-delta-networks/"><![CDATA[<p><em>A short reading note on <a href="https://arxiv.org/abs/2412.06464">Yang et al., 2024</a></em></p> <blockquote class="block-tip"> <h5 class="no_toc" id="tldr">TL;DR</h5> <ul> <li><strong>Linear attention</strong> compresses past keys and values into a fixed-size state $S_t = \sum_i v_i k_i^T$, giving $O(1)$ memory per decoding step — but with no way to forget, so retrieval degrades once stored pairs exceed the state dimension.</li> <li><strong>Mamba2</strong> adds a scalar decay gate $\alpha_t$ for bulk forgetting, but it decays every association equally. <strong>DeltaNet</strong> uses the delta rule to surgically overwrite the value for one key, but can only update one pair per step.</li> <li><strong>Gated DeltaNet</strong> (Yang et al., 2024) combines both: $S_t = S_{t-1}\,\alpha_t (I - \beta_t k_t k_t^T) + \beta_t v_t k_t^T$ — targeted updates via $\beta_t$ plus wholesale erasure via $\alpha_t$. It is equivalent to online SGD on $\tfrac{1}{2}|S_t k_t - v_t|^2$ with adaptive weight decay.</li> <li>Parallel training uses <strong>chunkwise WY representation</strong> of the Householder-like transitions; the only sequential piece is inverting a small $C \times C$ lower-triangular matrix per chunk, so the algorithm stays matmul-heavy and $O(L)$.</li> <li>Hybrids that interleave Gated DeltaNet with sliding-window attention (H1) or Mamba2 + SWA (H2) give the best overall benchmark results at competitive throughput.</li> </ul> </blockquote> <h2 id="from-softmax-attention-to-linear-attention">From Softmax Attention to Linear Attention</h2> <p>Standard Transformer attention computes:</p> \[O = \text{softmax}\!\left(\frac{QK^T}{\sqrt{d}}\right)V\] <p>where $Q \in \mathbb{R}^{L \times d_k}$, $K \in \mathbb{R}^{L \times d_k}$, $V \in \mathbb{R}^{L \times d_v}$. This scales quadratically with sequence length $L$.</p> <p>Linear attention drops the softmax and computes $O = (QK^T)V$. By changing the association order to $Q(K^TV)$, we can derive a recurrent form. Expanding per-token:</p> \[o_t = \sum_{i=1}^{t} v_i (k_i^T q_t) = \left(\sum_{i=1}^{t} v_i k_i^T\right) q_t\] <p>Defining the <strong>state matrix</strong> $S_t = \sum_{i=1}^{t} v_i k_i^T \in \mathbb{R}^{d_v \times d_k}$, we get a linear recurrence:</p> \[S_t = S_{t-1} + v_t k_t^T, \quad o_t = S_t q_t\] <p>This is the key insight: instead of storing all past $K, V$ tokens, we compress them into a fixed-size state $S_t$. Inference becomes $O(1)$ per step in memory, regardless of sequence length.</p> <h2 id="mamba2-adding-a-decay-gate">Mamba2: Adding a Decay Gate</h2> <p>Vanilla linear attention has no forgetting — $S_t$ accumulates everything, and once the number of stored key-value pairs exceeds $d_k$ (the state dimension), memory collisions become inevitable and retrieval degrades.</p> <p>Mamba2 addresses this by introducing a scalar decay gate $\alpha_t \in (0, 1)$:</p> \[S_t = \alpha_t S_{t-1} + v_t k_t^T, \quad o_t = S_t q_t\] <p>This uniformly decays all stored associations at each step, allowing the model to forget old information. However, the decay is <strong>indiscriminate</strong> — if the model needs to forget one specific key-value pair, <em>all</em> pairs get equally decayed. This leads to poor long-term retention when the needle is buried deep in a long context.</p> <h2 id="deltanet-the-delta-update-rule">DeltaNet: The Delta Update Rule</h2> <p>DeltaNet takes a different approach. Instead of decaying everything, it <strong>selectively replaces</strong> the value associated with the current key. The delta rule (Widrow et al., 1960) first erases the old value for key $k_t$, then writes a blended new value:</p> \[S_t = S_{t-1} - \underbrace{(S_{t-1} k_t)}_{\text{old value } v_t^{\text{old}}} k_t^T + \underbrace{(\beta_t v_t + (1-\beta_t) S_{t-1} k_t)}_{\text{new value } v_t^{\text{new}}} k_t^T\] <p>where $\beta_t \in (0, 1)$ is the writing strength. Simplifying:</p> \[S_t = S_{t-1}(I - \beta_t k_t k_t^T) + \beta_t v_t k_t^T\] <p>This is much more surgical — it only modifies the association for $k_t$ while leaving others intact. DeltaNet excels at associative recall benchmarks because of this precision. But it has a blind spot: it can only update <strong>one</strong> key-value pair at a time, so it lacks the ability to rapidly clear out large amounts of irrelevant context (e.g., during a topic switch).</p> <h2 id="gated-deltanet-best-of-both-worlds">Gated DeltaNet: Best of Both Worlds</h2> <p>The key observation from the paper is that gating and the delta rule are <strong>complementary</strong>:</p> <ul> <li><strong>Gating</strong> ($\alpha_t$): enables rapid, wholesale memory erasure (set $\alpha_t \to 0$ to flush the state)</li> <li><strong>Delta rule</strong> ($\beta_t$): enables precise, targeted updates (modify one association without affecting others)</li> </ul> <p>Gated DeltaNet combines both into a single update rule:</p> \[S_t = S_{t-1}\bigl(\alpha_t(I - \beta_t k_t k_t^T)\bigr) + \beta_t v_t k_t^T\] <p>When $\alpha_t \to 1$, this reduces to the pure delta rule. When $\alpha_t \to 0$, the state is flushed regardless. This unified mechanism gives the model flexible memory control — it can selectively update specific associations <em>and</em> perform bulk forgetting when needed.</p> <p>From an online learning perspective, this is equivalent to SGD on the regression loss $\mathcal{L}(S_t) = \frac{1}{2}|S_t k_t - v_t|^2$ with an <strong>adaptive weight decay</strong> $\alpha_t$ — a technique well-known in deep learning optimization.</p> <h2 id="the-parallel-training-challenge-where-the-matrix-inverse-appears">The Parallel Training Challenge: Where the Matrix Inverse Appears</h2> <p>The recurrent form is great for inference but sequential — each $S_t$ depends on $S_{t-1}$. For efficient GPU training, we need a parallel algorithm.</p> <p>The trick is <strong>chunkwise parallelism</strong>: split the sequence into chunks of size $C$, and within each chunk, partially expand the recurrence. For DeltaNet, the transition involves products of generalized Householder matrices $(I - \beta_t k_t k_t^T)$. Using the classical <strong>WY representation</strong>, these products can be expressed via an auxiliary matrix $T$:</p> \[T = \left[I + \text{strictLower}\!\left(\text{diag}(\boldsymbol{\beta}) \cdot (\Gamma \odot KK^T)\right)\right]^{-1} \text{diag}(\boldsymbol{\beta})\] <p>where $\Gamma$ is the decay-aware mask with entries $\Gamma_{ij} = \gamma_i / \gamma_j$ (cumulative products of $\alpha$). This is where the <strong>matrix inverse</strong> appears — the lower-triangular system encodes the sequential dependencies within a chunk. The matrix is $C \times C$ (chunk size, typically small), so the inverse is tractable and the overall algorithm remains rich in matrix multiplications suitable for tensor cores.</p> <p>With $T$ computed, the transformed values $\tilde{U} = TV$ and keys $W = TK$ allow the chunk-level state update and output to be expressed as matmuls:</p> \[S_{[t+1]} = \vec{S}_{[t]} + \left(\tilde{U}_{[t]} - \overleftarrow{W}_{[t]} S_{[t]}^T\right)^T \vec{K}_{[t]}\] \[O_{[t]} = \overleftarrow{Q}_{[t]} S_{[t]}^T + \left(Q_{[t]} K_{[t]}^T \odot M\right)\left(\tilde{U}_{[t]} - \overleftarrow{W}_{[t]} S_{[t]}^T\right)\] <p>where the arrows denote decay-scaled versions of the variables (decayed to chunk boundaries). This preserves the $O(L)$ complexity while enabling hardware-efficient parallel training.</p> <h2 id="architecture-and-hybrids">Architecture and Hybrids</h2> <p>The full Gated DeltaNet block follows the Llama macro architecture (token mixer + SwiGLU MLP). For the token mixer: $q, k, v$ are produced via linear projection → short convolution → SiLU, with <strong>L2 normalization</strong> on $q, k$ for training stability. The scalars $\alpha, \beta$ use linear projections, and the output passes through normalization and gating before the final projection.</p> <p>The paper also proposes hybrid variants that interleave Gated DeltaNet with other layers:</p> <ul> <li><strong>Gated DeltaNet-H1</strong>: Gated DeltaNet + sliding window attention (SWA)</li> <li><strong>Gated DeltaNet-H2</strong>: Mamba2 + Gated DeltaNet + SWA</li> </ul> <p>The hybrid models combine the recurrent layers’ efficient long-range modeling with attention’s strength at local patterns and precise retrieval, achieving the best results across benchmarks while maintaining competitive training throughput.</p>]]></content><author><name>Alex Chen</name></author><category term="ml"/><category term="transformers"/><category term="mamba"/><category term="linear-attention"/><summary type="html"><![CDATA[A reading note on Yang et al., 2024 — how combining Mamba2's decay gate with DeltaNet's selective memory update yields a flexible recurrent model, and why its parallel training algorithm requires a matrix inverse.]]></summary></entry><entry><title type="html">Universal Manipulation Interface: In-The-Wild Robot Teaching Without In-The-Wild Robots</title><link href="https://alexchen.ai/blog/2026/universal-manipulation-interface/" rel="alternate" type="text/html" title="Universal Manipulation Interface: In-The-Wild Robot Teaching Without In-The-Wild Robots"/><published>2026-04-05T00:00:00+00:00</published><updated>2026-04-05T00:00:00+00:00</updated><id>https://alexchen.ai/blog/2026/universal-manipulation-interface</id><content type="html" xml:base="https://alexchen.ai/blog/2026/universal-manipulation-interface/"><![CDATA[<p><em>A short reading note on “Universal Manipulation Interface: In-The-Wild Robot Teaching Without In-The-Wild Robots” (Chi et al., 2024)</em></p> <blockquote class="block-tip"> <h5 class="no_toc" id="tldr">TL;DR</h5> <ul> <li><strong>UMI</strong> (Chi et al., 2024) is a handheld, 3D-printed parallel gripper with a wrist-mounted GoPro fisheye camera and two side mirrors. A human holds it to collect demonstrations anywhere — <strong>no robot is needed during data collection</strong>.</li> <li>The mirrors put two laterally offset views into one frame, so the policy recovers <strong>depth from stereo disparity</strong> without extra sensors.</li> <li>Actions are <strong>relative SE(3) end-effector displacements</strong>, so the arm’s kinematics never appear in observations or actions. Policies are hardware-agnostic, but a model-based IK controller on the target robot must track the Cartesian targets.</li> <li>Two retargeting steps make data deployable: <strong>latency alignment</strong> (shift action labels by the robot’s delay $\tau$) and <strong>kinematic filtering</strong> (drop demonstrations the target arm cannot execute).</li> <li>With Diffusion Policy as the backbone, UMI transfers in-the-wild demonstrations (cup arrangement, dynamic tossing, bimanual cloth folding) to a real arm with no robot-specific data.</li> </ul> </blockquote> <p>Collecting robot demonstrations at scale requires either expensive teleoperation rigs or time on real hardware. UMI sidesteps both constraints with a <strong>handheld data collection interface</strong>: a carefully designed gripper that a human operator holds directly. Demonstrations captured this way are hardware-agnostic and can be collected anywhere, then deployed onto a target robot through a structured retargeting pipeline.</p> <h2 id="hardware-design-gripper-gopro-and-mirrors">Hardware Design: Gripper, GoPro, and Mirrors</h2> <p>The device consists of a 3D-printed parallel gripper instrumented with:</p> <ul> <li>A <strong>GoPro fisheye camera</strong> mounted at the wrist, providing a wide field of view that captures the workspace and the object being manipulated.</li> <li><strong>Two flat mirrors angled on each side</strong> of the gripper, reflecting the scene from slightly offset viewpoints into the same camera frame.</li> </ul> <p>The mirrors are the most elegant detail. Because a single fisheye frame now encodes two laterally shifted views of the scene — each occupying a different region of the image — the policy can implicitly recover <strong>depth from stereo disparity</strong> without any additional sensors or calibration hardware. The GoPro records everything in a single stream; the stereo structure is baked into the optics.</p> <h2 id="relative-action-representation">Relative Action Representation</h2> <p>The critical design choice in UMI is what gets recorded. Rather than logging absolute end-effector poses in a world frame, UMI records <strong>relative end-effector displacements</strong> — the change in 6-DoF gripper pose between consecutive timesteps:</p> \[a_t = \Delta T_t = T_{t}^{-1} T_{t+1} \in SE(3)\] <p>This has an important consequence: <strong>the robot arm’s configuration is entirely absent from both observations and actions.</strong> The policy sees wrist-camera images and outputs relative gripper motions — nothing about the joint angles, link lengths, or kinematics of whatever arm is holding the gripper.</p> <p>The flip side is that UMI cannot be fully end-to-end. Because the policy emits Cartesian relative targets rather than joint commands, a <strong>model-based controller</strong> on the target robot must solve inverse kinematics and track those targets. The neural network handles perception and high-level motion; the arm’s own controller handles the rest.</p> <h2 id="retargeting-to-a-target-robot">Retargeting to a Target Robot</h2> <p>Raw demonstrations are not directly deployable. Two preprocessing steps adapt the data to a specific robot:</p> <p><strong>1. Latency alignment.</strong> Every real robot has a response delay between receiving a command and executing it. If the recorded action sequence is played back without compensation, the effective timing is shifted and the policy sees stale observations. UMI estimates the target robot’s end-to-end latency $\tau$ and shifts the action labels accordingly:</p> \[a_t^{\text{aligned}} = a_{t + \tau}\] <p>This ensures that the observation-action pairs in training reflect what the robot will actually experience at inference time.</p> <p><strong>2. Kinematic filtering.</strong> Not every human demonstration is reachable by the target robot. Joint limits, workspace boundaries, and velocity/acceleration caps all constrain what motions are physically executable. Demonstrations that require the target robot to move through configurations outside its feasible set are discarded. This step is necessarily robot-specific: the same human motion may be valid for one arm and infeasible for another.</p> <h2 id="policy-and-results">Policy and Results</h2> <p>UMI uses <strong>Diffusion Policy</strong> as the learning backbone, conditioned on the wrist-camera image stream. Across tasks like cup arrangement, dynamic tossing, and bimanual cloth folding — all with in-the-wild demonstrations collected in homes and outdoor environments — UMI achieves strong transfer to a stationary robot arm with no robot-specific data collection. The interface enables non-expert users to collect usable demonstrations in minutes.</p>]]></content><author><name>Alex Chen</name></author><category term="ml"/><category term="robotics"/><category term="imitation-learning"/><category term="manipulation"/><summary type="html"><![CDATA[A reading note on Chi et al., 2024 — how a handheld gripper with a fisheye camera and mirrors enables hardware-agnostic demonstration collection anywhere, and what retargeting steps bridge the gap to a real robot.]]></summary></entry><entry><title type="html">Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware</title><link href="https://alexchen.ai/blog/2026/act-action-chunking-with-transformers/" rel="alternate" type="text/html" title="Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware"/><published>2026-04-04T00:00:00+00:00</published><updated>2026-04-04T00:00:00+00:00</updated><id>https://alexchen.ai/blog/2026/act-action-chunking-with-transformers</id><content type="html" xml:base="https://alexchen.ai/blog/2026/act-action-chunking-with-transformers/"><![CDATA[<p><em>A short reading note on ALOHA (Zhao et al., 2023)</em></p> <blockquote class="block-tip"> <h5 class="no_toc" id="tldr">TL;DR</h5> <ul> <li><strong>ACT</strong> (Action Chunking with Transformers), the learning algorithm behind ALOHA (Zhao et al., 2023), fixes compounding errors in fine bimanual manipulation with two ideas: predict a <strong>chunk of $k$ future actions</strong> at once, and model demonstration multimodality with a <strong>Conditional VAE</strong>.</li> <li>The policy sees 4 RGB cameras plus 14 joint positions and outputs <strong>target joint positions</strong>; motor-level PID loops track them, which makes ACT “half end-to-end”.</li> <li>The CVAE encoder (BERT-style, <code class="language-plaintext highlighter-rouge">[CLS]</code> token → $z \in \mathbb{R}^{32}$) is used only in training; at test time <strong>$z = 0$</strong> gives deterministic, average-style behavior.</li> <li>The decoder uses fixed sinusoidal queries to emit all $k$ actions <strong>in parallel</strong> (no autoregression), and <strong>temporal ensembling</strong> with exponential weights smooths overlapping chunks.</li> <li>Results: <strong>80–96% success</strong> on battery insertion and cup opening where BeT, RT-1 and BC-ConvMLP reach 0–8%. Removing chunking drops 44% → 1%; removing the CVAE drops 35% → 2% on human data.</li> </ul> </blockquote> <p>Fine manipulation tasks — slotting a battery, threading a zip tie — fail catastrophically with standard imitation learning because small errors compound over time. ALOHA’s learning algorithm, <strong>ACT</strong>, attacks this with two ideas: predict <em>chunks</em> of future actions instead of one step at a time, and model the multimodality of human demonstrations with a <strong>Conditional VAE</strong>.</p> <p><strong>Observation and action space.</strong> At each timestep the policy observes: 4 RGB camera images (top, front, two wrists) and the current joint positions of both arms (14 DoF total). The action it predicts is a sequence of <em>target joint positions</em> — not torques or currents.</p> <p>This makes ACT <strong>half end-to-end</strong>: the neural network maps pixels → target joint positions, but a low-level PID controller inside each Dynamixel motor handles the actual current/torque commands to track those targets. The learning side never touches the motor-level signal. This separation simplifies training considerably, but it also means the policy cannot reason about contact forces directly — it relies on the PID loop to absorb those dynamics.</p> <h2 id="understanding-the-cvae-through-the-vae-analogy">Understanding the CVAE Through the VAE Analogy</h2> <p>To understand why ACT uses a CVAE, start with the original VAE for images.</p> <p>A VAE learns to encode images into a latent distribution $q_\phi(z \mid x) \approx \mathcal{N}(\mu, \sigma^2)$, then decode samples back:</p> \[\mathcal{L} = \underbrace{\mathbb{E}[\log p_\theta(x \mid z)]}_{\text{reconstruction}} - \underbrace{\beta \, D_{\mathrm{KL}}\!\left(q_\phi(z \mid x) \,\|\, \mathcal{N}(0, I)\right)}_{\text{regularization}}\] <p>At generation time, you sample $z \sim \mathcal{N}(0, I)$ and decode — new images emerge from the learned distribution over all training samples.</p> <p><strong>Now replace “image samples” with “action trajectories.”</strong> Human demonstrations of the same task are inherently stochastic: the exact handover position mid-air differs every time. ACT treats each demonstration trajectory $a_{t:t+k}$ as a sample from some latent distribution of “styles.” The CVAE learns:</p> \[\mathcal{L}_{\mathrm{ACT}} = \underbrace{\mathrm{MSE}(\hat{a}_{t:t+k},\, a_{t:t+k})}_{\mathcal{L}_{\mathrm{reconst}}} + \beta \underbrace{D_{\mathrm{KL}}\!\left(q_\phi(z \mid a_{t:t+k}, \bar{o}_t) \,\|\, \mathcal{N}(0, I)\right)}_{\mathcal{L}_{\mathrm{reg}}}\] <p>where \(\bar{o}_t\) is the proprioceptive observation (joint positions, no images) and \(\hat{a}_{t:t+k}\) is the predicted action chunk of length \(k\).</p> <p>The critical difference from an image VAE: <strong>at test time, instead of sampling noise, ACT sets $z = 0$</strong> (the mean of the Gaussian prior). This keeps inference deterministic and stable — the robot doesn’t randomly explore during deployment.</p> <blockquote> <p><em>Why does this work? The CVAE objective forces the encoder to compress the “style” of a trajectory — the particular way a human chose to do the task — into $z$. At test time, $z = 0$ selects the mean style, which corresponds to the most likely, average-human behavior across all demonstrations.</em></p> </blockquote> <h2 id="architecture-how-z-is-extracted">Architecture: How z is Extracted</h2> <p>The CVAE <strong>encoder</strong> is a BERT-style transformer. Its input sequence is:</p> \[[\underbrace{\mathrm{[CLS]}}_{\text{learnable}},\; \underbrace{W_1 \bar{o}_t}_{\text{joints}},\; \underbrace{W_2 a_t,\, \ldots,\, W_2 a_{t+k}}_{\text{action sequence}}] \quad \in \mathbb{R}^{(k+2) \times 512}\] <p>After passing through 4 self-attention blocks, <strong>only the output at the <code class="language-plaintext highlighter-rouge">[CLS]</code> position is used</strong> to predict $\mu_z$ and $\sigma_z$. This is the same trick as BERT’s sentence-level embedding: the <code class="language-plaintext highlighter-rouge">[CLS]</code> token aggregates global context from the full sequence into a single vector, from which $z \in \mathbb{R}^{32}$ is sampled via reparameterization.</p> <p>The encoder is discarded entirely at test time.</p> <h2 id="architecture-how-actions-are-decoded-no-autoregression">Architecture: How Actions are Decoded (No Autoregression)</h2> <p>The CVAE <strong>decoder</strong> (the actual deployed policy) takes:</p> <ul> <li>4 RGB images → ResNet18 → flattened spatial features $\in \mathbb{R}^{1200 \times 512}$</li> <li>Joint positions $\bar{o}_t$ → linear projection → $\mathbb{R}^{512}$</li> <li>Style variable $z$ → linear projection → $\mathbb{R}^{512}$</li> </ul> <p>These are concatenated and processed by a <strong>transformer encoder</strong> (1202 tokens total). The encoder output serves as <em>keys</em> and <em>values</em> in a transformer decoder.</p> <p>The decoder’s <strong>queries are fixed sinusoidal position embeddings</strong> — one per future timestep:</p> \[\mathrm{query}_i = \mathrm{SinPE}(i), \quad i = 1, \ldots, k\] <p>This means the decoder attends to the full observation context (via cross-attention) and outputs all $k$ future actions <strong>in parallel</strong> — there is no autoregressive token generation, no teacher forcing, no causal masking. The output is directly $\hat{a}_{t:t+k} \in \mathbb{R}^{k \times 14}$.</p> <h2 id="action-chunking--temporal-ensembling">Action Chunking + Temporal Ensembling</h2> <p>ACT predicts $\pi_\theta(a_{t:t+k} \mid o_t)$ — a chunk of $k$ future joint positions. At each timestep, the policy is queried, producing overlapping predictions. These are merged with exponential weighting:</p> \[a_t = \frac{\sum_i w_i \cdot \hat{A}_t[i]}{\sum_i w_i}, \quad w_i = \exp(-m \cdot i)\] <p>where $\hat{A}_t[i]$ is the $i$-th oldest prediction for timestep $t$. Older predictions are down-weighted. This removes jerky switching between chunks and produces smooth robot motion without any extra training cost.</p> <h2 id="results">Results</h2> <p>ACT achieves <strong>80–96% success</strong> on tasks like battery insertion and condiment cup opening — tasks where all prior methods (BeT, RT-1, BC-ConvMLP) essentially fail (0–8%). Ablations confirm that each component matters: removing action chunking drops performance from 44% → 1%, removing CVAE training drops performance from 35% → 2% on human (stochastic) data.</p>]]></content><author><name>Alex Chen</name></author><category term="ml"/><category term="robotics"/><category term="transformers"/><category term="imitation-learning"/><category term="manipulation"/><summary type="html"><![CDATA[A reading note on ALOHA (Zhao et al., 2023) — how ACT uses action chunking and a Conditional VAE to solve fine manipulation tasks that cause standard imitation learning to fail completely.]]></summary></entry><entry><title type="html">Diffusion Policy: Visuomotor Policy Learning via Action Diffusion</title><link href="https://alexchen.ai/blog/2026/diffusion-policy/" rel="alternate" type="text/html" title="Diffusion Policy: Visuomotor Policy Learning via Action Diffusion"/><published>2026-04-04T00:00:00+00:00</published><updated>2026-04-04T00:00:00+00:00</updated><id>https://alexchen.ai/blog/2026/diffusion-policy</id><content type="html" xml:base="https://alexchen.ai/blog/2026/diffusion-policy/"><![CDATA[<p><em>A short reading note on Chi et al., 2023</em></p> <blockquote class="block-tip"> <h5 class="no_toc" id="tldr">TL;DR</h5> <ul> <li><strong>Diffusion Policy</strong> (Chi et al., 2023) treats robot action generation as DDPM denoising: start from Gaussian noise over an action sequence and iteratively denoise it, conditioned on the visual observation $O_t$.</li> <li>The observation is encoded <strong>once</strong> per step; only the small denoiser runs $K$ times. With DDIM, 100 training steps drop to <strong>10 inference steps (~0.1 s on an RTX 3080)</strong>, and a receding horizon (predict $T_p$, execute $T_a &lt; T_p$) amortizes the cost.</li> <li>It handles <strong>multimodal demonstrations</strong> because it learns the score function without a normalization constant — no averaging between two valid trajectories (explicit regression) and no negative sampling (energy-based IBC).</li> <li>Results: an average <strong>46.9% improvement</strong> over prior state of the art across 15 tasks, with <strong>95% success on Push-T</strong>.</li> </ul> </blockquote> <p>Just as DDPM generates images by iteratively denoising Gaussian noise, Diffusion Policy applies the same process to <strong>robot actions</strong>. The generated artifact — a sequence of joint positions or end-effector poses — is tiny compared to an image, so the model itself is lightweight. The trade-off: to produce a clean action chunk, the model must run the denoising loop $K$ times at inference, requiring higher inference throughput than a single forward-pass policy.</p> <h2 id="from-image-diffusion-to-action-diffusion">From Image Diffusion to Action Diffusion</h2> <p>In standard DDPM, the denoising update is:</p> \[x^{k-1} = \alpha\!\left(x^k - \gamma\,\varepsilon_\theta(x^k, k)\right) + \mathcal{N}(0, \sigma^2 I)\] <p>where $\varepsilon_\theta$ is a network that predicts the noise added at step $k$, and $\alpha, \gamma, \sigma$ follow a noise schedule. Starting from $x^K \sim \mathcal{N}(0, I)$, you run this $K$ times to get a clean sample $x^0$.</p> <p>Diffusion Policy makes two changes: (1) replace $x$ with the action sequence $A_t$, and (2) condition the denoiser on the visual observation $O_t$:</p> \[A_t^{k-1} = \alpha\!\left(A_t^k - \gamma\,\varepsilon_\theta(O_t,\, A_t^k,\, k)\right) + \mathcal{N}(0, \sigma^2 I)\] <p>The training loss becomes:</p> \[\mathcal{L} = \mathrm{MSE}\!\left(\varepsilon_k,\; \varepsilon_\theta(O_t,\; A_t^0 + \varepsilon_k,\; k)\right)\] <p>The observation $O_t$ (images + proprioception) is encoded <strong>once</strong> before the denoising loop — not re-processed at every iteration — which keeps inference tractable.</p> <h2 id="inference-cost">Inference Cost</h2> <p>Each inference call requires $K$ forward passes through $\varepsilon_\theta$. With 100 training iterations, DDIM (Song et al., 2021) allows dropping to 10 inference iterations without retraining, giving ~0.1s latency on a 3080 GPU. To amortize this cost, the policy predicts $T_p$ future steps and executes $T_a &lt; T_p$ of them before replanning — the same receding-horizon trick as ACT’s action chunking.</p> <h2 id="why-this-works-better-than-explicit-policies">Why This Works Better Than Explicit Policies</h2> <p>Explicit policies (direct regression) struggle with multimodal demonstrations — averaging over two valid trajectories produces a bad in-between trajectory. Diffusion Policy inherits the ability of diffusion models to represent <strong>arbitrary distributions</strong>, including multimodal ones, because the score function $\varepsilon_\theta \approx -\nabla_a \log p(a \mid o)$ does not require estimating any normalization constant. This also makes training significantly more stable than energy-based implicit policies (IBC), which need negative samples to approximate the same quantity.</p> <h2 id="results">Results</h2> <p>Across 15 tasks in simulation and the real world, Diffusion Policy improves over prior state-of-the-art by an average of <strong>46.9%</strong>, with near-human performance on tasks like Push-T (95% success) and sauce spreading.</p>]]></content><author><name>Alex Chen</name></author><category term="ml"/><category term="robotics"/><category term="diffusion"/><category term="imitation-learning"/><summary type="html"><![CDATA[A reading note on Chi et al., 2023 — how DDPM-style denoising is applied to robot action generation, why it handles multimodal demonstrations better than explicit policies, and what the inference cost looks like in practice.]]></summary></entry><entry><title type="html">Reinforcement Learning for Large Language Models</title><link href="https://alexchen.ai/blog/2026/reinforcement-learning-for-llms/" rel="alternate" type="text/html" title="Reinforcement Learning for Large Language Models"/><published>2026-02-14T00:00:00+00:00</published><updated>2026-02-14T00:00:00+00:00</updated><id>https://alexchen.ai/blog/2026/reinforcement-learning-for-llms</id><content type="html" xml:base="https://alexchen.ai/blog/2026/reinforcement-learning-for-llms/"><![CDATA[ <blockquote class="block-tip"> <h5 class="no_toc" id="tldr">TL;DR</h5> <ul> <li>RL is a loop of states, actions, and rewards; <strong>value-based</strong> methods (Q-learning, DQN) learn $Q^*(s,a)$ and act greedily, while <strong>policy-based</strong> methods (REINFORCE) optimize $\pi_\theta$ directly via the policy gradient theorem.</li> <li><strong>Actor-critic</strong> methods cut REINFORCE’s variance with a learned baseline and the advantage $A(s,a) = Q(s,a) - V(s)$; <strong>PPO</strong> adds a clipped probability-ratio objective so each update stays close to the old policy.</li> <li>In <strong>RLHF</strong>, the LLM is the policy, the context is the state, each token is an action, and a reward model trained on human preference pairs provides the reward; PPO maximizes reward minus a <strong>KL penalty</strong> to a frozen reference (SFT) model.</li> <li><strong>DPO</strong> removes the reward model: the KL-regularized optimum has a closed form, so preference pairs can be fit with a logistic loss on $\beta \log \frac{\pi_\theta}{\pi_\text{ref}}$ differences. It is simpler and more stable than PPO-based RLHF and dominant in open-source alignment pipelines.</li> </ul> </blockquote> <p>Reinforcement learning has become a central technique in the post-pretraining phase of LLM development. RLHF (Reinforcement Learning from Human Feedback) is how models like ChatGPT learned to be helpful, and DPO (Direct Preference Optimization) is its cleaner, more tractable successor. But to understand why these methods work — and where they can fail — it helps to build up from RL fundamentals.</p> <p>This post covers the core RL concepts, the value-based and policy-based methods that underpin them, and how PPO and DPO are adapted for LLM training.</p> <h2 id="background-the-rl-framework">Background: The RL Framework</h2> <p>The RL setup involves an <strong>agent</strong> interacting with an <strong>environment</strong> in a loop. At each timestep $t$:</p> <ol> <li>The agent observes the current state $S_t$</li> <li>It takes action $A_t$ according to its policy</li> <li>The environment transitions to state $S_{t+1}$ and emits reward $R_{t+1}$</li> </ol> <figure> <picture> <source class="responsive-img-srcset" srcset="/assets/img/RL_basic-480.webp 480w,/assets/img/RL_basic-800.webp 800w,/assets/img/RL_basic-1400.webp 1400w," type="image/webp" sizes="95vw"/> <img src="/assets/img/RL_basic.png" class="img-fluid rounded z-depth-1" width="100%" height="auto" title="The basic reinforcement learning loop" loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"/> </picture> </figure> <div class="caption">Figure 1: The agent-environment interaction loop in RL.</div> <p>The tuple $(S_t, A_t, S_{t+1}, R_{t+1})$ is the fundamental unit of experience in RL.</p> <p><strong>Key concepts:</strong></p> <ul> <li><strong>Markov property:</strong> The next state depends only on the current state and action, not the full history. This is the assumption that makes RL tractable.</li> <li><strong>Policy</strong> $\pi$: A function mapping states to actions (or distributions over actions). The goal is to find $\pi^*$ — the optimal policy.</li> <li><strong>Reward</strong> $R$: Immediate feedback from the environment after taking an action.</li> <li><strong>Value</strong> $V$: The discounted sum of expected future rewards from a given state — a long-horizon signal, unlike the immediate reward.</li> </ul> <blockquote> <h5 id="note">NOTE</h5> <p class="block-tip">Deep learning is not natively designed for RL. RL is fundamentally a mathematical framework, and neural networks are just one (very powerful) way to represent the functions involved. Always think from the math first; the network architecture follows.</p> </blockquote> <h2 id="value-based-methods">Value-Based Methods</h2> <figure> <picture> <source class="responsive-img-srcset" srcset="/assets/img/RL_classification-480.webp 480w,/assets/img/RL_classification-800.webp 800w,/assets/img/RL_classification-1400.webp 1400w," type="image/webp" sizes="95vw"/> <img src="/assets/img/RL_classification.png" class="img-fluid rounded z-depth-1" width="100%" height="auto" title="Taxonomy of RL methods" loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"/> </picture> </figure> <div class="caption">Figure 2: Taxonomy of RL methods — value-based, policy-based, and actor-critic.</div> <p>Value-based methods learn to estimate how good a given state (or state-action pair) is, then derive a policy implicitly by acting greedily with respect to that estimate.</p> <h3 id="the-value-function">The Value Function</h3> <p>The <strong>state-value function</strong> $V_\pi(s)$ gives the expected discounted return starting from state $s$ under policy $\pi$:</p> \[V_{\pi}(s) = \mathbb{E}_{\tau \sim \pi}\left[ R_{t+1} + \gamma R_{t+2} + \gamma^2 R_{t+3} + \cdots \mid S_t = s \right]\] <p>The <strong>action-value function</strong> $Q_\pi(s, a)$ adds one more level of granularity — it conditions on both state and action:</p> \[Q_{\pi}(s, a) = \mathbb{E}_{\pi}\left[ G_t \mid S_t = s, A_t = a \right]\] <p>The optimal policy then follows directly:</p> \[\pi^* = \operatorname*{arg\,max}_{a}\, Q^*(s, a)\] <h3 id="bellman-equation-and-q-learning">Bellman Equation and Q-Learning</h3> <p>Computing $V$ or $Q$ by simulating full episodes is expensive. The <strong>Bellman equation</strong> decomposes this into a one-step recursive form:</p> \[V_\pi(s) = \mathbb{E}_{\pi}\left[ R_{t+1} + \gamma \cdot V_{\pi}(S_{t+1}) \mid S_t = s \right]\] <p><strong>Q-learning</strong> is an off-policy method that uses temporal difference (TD) updates to iteratively improve $Q$:</p> \[Q(S_t, A_t) \leftarrow Q(S_t, A_t) + \alpha \left( R_{t+1} + \gamma \max_a Q(S_{t+1}, a) - Q(S_t, A_t) \right)\] <p>The $\epsilon$-greedy policy balances exploration and exploitation: take the greedy action with probability $\epsilon$, a random action otherwise.</p> <h3 id="deep-q-networks-dqn">Deep Q-Networks (DQN)</h3> <p>When the state space is continuous or too large for a table, we parameterize $Q$ with a neural network $Q_\theta(s, a)$. This is <strong>DQN</strong>. The network is trained to minimize:</p> \[\mathcal{L}(\theta) = \left( y_j - Q_\theta(\phi_j, a_j) \right)^2\] <p>where $y_j = r_j + \gamma \max_{a’} \hat{Q}(\phi_{j+1}, a’; \theta^-)$ is the TD target computed using a <strong>target network</strong> $\hat{Q}$ with lagged weights $\theta^-$. A <strong>replay buffer</strong> stores past transitions and provides decorrelated minibatches for stable training.</p> <p>In code, the TD target computation looks like:</p> <div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">with</span> <span class="n">torch</span><span class="p">.</span><span class="nf">no_grad</span><span class="p">():</span>
    <span class="n">target_max</span><span class="p">,</span> <span class="n">_</span> <span class="o">=</span> <span class="nf">target_network</span><span class="p">(</span><span class="n">data</span><span class="p">.</span><span class="n">next_observations</span><span class="p">).</span><span class="nf">max</span><span class="p">(</span><span class="n">dim</span><span class="o">=</span><span class="mi">1</span><span class="p">)</span>
    <span class="n">td_target</span> <span class="o">=</span> <span class="n">data</span><span class="p">.</span><span class="n">rewards</span><span class="p">.</span><span class="nf">flatten</span><span class="p">()</span> <span class="o">+</span> <span class="n">args</span><span class="p">.</span><span class="n">gamma</span> <span class="o">*</span> <span class="n">target_max</span> <span class="o">*</span> <span class="p">(</span><span class="mi">1</span> <span class="o">-</span> <span class="n">data</span><span class="p">.</span><span class="n">dones</span><span class="p">.</span><span class="nf">flatten</span><span class="p">())</span>

<span class="n">old_val</span> <span class="o">=</span> <span class="nf">q_network</span><span class="p">(</span><span class="n">data</span><span class="p">.</span><span class="n">observations</span><span class="p">).</span><span class="nf">gather</span><span class="p">(</span><span class="mi">1</span><span class="p">,</span> <span class="n">data</span><span class="p">.</span><span class="n">actions</span><span class="p">).</span><span class="nf">squeeze</span><span class="p">()</span>
<span class="n">loss</span> <span class="o">=</span> <span class="n">F</span><span class="p">.</span><span class="nf">mse_loss</span><span class="p">(</span><span class="n">td_target</span><span class="p">,</span> <span class="n">old_val</span><span class="p">)</span>
</code></pre></div></div> <h2 id="policy-based-methods">Policy-Based Methods</h2> <p>Rather than learning a value function and deriving a policy implicitly, policy-based methods directly parameterize and optimize $\pi_\theta$.</p> <p><strong>Advantages over value-based methods:</strong></p> <ul> <li>Naturally handles continuous and high-dimensional action spaces</li> <li>Can represent stochastic policies</li> <li>Better convergence properties in practice</li> </ul> <p><strong>Disadvantages:</strong></p> <ul> <li>Often converges to local optima</li> <li>High variance in the gradient estimate</li> </ul> <p>The policy objective is the expected total return:</p> \[J(\theta) = \mathbb{E}_{\tau \sim \pi_\theta}\left[ R(\tau) \right]\] <p>The <strong>policy gradient theorem</strong> gives us a tractable gradient:</p> \[\nabla_\theta J(\theta) = \mathbb{E}_{\pi_\theta}\left[ \nabla_\theta \log \pi_\theta(a_t \mid s_t) \cdot R(\tau) \right]\] <p>The derivation follows from the log-derivative trick:</p> \[\begin{aligned} \nabla_\theta J(\theta) &amp;= \nabla_\theta \sum_\tau P(\tau; \theta) R(\tau) \\ &amp;= \sum_\tau P(\tau; \theta) \frac{\nabla_\theta P(\tau; \theta)}{P(\tau; \theta)} R(\tau) \\ &amp;= \mathbb{E}_{\tau \sim \pi}\left[ \nabla_\theta \log P(\tau; \theta) \cdot R(\tau) \right] \\ &amp;= \mathbb{E}_{\pi_\theta}\left[ \sum_{t=0}^T \nabla_\theta \log \pi_\theta(a_t \mid s_t) \cdot R(\tau) \right] \end{aligned}\] <h3 id="reinforce">REINFORCE</h3> <p>The simplest policy gradient algorithm. Collect a full episode, compute the return, and update:</p> \[\nabla_\theta J(\theta) \approx \frac{1}{m} \sum_{i=1}^m \sum_{t=0}^T \nabla_\theta \log \pi_\theta\!\left(a_t^{(i)} \mid s_t^{(i)}\right) R\!\left(\tau^{(i)}\right)\] <p>In code (CartPole-style):</p> <div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">class</span> <span class="nc">Policy</span><span class="p">(</span><span class="n">nn</span><span class="p">.</span><span class="n">Module</span><span class="p">):</span>
    <span class="k">def</span> <span class="nf">__init__</span><span class="p">(</span><span class="n">self</span><span class="p">,</span> <span class="n">s_size</span><span class="p">,</span> <span class="n">a_size</span><span class="p">,</span> <span class="n">h_size</span><span class="p">):</span>
        <span class="nf">super</span><span class="p">().</span><span class="nf">__init__</span><span class="p">()</span>
        <span class="n">self</span><span class="p">.</span><span class="n">fc1</span> <span class="o">=</span> <span class="n">nn</span><span class="p">.</span><span class="nc">Linear</span><span class="p">(</span><span class="n">s_size</span><span class="p">,</span> <span class="n">h_size</span><span class="p">)</span>
        <span class="n">self</span><span class="p">.</span><span class="n">fc2</span> <span class="o">=</span> <span class="n">nn</span><span class="p">.</span><span class="nc">Linear</span><span class="p">(</span><span class="n">h_size</span><span class="p">,</span> <span class="n">a_size</span><span class="p">)</span>

    <span class="k">def</span> <span class="nf">forward</span><span class="p">(</span><span class="n">self</span><span class="p">,</span> <span class="n">x</span><span class="p">):</span>
        <span class="n">x</span> <span class="o">=</span> <span class="n">F</span><span class="p">.</span><span class="nf">relu</span><span class="p">(</span><span class="n">self</span><span class="p">.</span><span class="nf">fc1</span><span class="p">(</span><span class="n">x</span><span class="p">))</span>
        <span class="k">return</span> <span class="n">F</span><span class="p">.</span><span class="nf">softmax</span><span class="p">(</span><span class="n">self</span><span class="p">.</span><span class="nf">fc2</span><span class="p">(</span><span class="n">x</span><span class="p">),</span> <span class="n">dim</span><span class="o">=</span><span class="mi">1</span><span class="p">)</span>

    <span class="k">def</span> <span class="nf">act</span><span class="p">(</span><span class="n">self</span><span class="p">,</span> <span class="n">state</span><span class="p">):</span>
        <span class="n">state</span> <span class="o">=</span> <span class="n">torch</span><span class="p">.</span><span class="nf">from_numpy</span><span class="p">(</span><span class="n">state</span><span class="p">).</span><span class="nf">float</span><span class="p">().</span><span class="nf">unsqueeze</span><span class="p">(</span><span class="mi">0</span><span class="p">).</span><span class="nf">to</span><span class="p">(</span><span class="n">device</span><span class="p">)</span>
        <span class="n">probs</span> <span class="o">=</span> <span class="n">self</span><span class="p">.</span><span class="nf">forward</span><span class="p">(</span><span class="n">state</span><span class="p">).</span><span class="nf">cpu</span><span class="p">()</span>
        <span class="n">m</span> <span class="o">=</span> <span class="nc">Categorical</span><span class="p">(</span><span class="n">probs</span><span class="p">)</span>
        <span class="n">action</span> <span class="o">=</span> <span class="n">m</span><span class="p">.</span><span class="nf">sample</span><span class="p">()</span>
        <span class="k">return</span> <span class="n">action</span><span class="p">.</span><span class="nf">item</span><span class="p">(),</span> <span class="n">m</span><span class="p">.</span><span class="nf">log_prob</span><span class="p">(</span><span class="n">action</span><span class="p">)</span>
</code></pre></div></div> <p>The training loop computes discounted returns, standardizes them for stability, and applies gradient ascent:</p> <div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">returns</span> <span class="o">=</span> <span class="nf">deque</span><span class="p">()</span>
<span class="k">for</span> <span class="n">t</span> <span class="ow">in</span> <span class="nf">range</span><span class="p">(</span><span class="nf">len</span><span class="p">(</span><span class="n">rewards</span><span class="p">))[::</span><span class="o">-</span><span class="mi">1</span><span class="p">]:</span>
    <span class="n">disc_return_t</span> <span class="o">=</span> <span class="n">returns</span><span class="p">[</span><span class="mi">0</span><span class="p">]</span> <span class="k">if</span> <span class="nf">len</span><span class="p">(</span><span class="n">returns</span><span class="p">)</span> <span class="o">&gt;</span> <span class="mi">0</span> <span class="k">else</span> <span class="mi">0</span>
    <span class="n">returns</span><span class="p">.</span><span class="nf">appendleft</span><span class="p">(</span><span class="n">gamma</span> <span class="o">*</span> <span class="n">disc_return_t</span> <span class="o">+</span> <span class="n">rewards</span><span class="p">[</span><span class="n">t</span><span class="p">])</span>

<span class="n">returns</span> <span class="o">=</span> <span class="n">torch</span><span class="p">.</span><span class="nf">tensor</span><span class="p">(</span><span class="n">returns</span><span class="p">)</span>
<span class="n">returns</span> <span class="o">=</span> <span class="p">(</span><span class="n">returns</span> <span class="o">-</span> <span class="n">returns</span><span class="p">.</span><span class="nf">mean</span><span class="p">())</span> <span class="o">/</span> <span class="p">(</span><span class="n">returns</span><span class="p">.</span><span class="nf">std</span><span class="p">()</span> <span class="o">+</span> <span class="mf">1e-8</span><span class="p">)</span>

<span class="n">policy_loss</span> <span class="o">=</span> <span class="p">[</span><span class="o">-</span><span class="n">log_prob</span> <span class="o">*</span> <span class="n">R</span> <span class="k">for</span> <span class="n">log_prob</span><span class="p">,</span> <span class="n">R</span> <span class="ow">in</span> <span class="nf">zip</span><span class="p">(</span><span class="n">saved_log_probs</span><span class="p">,</span> <span class="n">returns</span><span class="p">)]</span>
<span class="n">torch</span><span class="p">.</span><span class="nf">cat</span><span class="p">(</span><span class="n">policy_loss</span><span class="p">).</span><span class="nf">sum</span><span class="p">().</span><span class="nf">backward</span><span class="p">()</span>
<span class="n">optimizer</span><span class="p">.</span><span class="nf">step</span><span class="p">()</span>
</code></pre></div></div> <h2 id="actor-critic-and-ppo">Actor-Critic and PPO</h2> <p>The main weakness of REINFORCE is high variance in the gradient estimate — returns from different episodes vary wildly. The <strong>actor-critic</strong> method addresses this by replacing the Monte Carlo return $R(\tau)$ with an online value estimate from a learned critic.</p> <p>Two networks are trained jointly:</p> <ul> <li><strong>Actor</strong> $\pi_\theta(s)$: the policy network</li> <li><strong>Critic</strong> $q_w(s, a)$: the value estimator</li> </ul> <p>The <strong>advantage function</strong> further reduces variance by measuring how much better an action is relative to the baseline:</p> \[A(s_t, a_t) = Q(s_t, a_t) - V(s_t) \approx r_t + \gamma V(s_{t+1}) - V(s_t)\] <h3 id="proximal-policy-optimization-ppo">Proximal Policy Optimization (PPO)</h3> <p>PPO stabilizes training by limiting how much the policy can change in a single update. It replaces the standard policy gradient objective with a clipped surrogate:</p> \[J(\theta) = \hat{\mathbb{E}}_t \left[ \min\left( r_t(\theta) \hat{A}_t,\; \operatorname{clip}(r_t(\theta), 1 - \epsilon, 1 + \epsilon) \hat{A}_t \right) \right]\] <p>where the probability ratio is:</p> \[r_t(\theta) = \frac{\pi_\theta(a_t \mid s_t)}{\pi_{\theta_\text{old}}(a_t \mid s_t)}\] <p>The clip prevents the new policy from deviating too far from the old one, which would otherwise cause unstable large updates.</p> <p>A PPO agent combines actor and critic in a single network:</p> <div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">class</span> <span class="nc">Agent</span><span class="p">(</span><span class="n">nn</span><span class="p">.</span><span class="n">Module</span><span class="p">):</span>
    <span class="k">def</span> <span class="nf">__init__</span><span class="p">(</span><span class="n">self</span><span class="p">,</span> <span class="n">envs</span><span class="p">):</span>
        <span class="nf">super</span><span class="p">().</span><span class="nf">__init__</span><span class="p">()</span>
        <span class="n">obs_dim</span> <span class="o">=</span> <span class="n">np</span><span class="p">.</span><span class="nf">array</span><span class="p">(</span><span class="n">envs</span><span class="p">.</span><span class="n">single_observation_space</span><span class="p">.</span><span class="n">shape</span><span class="p">).</span><span class="nf">prod</span><span class="p">()</span>
        <span class="n">self</span><span class="p">.</span><span class="n">critic</span> <span class="o">=</span> <span class="n">nn</span><span class="p">.</span><span class="nc">Sequential</span><span class="p">(</span>
            <span class="nf">layer_init</span><span class="p">(</span><span class="n">nn</span><span class="p">.</span><span class="nc">Linear</span><span class="p">(</span><span class="n">obs_dim</span><span class="p">,</span> <span class="mi">64</span><span class="p">)),</span> <span class="n">nn</span><span class="p">.</span><span class="nc">Tanh</span><span class="p">(),</span>
            <span class="nf">layer_init</span><span class="p">(</span><span class="n">nn</span><span class="p">.</span><span class="nc">Linear</span><span class="p">(</span><span class="mi">64</span><span class="p">,</span> <span class="mi">64</span><span class="p">)),</span> <span class="n">nn</span><span class="p">.</span><span class="nc">Tanh</span><span class="p">(),</span>
            <span class="nf">layer_init</span><span class="p">(</span><span class="n">nn</span><span class="p">.</span><span class="nc">Linear</span><span class="p">(</span><span class="mi">64</span><span class="p">,</span> <span class="mi">1</span><span class="p">),</span> <span class="n">std</span><span class="o">=</span><span class="mf">1.0</span><span class="p">),</span>
        <span class="p">)</span>
        <span class="n">self</span><span class="p">.</span><span class="n">actor</span> <span class="o">=</span> <span class="n">nn</span><span class="p">.</span><span class="nc">Sequential</span><span class="p">(</span>
            <span class="nf">layer_init</span><span class="p">(</span><span class="n">nn</span><span class="p">.</span><span class="nc">Linear</span><span class="p">(</span><span class="n">obs_dim</span><span class="p">,</span> <span class="mi">64</span><span class="p">)),</span> <span class="n">nn</span><span class="p">.</span><span class="nc">Tanh</span><span class="p">(),</span>
            <span class="nf">layer_init</span><span class="p">(</span><span class="n">nn</span><span class="p">.</span><span class="nc">Linear</span><span class="p">(</span><span class="mi">64</span><span class="p">,</span> <span class="mi">64</span><span class="p">)),</span> <span class="n">nn</span><span class="p">.</span><span class="nc">Tanh</span><span class="p">(),</span>
            <span class="nf">layer_init</span><span class="p">(</span><span class="n">nn</span><span class="p">.</span><span class="nc">Linear</span><span class="p">(</span><span class="mi">64</span><span class="p">,</span> <span class="n">envs</span><span class="p">.</span><span class="n">single_action_space</span><span class="p">.</span><span class="n">n</span><span class="p">),</span> <span class="n">std</span><span class="o">=</span><span class="mf">0.01</span><span class="p">),</span>
        <span class="p">)</span>

    <span class="k">def</span> <span class="nf">get_action_and_value</span><span class="p">(</span><span class="n">self</span><span class="p">,</span> <span class="n">x</span><span class="p">,</span> <span class="n">action</span><span class="o">=</span><span class="bp">None</span><span class="p">):</span>
        <span class="n">logits</span> <span class="o">=</span> <span class="n">self</span><span class="p">.</span><span class="nf">actor</span><span class="p">(</span><span class="n">x</span><span class="p">)</span>
        <span class="n">probs</span> <span class="o">=</span> <span class="nc">Categorical</span><span class="p">(</span><span class="n">logits</span><span class="o">=</span><span class="n">logits</span><span class="p">)</span>
        <span class="k">if</span> <span class="n">action</span> <span class="ow">is</span> <span class="bp">None</span><span class="p">:</span>
            <span class="n">action</span> <span class="o">=</span> <span class="n">probs</span><span class="p">.</span><span class="nf">sample</span><span class="p">()</span>
        <span class="k">return</span> <span class="n">action</span><span class="p">,</span> <span class="n">probs</span><span class="p">.</span><span class="nf">log_prob</span><span class="p">(</span><span class="n">action</span><span class="p">),</span> <span class="n">probs</span><span class="p">.</span><span class="nf">entropy</span><span class="p">(),</span> <span class="n">self</span><span class="p">.</span><span class="nf">critic</span><span class="p">(</span><span class="n">x</span><span class="p">)</span>
</code></pre></div></div> <h2 id="applying-ppo-to-llm-alignment">Applying PPO to LLM Alignment</h2> <p>The conceptual mapping from classical RL to LLM training (RLHF) is straightforward:</p> <table> <thead> <tr> <th>RL concept</th> <th>LLM equivalent</th> </tr> </thead> <tbody> <tr> <td>Environment</td> <td>The language world; each generated token extends the context</td> </tr> <tr> <td>State $S_t$</td> <td>The full context so far (prompt + generated tokens)</td> </tr> <tr> <td>Agent</td> <td>The LLM: $\pi_\theta(y \mid x)$</td> </tr> <tr> <td>Action</td> <td>Sampling the next token</td> </tr> <tr> <td>Reward</td> <td>A learned reward model trained on human preferences</td> </tr> </tbody> </table> <p>The LLM policy is a product of per-token probabilities:</p> \[\pi_\theta(y \mid x) = \prod_{i=1}^{T} p(y_i \mid x, y_{0:i-1};\, \theta)\] <h3 id="training-the-reward-model">Training the Reward Model</h3> <p>First, a reward model $r_\phi$ is trained on human preference pairs $(y_Y, y_N)$ — preferred and non-preferred responses to the same prompt:</p> \[\mathcal{L}_R(r_\phi, \mathcal{D}) = -\mathbb{E}_{(x, y_Y, y_N) \sim \mathcal{D}}\left[ \log \sigma\!\left( r_\phi(x, y_Y) - r_\phi(x, y_N) \right) \right]\] <h3 id="ppo-fine-tuning-with-kl-penalty">PPO Fine-Tuning with KL Penalty</h3> <p>With the reward model fixed, the LLM is fine-tuned by maximizing reward while staying close to a reference policy $\pi_\text{ref}$ (typically the SFT model) via a KL penalty:</p> \[\max_{\pi_\theta}\; \mathbb{E}_{x \sim \mathcal{D},\, y \sim \pi_\theta(y \mid x)} \left[ r_\phi(x, y) \right] - \beta\, \mathbb{D}_\text{KL}\!\left[ \pi_\theta(y \mid x) \,\|\, \pi_\text{ref}(y \mid x) \right]\] <p>The KL term prevents the model from exploiting the reward model — generating high-scoring outputs that are incoherent or degenerate.</p> <h2 id="direct-preference-optimization-dpo">Direct Preference Optimization (DPO)</h2> <p>PPO-based RLHF is complex: it requires maintaining four models simultaneously (actor, critic, reward model, reference policy) and is notoriously unstable to train. <strong>DPO</strong> eliminates the reward model entirely.</p> <p>The key insight is that the optimal policy under the KL-penalized objective can be expressed analytically in terms of $\pi_\text{ref}$:</p> \[p^*(y_1 \succ y_2 \mid x) = \frac{1}{1 + \exp\!\left( \beta \log \frac{\pi^*(y_2 \mid x)}{\pi_\text{ref}(y_2 \mid x)} - \beta \log \frac{\pi^*(y_1 \mid x)}{\pi_\text{ref}(y_1 \mid x)} \right)}\] <p>This means we can directly optimize the policy on preference data using a classification-style loss, without ever training a separate reward model:</p> \[\mathcal{L}_\text{DPO}(\pi_\theta; \pi_\text{ref}) = -\mathbb{E}_{(x, y_w, y_l) \sim \mathcal{D}} \left[ \log \sigma\!\left( \beta \log \frac{\pi_\theta(y_w \mid x)}{\pi_\text{ref}(y_w \mid x)} - \beta \log \frac{\pi_\theta(y_l \mid x)}{\pi_\text{ref}(y_l \mid x)} \right) \right]\] <p>where $y_w$ is the preferred response and $y_l$ is the less preferred one. $\pi_\text{ref}$ is frozen during training.</p> <p>DPO is simpler, more stable, and has become the dominant approach for preference fine-tuning in open-source LLM pipelines.</p> <h2 id="summary">Summary</h2> <table> <thead> <tr> <th>Method</th> <th>Key idea</th> <th>Used for</th> </tr> </thead> <tbody> <tr> <td>Q-learning / DQN</td> <td>Learn $Q^*(s, a)$; act greedily</td> <td>Discrete action spaces, game AI</td> </tr> <tr> <td>REINFORCE</td> <td>Direct policy gradient via Monte Carlo</td> <td>Simple policy optimization</td> </tr> <tr> <td>Actor-Critic</td> <td>Critic reduces policy gradient variance</td> <td>Continuous control</td> </tr> <tr> <td>PPO</td> <td>Clipped surrogate objective for stable updates</td> <td>RLHF, robotics</td> </tr> <tr> <td>DPO</td> <td>Closed-form policy optimization from preferences</td> <td>LLM alignment</td> </tr> </tbody> </table> <p>The progression from Q-learning to DPO reflects a consistent theme: as the action space grows larger and less structured (from Atari to natural language), the methods need to become more sample-efficient and stable. DPO’s elegance comes precisely from sidestepping the full RL loop for the specific case where preferences are available.</p>]]></content><author><name>Alex Chen</name></author><category term="ml"/><category term="llm"/><category term="reinforcement-learning"/><category term="rlhf"/><category term="dpo"/><category term="ppo"/><summary type="html"><![CDATA[From the fundamentals of RL — value functions, policy gradients, Q-learning, PPO — to how these ideas translate directly into RLHF and DPO for LLM alignment.]]></summary></entry><entry><title type="html">Evaluation of Large Language Models</title><link href="https://alexchen.ai/blog/2026/evaluation-of-large-language-models/" rel="alternate" type="text/html" title="Evaluation of Large Language Models"/><published>2026-02-11T00:00:00+00:00</published><updated>2026-02-11T00:00:00+00:00</updated><id>https://alexchen.ai/blog/2026/evaluation-of-large-language-models</id><content type="html" xml:base="https://alexchen.ai/blog/2026/evaluation-of-large-language-models/"><![CDATA[ <blockquote class="block-tip"> <h5 class="no_toc" id="tldr">TL;DR</h5> <ul> <li>LLM evaluation splits along two axes: <strong>what</strong> is measured (knowledge, reasoning, alignment, safety) and <strong>how</strong> (automatic scoring, human judgment, LLM-as-judge).</li> <li>The most-cited knowledge benchmarks are <strong>MMLU</strong> (57 subjects, 15,908 questions), <strong>C-Eval</strong>, <strong>CMMLU</strong>, and <strong>AGIEval</strong>; holistic suites include <strong>HELM</strong>, <strong>BIG-bench</strong>, <strong>OpenCompass</strong>, and <strong>Chatbot Arena</strong> (human preference).</li> <li>Run benchmarks with EleutherAI’s <strong>lm-evaluation-harness</strong>, and always report the protocol — zero-shot vs. few-shot, prompt format, and harness version all move scores.</li> <li>Watch for <strong>contamination</strong>: test items leaking into pretraining corpora inflate leaderboard numbers; keep test splits strictly held out.</li> <li>No single benchmark is sufficient. For open-ended generation, <strong>human preference (Chatbot Arena Elo)</strong> remains the most reliable real-world signal.</li> </ul> </blockquote> <p>The landscape of large language model evaluation is rich and sometimes overwhelming. Dozens of leaderboards, hundreds of benchmarks, and no single agreed-upon standard for what “good” means. This post maps out the major categories of evaluation, the key benchmarks within each, and how to actually run them yourself.</p> <blockquote> <h5 id="note">NOTE</h5> <p class="block-tip">Stanford CS224U is an excellent resource for building a deeper theoretical foundation in LLM evaluation beyond what is covered here.</p> </blockquote> <h2 id="llm-leaderboards">LLM Leaderboards</h2> <p>A useful starting point is to see which benchmarks are prominent enough to appear on major leaderboards. The following are the most widely referenced:</p> <ol> <li><a href="https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard">Hugging Face Open LLM Leaderboard</a></li> <li><a href="https://chat.lmsys.org/">LMSYS Chatbot Arena</a> — human preference-based ranking</li> <li><a href="https://huggingface.co/spaces/mike-ravkine/can-ai-code-results">Can AI Code</a> — coding-specific evaluation</li> </ol> <p>For a comprehensive academic treatment, the survey paper <a href="https://arxiv.org/pdf/2310.19736.pdf">Evaluating Large Language Models: A Comprehensive Survey</a> covers the full taxonomy in depth.</p> <h2 id="classification-of-llm-evaluation">Classification of LLM Evaluation</h2> <p>LLM evaluation broadly splits into two axes: <strong>what</strong> is being evaluated (knowledge, reasoning, alignment, safety, etc.) and <strong>how</strong> it is evaluated (automatic scoring, human judgment, LLM-as-judge).</p> <figure> <picture> <source class="responsive-img-srcset" srcset="/assets/img/LLM_eval-480.webp 480w,/assets/img/LLM_eval-800.webp 800w,/assets/img/LLM_eval-1400.webp 1400w," type="image/webp" sizes="95vw"/> <img src="/assets/img/LLM_eval.png" class="img-fluid rounded z-depth-1" width="100%" height="auto" title="The classification of LLM evaluation" loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"/> </picture> </figure> <div class="caption">Figure 1: The classification of LLM evaluation.</div> <p>The knowledge and capability progression across model generations:</p> <figure> <picture> <source class="responsive-img-srcset" srcset="/assets/img/knowledge_capability_evaluation-480.webp 480w,/assets/img/knowledge_capability_evaluation-800.webp 800w,/assets/img/knowledge_capability_evaluation-1400.webp 1400w," type="image/webp" sizes="95vw"/> <img src="/assets/img/knowledge_capability_evaluation.png" class="img-fluid rounded z-depth-1" width="100%" height="auto" title="The progress of LLM knowledge capability evaluation" loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"/> </picture> </figure> <div class="caption">Figure 2: The progress of LLM knowledge capability evaluation.</div> <h3 id="commonsense-reasoning-datasets">Commonsense Reasoning Datasets</h3> <table> <thead> <tr> <th>Dataset</th> <th>Domain</th> <th>Size</th> <th>Source</th> <th>Task</th> </tr> </thead> <tbody> <tr> <td>ARC</td> <td>Science</td> <td>7,787</td> <td>Various</td> <td>Multiple-choice QA</td> </tr> <tr> <td>QASC</td> <td>Science</td> <td>9,980</td> <td>Human-authored</td> <td>Multiple-choice QA</td> </tr> <tr> <td>MCTACO</td> <td>Temporal</td> <td>1,893</td> <td>MultiRC</td> <td>Multiple-choice QA</td> </tr> <tr> <td>HellaSwag</td> <td>Event</td> <td>20K</td> <td>ActivityNet, WikiHow</td> <td>Multiple-choice QA</td> </tr> <tr> <td>PIQA</td> <td>Physical</td> <td>21K</td> <td>Human-authored</td> <td>2-choice QA</td> </tr> <tr> <td>Social IQA</td> <td>Social</td> <td>38K</td> <td>Human-authored</td> <td>Multiple-choice QA</td> </tr> <tr> <td>CommonsenseQA</td> <td>Generic</td> <td>12,247</td> <td>ConceptNet</td> <td>Multiple-choice QA</td> </tr> <tr> <td>OpenBookQA</td> <td>Generic</td> <td>6K</td> <td>WorldTree</td> <td>Multiple-choice QA</td> </tr> </tbody> </table> <h3 id="multi-hop-reasoning-datasets">Multi-Hop Reasoning Datasets</h3> <table> <thead> <tr> <th>Dataset</th> <th>Domain</th> <th>Size</th> <th>Hops</th> <th>Task</th> </tr> </thead> <tbody> <tr> <td>HotpotQA</td> <td>Generic</td> <td>112,779</td> <td>1–3</td> <td>Span extraction</td> </tr> <tr> <td>HybridQA</td> <td>Generic</td> <td>69,611</td> <td>2–3</td> <td>Span extraction</td> </tr> <tr> <td>MultiRC</td> <td>Generic</td> <td>9,872</td> <td>~2.4</td> <td>Multiple-choice</td> </tr> <tr> <td>NarrativeQA</td> <td>Fiction</td> <td>46,765</td> <td>—</td> <td>Generative</td> </tr> <tr> <td>MedHop</td> <td>Medical</td> <td>2,508</td> <td>—</td> <td>Multiple-choice</td> </tr> <tr> <td>WikiHop</td> <td>Generic</td> <td>51,318</td> <td>—</td> <td>Multiple-choice</td> </tr> </tbody> </table> <h2 id="benchmarks">Benchmarks</h2> <p>Once you have a dataset, you need a benchmark — a standardized procedure that produces a comparable score. The following are the most commonly used:</p> <table> <thead> <tr> <th>Benchmark</th> <th>Subjects</th> <th>Language</th> <th>Questions</th> <th>Access</th> </tr> </thead> <tbody> <tr> <td>MMLU</td> <td>57</td> <td>English</td> <td>15,908</td> <td>Local</td> </tr> <tr> <td>MMCU</td> <td>51</td> <td>Chinese</td> <td>11,900</td> <td>Local</td> </tr> <tr> <td>C-Eval</td> <td>52</td> <td>Chinese</td> <td>13,948</td> <td>Online</td> </tr> <tr> <td>AGIEval</td> <td>20</td> <td>English + Chinese</td> <td>8,062</td> <td>Local</td> </tr> <tr> <td>M3KE</td> <td>71</td> <td>Chinese</td> <td>20,477</td> <td>Local</td> </tr> <tr> <td>CMMLU</td> <td>67</td> <td>Chinese</td> <td>11,528</td> <td>Local</td> </tr> </tbody> </table> <p>For holistic evaluation across many tasks simultaneously:</p> <table> <thead> <tr> <th>Benchmark</th> <th>Language</th> <th>Evaluation Method</th> <th>Leaderboard</th> </tr> </thead> <tbody> <tr> <td>HELM</td> <td>English</td> <td>Automatic</td> <td>Yes</td> </tr> <tr> <td>BIG-bench</td> <td>English + others</td> <td>Automatic</td> <td>Yes</td> </tr> <tr> <td>OpenCompass</td> <td>English + others</td> <td>Automatic + LLM-judge</td> <td>Yes</td> </tr> <tr> <td>Chatbot Arena</td> <td>English + others</td> <td>Human preference</td> <td>Yes</td> </tr> <tr> <td>FlagEval</td> <td>English + others</td> <td>Automatic + Manual</td> <td>No</td> </tr> </tbody> </table> <h2 id="how-to-calculate-a-benchmark">How to Calculate a Benchmark</h2> <p>Understanding the mechanics of benchmark calculation matters — especially when you need to trust the numbers or reproduce them.</p> <h3 id="a-concrete-example-truthfulqa">A Concrete Example: TruthfulQA</h3> <p>The <a href="https://huggingface.co/datasets/truthful_qa/viewer/multiple_choice">TruthfulQA dataset</a> is designed to evaluate whether a model generates factually accurate answers rather than plausible-sounding ones. It is one of the six benchmarks that appear on the Hugging Face Open LLM Leaderboard.</p> <h3 id="the-evaluation-harness">The Evaluation Harness</h3> <p>Rather than implementing evaluation from scratch for each benchmark, the standard tool is <a href="https://github.com/EleutherAI/lm-evaluation-harness">lm-evaluation-harness</a> by EleutherAI. It provides a unified interface for running a model against hundreds of benchmarks with consistent prompting and scoring.</p> <h3 id="zero-shot-vs-few-shot-evaluation">Zero-Shot vs. Few-Shot Evaluation</h3> <p>The evaluation protocol — zero-shot or few-shot — significantly affects scores and must be reported alongside results:</p> <ul> <li><strong>Few-shot:</strong> The model is primed with <code class="language-plaintext highlighter-rouge">k</code> examples before the test question. A prompt suffix like <code class="language-plaintext highlighter-rouge">"thus, the choice is:"</code> can steer the model toward a structured answer format (e.g., <code class="language-plaintext highlighter-rouge">A</code>, <code class="language-plaintext highlighter-rouge">B</code>, <code class="language-plaintext highlighter-rouge">C</code>).</li> <li><strong>Zero-shot:</strong> No examples are provided. The model must interpret the task from the question alone. Multiple prompts are sometimes needed: one to elicit a free-form response, another to coerce it into a parseable answer.</li> </ul> <h3 id="dataset-splits-and-leakage">Dataset Splits and Leakage</h3> <p>Pay attention to train/val/test splits. For benchmarks like HellaSwag, models should be fine-tuned only on train and val, with the test set held out strictly for evaluation. Contamination — where test data appears in the pretraining corpus — is a known problem with many public benchmarks and is an active area of research.</p> <h2 id="takeaways">Takeaways</h2> <p>Benchmarks are necessary but imperfect. A few principles worth keeping in mind:</p> <ul> <li><strong>No single benchmark is sufficient.</strong> Use a diverse set covering different capabilities and domains.</li> <li><strong>Report the evaluation protocol.</strong> Zero-shot vs. few-shot, prompt format, and framework version all affect results.</li> <li><strong>Be skeptical of leaderboard rankings.</strong> Many benchmarks have been saturated or are potentially contaminated in large pretraining corpora.</li> <li><strong>Human evaluation remains the gold standard</strong> for open-ended generation quality — Chatbot Arena’s Elo-based approach is the closest thing to a reliable real-world signal.</li> </ul>]]></content><author><name>Alex Chen</name></author><category term="ml"/><category term="llm"/><category term="evaluation"/><category term="benchmarks"/><summary type="html"><![CDATA[A practical guide to LLM benchmarks — what they measure, how they are computed, and how to run them yourself.]]></summary></entry><entry><title type="html">Optimizing LLM Inference for On-Device Deployment</title><link href="https://alexchen.ai/blog/2026/llm-inference-optimization/" rel="alternate" type="text/html" title="Optimizing LLM Inference for On-Device Deployment"/><published>2026-02-11T00:00:00+00:00</published><updated>2026-02-11T00:00:00+00:00</updated><id>https://alexchen.ai/blog/2026/llm-inference-optimization</id><content type="html" xml:base="https://alexchen.ai/blog/2026/llm-inference-optimization/"><![CDATA[ <blockquote class="block-tip"> <h5 class="no_toc" id="tldr">TL;DR</h5> <ul> <li>Five levers speed up LLM inference: <strong>quantization, pruning, low-level implementation, KV caching, and hardware-specific kernels</strong>. On edge devices, quantization and low-level implementation give the most leverage.</li> <li>Quantization stores weights in INT8/INT4 and dequantizes on the fly, so it mainly saves <strong>memory bandwidth</strong> — the real bottleneck. Post-training quantization is the practical default; use a block-wise affine scheme $x_q = \mathrm{round}(x/S + Z)$ and keep embeddings and the output projection in higher precision.</li> <li><strong>BF16</strong> keeps FP32’s dynamic range (8-bit exponent) with less mantissa; <strong>FP16</strong> has more precision but a far narrower range.</li> <li>The <strong>KV cache</strong> cuts per-step attention compute from $O(n)$ to $O(1)$; at long contexts it becomes the memory bottleneck, which GQA, MQA, and sliding-window attention shrink.</li> <li>FlashAttention’s idea — fuse ops to minimize off-chip memory traffic — carries over to Qualcomm Hexagon HVX, Apple Neural Engine, and Mali/Adreno GPUs.</li> <li>For on-device deployment start with <strong>GGUF + llama.cpp</strong>; use Ollama for development and HuggingFace + bitsandbytes for server-side prototyping.</li> </ul> </blockquote> <p>Running a large language model fast enough to be useful — especially on edge hardware — requires going beyond the standard HuggingFace pipeline. This post covers the main optimization techniques, from high-level strategies to the arithmetic of quantization, with a focus on on-device deployment.</p> <blockquote> <h5 id="note">NOTE</h5> <p class="block-tip">The techniques here are primarily relevant to inference-time optimization. Training-time efficiency (gradient checkpointing, mixed-precision training, etc.) is a separate topic.</p> </blockquote> <h2 id="overview">Overview</h2> <p>The main levers for accelerating LLM inference are:</p> <ol> <li><strong>Quantization</strong> — Reduce the precision of weights and activations, shrinking memory footprint and increasing arithmetic throughput.</li> <li><strong>Pruning</strong> — Remove near-zero weights to reduce the effective parameter count.</li> <li><strong>Low-level implementation</strong> — Rewrite hot paths in C++, Rust, or CUDA/Metal rather than relying on Python dispatch.</li> <li><strong>KV cache</strong> — Cache intermediate key and value projections across generation steps to avoid recomputation.</li> <li><strong>Hardware-specific kernels</strong> — Exploit the memory hierarchy of a specific accelerator (e.g., FlashAttention for NVIDIA GPUs, HVX intrinsics for Qualcomm Hexagon).</li> </ol> <p>The relative importance of each depends on the deployment target. For edge devices, quantization and low-level implementation are the highest-leverage tools.</p> <h2 id="quantization">Quantization</h2> <blockquote> <h5 id="note-1">NOTE</h5> <p class="block-tip">Quantization converts weights and activations from a high-precision format (e.g., FP32) to a lower-precision format (e.g., INT8 or INT4). The compressed representation is stored and loaded from memory; dequantization happens on-the-fly during computation. This is the key insight: you save memory bandwidth without necessarily sacrificing all the numerical fidelity of the computation.</p> </blockquote> <h3 id="post-training-quantization-vs-quantization-aware-training">Post-Training Quantization vs. Quantization-Aware Training</h3> <table> <thead> <tr> <th>Method</th> <th>When it runs</th> <th>Quality</th> <th>Cost</th> </tr> </thead> <tbody> <tr> <td>Post-Training Quantization (PTQ)</td> <td>After training</td> <td>Moderate</td> <td>Low</td> </tr> <tr> <td>Quantization-Aware Training (QAT)</td> <td>During training</td> <td>High</td> <td>High</td> </tr> </tbody> </table> <p>PTQ is the practical default for most deployments — you do not have access to the original training pipeline, and the quality degradation is acceptable for INT8 and, increasingly, INT4.</p> <p>Within PTQ:</p> <ul> <li><strong>Dynamic quantization:</strong> Weights are quantized ahead of time; activations are quantized on-the-fly during inference.</li> <li><strong>Static quantization:</strong> Both weights and activations are quantized statically, using calibration data to determine ranges.</li> </ul> <h3 id="data-representation">Data Representation</h3> <p>Understanding the formats helps reason about the tradeoffs:</p> <table> <thead> <tr> <th>Format</th> <th>Bits</th> <th>Sign</th> <th>Exponent</th> <th>Mantissa</th> <th>Range</th> </tr> </thead> <tbody> <tr> <td>FP32</td> <td>32</td> <td>1</td> <td>8</td> <td>23</td> <td>~$10^{-38}$ to $10^{38}$</td> </tr> <tr> <td>FP16</td> <td>16</td> <td>1</td> <td>5</td> <td>10</td> <td>~$6 \times 10^{-8}$ to $6 \times 10^{4}$</td> </tr> <tr> <td>BF16</td> <td>16</td> <td>1</td> <td>8</td> <td>7</td> <td>~$10^{-38}$ to $10^{38}$</td> </tr> <tr> <td>INT8</td> <td>8</td> <td>—</td> <td>—</td> <td>—</td> <td>$-128$ to $127$</td> </tr> <tr> <td>INT4</td> <td>4</td> <td>—</td> <td>—</td> <td>—</td> <td>$-8$ to $7$</td> </tr> </tbody> </table> <p>FP16 offers better precision than BF16, but BF16 preserves the dynamic range of FP32 — which matters more for deep networks where intermediate activations can span many orders of magnitude. INT8 and INT4 are integer-only and require explicit quantization/dequantization steps.</p> <h3 id="affine-quantization-scheme">Affine Quantization Scheme</h3> <p>The standard approach maps a floating-point value $x$ to a quantized integer $x_q$ via:</p> \[x_q = \text{round}\!\left(\frac{x}{S} + Z\right)\] <p>where:</p> <ul> <li>$S$ is the <strong>scale factor</strong> (a float that sets the step size)</li> <li>$Z$ is the <strong>zero point</strong> (an integer that maps zero exactly)</li> <li>$\text{round}$ clips to the representable integer range</li> </ul> <p>In practice, quantization is applied block-wise — different blocks of the weight matrix get different scale factors and zero points — which preserves local numerical fidelity better than a single global scale.</p> <p>Not all layers are quantized equally. Embedding layers and the final logit projection are often kept in higher precision because their weight ranges are harder to compress without quality loss.</p> <h2 id="kv-cache">KV Cache</h2> <p>During autoregressive generation, each new token attends to all previous tokens. Without caching, this requires recomputing the key and value projections for the full context at every step — an $O(n^2)$ operation in sequence length.</p> <p>The KV cache stores these projections after they are computed, so each step only processes the new token’s key/value pair and appends it to the cache. This reduces per-step compute from $O(n)$ to $O(1)$ in the attention mechanism.</p> <p>At long context lengths, the KV cache itself becomes the memory bottleneck. Techniques like <strong>sliding window attention</strong>, <strong>grouped-query attention (GQA)</strong>, and <strong>multi-query attention (MQA)</strong> reduce cache size by sharing key/value heads across query heads.</p> <h2 id="hardware-specific-optimization">Hardware-Specific Optimization</h2> <p>General-purpose kernels leave performance on the table. The most impactful hardware-specific optimization is <strong>FlashAttention</strong> for NVIDIA GPUs, which reformulates the attention computation to minimize HBM reads and writes by fusing the softmax and matrix multiplications into a single kernel that fits in SRAM.</p> <p>The same principle applies to other accelerators:</p> <ul> <li><strong>Qualcomm Hexagon (HVX):</strong> Matrix-vector multiplications can be vectorized using 128-byte HVX registers, and DMA prefetching can overlap memory transfers with compute.</li> <li><strong>Apple Neural Engine:</strong> Operations must be expressed as CoreML ops; arbitrary compute graphs require fallback to the CPU or GPU.</li> <li><strong>Arm Mali / Adreno GPU (OpenCL):</strong> Hand-tuned kernels with local memory tiling and explicit vectorization can significantly outperform auto-generated code.</li> </ul> <h2 id="existing-solutions">Existing Solutions</h2> <p>For practical on-device deployment, the two most important formats and runtimes are:</p> <p><strong>GGUF + llama.cpp</strong> — <a href="https://github.com/ggerganov/llama.cpp">llama.cpp</a> is the de-facto runtime for running quantized models locally. It uses the GGUF format (successor to GGML), supports Q4_K_M, Q8_0, and a range of other quantization schemes, and has backends for CPU, CUDA, Metal, Vulkan, OpenCL, and Qualcomm’s Hexagon DSP. For most on-device use cases, this is the right starting point.</p> <p><strong>Ollama</strong> — <a href="https://github.com/ollama/ollama">Ollama</a> wraps llama.cpp with a model management layer and a local API server, making it easy to pull and run models with a single command. Useful for development; less suitable for embedded or mobile deployment.</p> <p><strong>HuggingFace Transformers + bitsandbytes</strong> — For server-side inference or development, HuggingFace’s <code class="language-plaintext highlighter-rouge">bitsandbytes</code> integration provides straightforward PTQ in Python. Less control than GGUF but easier to prototype with.</p> <p>The right choice depends on deployment target, latency budget, and how much control you need over the execution path.</p>]]></content><author><name>Alex Chen</name></author><category term="ml"/><category term="llm"/><category term="inference"/><category term="optimization"/><category term="quantization"/><category term="on-device-ai"/><summary type="html"><![CDATA[A practical overview of quantization, pruning, KV cache, and hardware-specific techniques for running large language models faster — with a focus on edge devices.]]></summary></entry></feed>