<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en"><generator uri="https://jekyllrb.com/" version="4.4.1">Jekyll</generator><link href="https://gangliao.me/feed.xml" rel="self" type="application/atom+xml" /><link href="https://gangliao.me/" rel="alternate" type="text/html" hreflang="en" /><updated>2026-10-06T13:37:32-07:00</updated><id>https://gangliao.me/feed.xml</id><title type="html">Gang Liao</title><subtitle>Gang Liao builds the systems that make ML work at scale. Source-backed deep dives on LLMs, GPUs, and AI infrastructure, in English and Chinese.</subtitle><author><name>Gang Liao</name></author><entry xml:lang="en"><title type="html">DeepSeek-V4 and V4-Flash, Layer by Layer: Compressing Attention to Reach a Million Tokens</title><link href="https://gangliao.me/blog/2026/deepseek-v4-deep-dive/" rel="alternate" type="text/html" title="DeepSeek-V4 and V4-Flash, Layer by Layer: Compressing Attention to Reach a Million Tokens" /><published>2026-10-05T00:00:00-07:00</published><updated>2026-10-05T00:00:00-07:00</updated><id>https://gangliao.me/blog/2026/deepseek-v4-deep-dive</id><author><name>Gang Liao</name></author><category term="llm" /><category term="model-architecture" /><category term="ml-systems" /><category term="moe" /><category term="attention" /><summary type="html">At a million tokens of context, the model’s weights stop being the expensive part. The bill is the key-value (KV) cache, the per-token memory that every attention layer keeps so later tokens can look back, plus the work each new token does to read it. DeepSeek-V3.2 already cut the reading cost: its DeepSeek Sparse Attention (DSA) lets each query read only a small, selected set of past tokens. But it still cached an entry for every token in every layer.</summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://gangliao.me/assets/img/og/default.png" /><media:content medium="image" url="https://gangliao.me/assets/img/og/default.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry xml:lang="en"><title type="html">GLM-5.3 and GLM-5.3-Flash, Layer by Layer: Same Skeleton, New Skills — and a Brand-New Hybrid</title><link href="https://gangliao.me/blog/2026/glm-5-3-deep-dive/" rel="alternate" type="text/html" title="GLM-5.3 and GLM-5.3-Flash, Layer by Layer: Same Skeleton, New Skills — and a Brand-New Hybrid" /><published>2026-10-04T00:00:00-07:00</published><updated>2026-10-04T00:00:00-07:00</updated><id>https://gangliao.me/blog/2026/glm-5-3-deep-dive</id><author><name>Gang Liao</name></author><category term="llm" /><category term="model-architecture" /><category term="ml-systems" /><category term="moe" /><category term="attention" /><summary type="html">Z.ai created the Hugging Face repos for two models on the same day, August 25, 2026: GLM-5.3 and GLM-5.3-Flash. On paper they look like siblings. Both use a 154,880-token vocabulary and a 1,048,576-position context window. Both pick the 2,048 most relevant past tokens with a small scorer of the same shape, the lightning indexer (32 heads of 128 dimensions). Both are mixture-of-experts (MoE) models that send each token to 8 experts chosen by sigmoid gating (each expert’s score is squashed to 0–1 on its own, rather than through a softmax across all experts).</summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://gangliao.me/assets/img/og/default.png" /><media:content medium="image" url="https://gangliao.me/assets/img/og/default.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry xml:lang="en"><title type="html">Miles v0.1 Deep Dive: From Train–Rollout Mismatch to Trillion-Parameter Agentic RL</title><link href="https://gangliao.me/blog/2026/miles-v0-1-deep-dive/" rel="alternate" type="text/html" title="Miles v0.1 Deep Dive: From Train–Rollout Mismatch to Trillion-Parameter Agentic RL" /><published>2026-09-29T00:00:00-07:00</published><updated>2026-09-29T00:00:00-07:00</updated><id>https://gangliao.me/blog/2026/miles-v0-1-deep-dive</id><author><name>Gang Liao</name></author><category term="ml-systems" /><category term="reinforcement-learning" /><category term="llm" /><category term="distributed-systems" /><summary type="html">The most dangerous failure in production reinforcement learning (RL) is often not a crashed job. It is a job that keeps running while it optimizes a trajectory the system never actually generated.</summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://gangliao.me/assets/img/og/default.png" /><media:content medium="image" url="https://gangliao.me/assets/img/og/default.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">AMD MI350 GPU Architecture Deep Dive</title><link href="https://gangliao.me/blog/2026/amd-gpu-architecture-deep-dive/" rel="alternate" type="text/html" title="AMD MI350 GPU Architecture Deep Dive" /><published>2026-04-01T00:00:00-07:00</published><updated>2026-04-01T00:00:00-07:00</updated><id>https://gangliao.me/blog/2026/amd-gpu-architecture-deep-dive</id><author><name>Gang Liao</name></author><category term="gpu" /><category term="amd" /><category term="architecture" /><category term="ml-infra" /><summary type="html">AMD’s MI350 series, built on the CDNA 4 architecture, marks a significant leap in datacenter GPU compute. This post breaks down its architecture, programming model, hardware features, and what makes it distinct — both from its predecessor MI300X and from NVIDIA’s competing Blackwell lineup.</summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://gangliao.me/assets/img/og/default.png" /><media:content medium="image" url="https://gangliao.me/assets/img/og/default.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry></feed>