Gang Liao
I build the systems that make ML work at scale — from silicon to software. At Meta
, my work spans AI accelerator pre-silicon validation (ASIC), ads model-hardware co-design, GPU kernel & compiler optimization, model training & serving, data infrastructure, and agentic AI systems.
Trained by Daniel Abadi
to never trust a single point of failure (Ph.D., University of Maryland).
Agentic kernel coding (KernelEvolve
, Meta Blog, Import AI, ISCA'26) · distributed serving · across NVIDIA · AMD · MTIA@ Meta
Post-training optimization · agentic RL infrastructure · data curation & synthesis · evals · experience graphs (Trellis
)@ Meta
KV stores · ML column stores (Bullion, CIDR'25
) · file systems (FileScale, SoCC'23
) · schema evolution (BullFrog, SIGMOD'21
) · HTAP · vector databases@ Meta · ByteDance · Microsoft Research
Core contributor to PaddlePaddle (22k+ ⭐) distributed parallel deep learning training (LLM) framework
@ Baidu Research
Latest writing
-
DeepSeek-V4 and V4-Flash, Layer by Layer: Compressing Attention to Reach a Million Tokens
How DeepSeek-V4-Pro and V4-Flash compress the KV cache with CSA and HCA: at 1M tokens, V4-Pro needs 27% of V3.2's per-token FLOPs and 10% of its KV cache.
-
GLM-5.3 and GLM-5.3-Flash, Layer by Layer: Same Skeleton, New Skills — and a Brand-New Hybrid
Layer-by-layer tour of GLM-5.3 (744B, 78 layers of MLA + DSA with IndexShare) and GLM-5.3-Flash (320B, a 45-layer KDA + DSA hybrid with NoPE and mHC).
-
Miles v0.1 Deep Dive: From Train–Rollout Mismatch to Trillion-Parameter Agentic RL
Miles v0.1 from source: TITO, R3, and TIS for train–rollout mismatch; async staleness; P2P and disk-delta weight sync; and on-policy distillation.