Gang Liao
I build the systems that make ML work at scale — from silicon to software. At Meta
, my work spans AI accelerator pre-silicon validation (ASIC), ads model-hardware co-design, GPU kernel & compiler optimization, model training & serving, data infrastructure, and agentic AI systems.
Trained by Daniel Abadi
to never trust a single point of failure (Ph.D., University of Maryland).
Agentic kernel coding (KernelEvolve
, Meta Blog, Import AI, ISCA'26) · distributed serving · across NVIDIA · AMD · MTIA@ Meta
Post-training optimization · agentic RL infrastructure · data curation & synthesis · evals · experience graphs (Trellis, CIDR'27
)@ Meta
KV stores · ML column stores (Bullion, CIDR'25
) · file systems (FileScale, SoCC'23
) · schema evolution (BullFrog, SIGMOD'21
) · HTAP · vector databases@ Meta · ByteDance · Microsoft Research
Core contributor to PaddlePaddle (22k+ ⭐) distributed parallel deep learning training (LLM) framework
@ Baidu Research
Latest writing
-
CAKE from Scratch: Agents Write GPU Kernels, and the Compiler Learns from Their Mistakes
A beginner's guide to CAKE: AI agents write a checkable GPU schedule language, and the compiler turns their repeated failures into new checks.
-
DeepSeek-V4 and V4-Flash, Layer by Layer: Compressing Attention to Reach a Million Tokens
How DeepSeek-V4-Pro and V4-Flash compress the KV cache with CSA and HCA: at 1M tokens, V4-Pro needs 27% of V3.2's per-token FLOPs and 10% of its KV cache.
-
GLM-5.3 and GLM-5.3-Flash, Layer by Layer: Same Skeleton, New Skills — and a Brand-New Hybrid
Layer-by-layer tour of GLM-5.3 (744B, 78 layers of MLA + DSA with IndexShare) and GLM-5.3-Flash (320B, a 45-layer KDA + DSA hybrid with NoPE and mHC).