A practical curriculum with courses, causal experiments, Mac workflows and a research portfolio path.
Writing
Notes on systems, ML infrastructure, and things I find interesting.
A practical model-to-RTL workflow covering quantization, dataflow, bandwidth, cycle modeling and verification.
An interactive walk through tiled matrix multiplication, memory locality, outer products, and thread-owned output tiles.
Memory demand is surging, but energized power, transformers, turbines, and delivery timelines may decide what actually ships.
A visual explanation of how large matrix products are split into cache-sized tiles, accumulated across K, and reorganized for efficient memory access.
*August 2026 — a data-driven working thesis with source-labeled models and a monitoring dashboard*
How owner-computes scheduling splits a large result matrix into disjoint C tiles so threads can use outer-product updates without atomics or competing writes.
Kimi Delta Attention (KDA) is the core linear-attention mechanism introduced in Moonshot AI’s **Kimi Linear** architecture. Its goal is ambitious: keep the inference efficiency of recurrent or linear
A systems-level explainer for transformer computation, attention, dense layers, prefill, and decode.
Interactive calculator for linear-layer work, quadratic attention, prefill, and decode.
Live KV-cache memory sizing for MHA, GQA, and MQA.