Apoorva Karnik

builder. maker. tinkerer.

I work at the intersection of hardware and software —
co-designing systems where the two meet for best performance.
When I'm not doing that, I'm reverse engineering things for fun.

Writing →
Learning Mechanistic Interpretability Through ExperimentsA practical curriculum with courses, causal experiments, Mac workflows and a research portfolio path.From a Model to a Verified RTL AcceleratorA practical model-to-RTL workflow covering quantization, dataflow, bandwidth, cycle modeling and verification.Matrix Tiling, Step by Step: From Cache Lines to Race-Free GEMMInteractive tiled matrix multiplication, cache locality, and race-free output ownershipThe AI Memory Boom Has a Hidden Bottleneck: Can We Actually Power All That RAM?Memory demand is surging, but energized power, transformers, turbines, and delivery timelines may decide what actually ships.How Tiled Matrix Multiplication WorksA visual explanation of how large matrix products are split into cache-sized tiles, accumulated across K, and reorganized for efficient memory access.The AI Memory Boom Has a Hidden Bottleneck: Can We Actually Power All That RAM?*August 2026 — a data-driven working thesis with source-labeled models and a monitoring dashboard*Parallel Matrix Multiplication Without Write RacesHow owner-computes scheduling splits a large result matrix into disjoint C tiles so threads can use outer-product updates without atomics or competing writes.Kimi Delta Attention: How It Works, Step by StepKimi Delta Attention (KDA) is the core linear-attention mechanism introduced in Moonshot AI’s **Kimi Linear** architecture. Its goal is ambitious: keep the inference efficiency of recurrent or linear Transformer Inference MathSystems-level explainer on FLOPs, prefill vs decode, and KV cacheFLOPs CalculatorInteractive prefill & decode FLOPs estimatorKV Cache CalculatorLive memory sizing for MHA / GQA / MQA