← Writing

From a Model to a Verified RTL Accelerator

Oct 7, 2026

The practical path to an accelerator starts with an executable workload and ends with a verified hardware contract. This guide stops at RTL, cycle simulation and optional synthesis sanity checks. It does not require an FPGA board or physical design.

The first project is a small quantized classifier with an eight-lane dot-product engine. It deliberately exposes the decisions that matter: numeric range, weight reuse, memory ports, dataflow, control and end-to-end latency. After completing it, extend the same process to an LLM projection or a sensor-processing kernel.

Resources checked October 7, 2026. Project dimensions and performance targets below are proposed teaching assumptions, not measured silicon results.

The best learning resources for this route

ID Resource Practical contribution Access or constraint
A01 MIT 6.5930 Hardware Architecture for Deep Learning Mapping, dataflow, partitioning, fusion, sparsity and numeric precision Public notes; Spring 2026 schedule says lectures are not recorded. Labs link to GitHub Classroom; external access is not assured.
A02 MIT 6.5940 Efficient AI Computing Compression, quantization and deployment experiments Public lecture links and selected lab links. Fall 2026 forbids cross-registration and prerequisite-waiver petitions. Some later links are placeholders.
A03 Cornell ECE5760 Concrete digital-design lectures, labs and student hardware projects Public course material; FPGA-specific portions require the relevant vendor tools and board.
A04 AccelForge Architecture and mapping exploration; linked by MIT's 2026 course Use the current tutorial and environment instructions.
A05 Timeloop and Accelergy Workload-to-architecture mapping, performance/energy estimation Model-based estimates; validate selected assumptions against your own cycle model.
A06 Berkeley Gemmini Real accelerator organization: array, scratchpad, accumulation, DMA and queues Chisel/Chipyard reference; integration is substantially heavier than standalone RTL.
A07 FINN documentation Quantized-network dataflow generation and folding tradeoffs Build requirements specify Linux, Docker and Vitis/Vivado; do not assume native Apple Silicon support.
A08 hls4ml A comparison route from trained model to HLS hardware Useful for precision/reuse exploration; HLS is not a substitute for learning the microarchitecture.
A09 Verilator Compile and simulate SystemVerilog locally Use documented platform/build requirements and record the installed version.
A10 cocotb Python scoreboards and randomized protocol testing Pair a supported cocotb release with the simulator; use version-matched examples.
A11 Machine Learning Systems book System-level constraints and deployment context Supporting reading rather than an RTL lab sequence.

Recommended combination: MIT 6.5930 for architectural reasoning, selected MIT 6.5940 material for numeric/model choices, Cornell for implementation examples, and your own SystemVerilog plus Python verification. Gemmini becomes more valuable after you can explain your simpler engine's bottleneck.

Step 1 Define the workload contract

Start with a handwritten-digit classifier using the public MNIST interface in torchvision. Use a 784 → 64 → 10 MLP with ReLU between layers. Train it on the Mac or reuse an explicitly licensed checkpoint with a fixed revision. Record the training seed, preprocessing and held-out accuracy.

The design brief is:

  • One image per request, batch size one.
  • Pixel normalization is explicitly defined and identical in training and deployment.
  • INT8 symmetric weights and activations; INT32 accumulation initially.
  • Weights remain resident between inferences.
  • Output is ten scores; classify by argmax, with no hardware softmax.
  • Proposed accuracy budget: at most one percentage point below the floating baseline.
  • Proposed application deadline: one millisecond per inference.
  • RTL first; any clock frequency used in planning is an assumption.

This MLP has 50,816 weight MACs per inference: 784 × 64 + 64 × 10. Its INT8 weights occupy 50,816 bytes, about 49.6 KiB, before biases, scales and layout padding. The layer shapes and simplicity make failures easy to localize.

Step 2 Measure the software baseline

Run a reproducible benchmark with warm-up and fixed test samples. Measure preprocessing, each layer and end-to-end latency separately. For MPS, synchronize around timed work; otherwise you may measure command submission instead of completion. Save median and tail latency and document whether weights are already loaded.

Do not conclude that the largest arithmetic count must be the largest observed Mac latency. Small kernels can be dominated by dispatch overhead. Hardware selection should combine workload structure with measurements and a target-system cost model.

Create three artifacts: a model checkpoint, a shape/operator inventory, and per-layer input/output vectors. Those vectors will become the RTL test corpus.

Step 3 Build the bit-exact numeric reference

Before RTL, implement the actual integer arithmetic in Python or NumPy. For symmetric quantization:

qx = clip(round(x / sx), -127, 127)
qw[j] = clip(round(w[j] / sw[j]), -127, 127)
bias_q[j] = round(bias[j] / (sx * sw[j]))
acc[j] = bias_q[j] + sum_i qx[i] * qw[j,i]
qy[j] = clip(round(acc[j] * sx * sw[j] / sy), output_range)

Use one activation scale per layer and per-output-channel weight scales initially. Specify the exact rounding and saturation rule; Python defaults are not a hardware specification. Translate the rescale ratio to an integer multiplier and shift, then make the reference use that same multiplier, shift and tie rule.

An INT8 dot product of length 784 has worst-case product-sum magnitude at most 784 × 127² = 12,645,136 for this symmetric range. A signed 25-bit accumulator can hold that sum before bias; INT32 gives margin, but validate bias and every supported length explicitly. Widen the rescale multiply before shifting.

There is a subtle final-layer issue: with per-channel weight scales, raw INT32 accumulators are in different units. Requantize all ten outputs to a common score scale before argmax, or compare rescaled values. Comparing raw per-channel accumulators is incorrect.

Exit gate: held-out integer-reference accuracy meets the budget. If not, adjust calibration, scale granularity or precision before writing hardware. Hardware that perfectly implements an unacceptable approximation is not a successful accelerator.

Step 4 Choose a dataflow from bandwidth

Begin with eight parallel lanes reducing one output neuron at a time. Store each layer's weights in output-major order so each cycle reads eight adjacent weights. Pack or bank activation SRAM to supply eight corresponding activations.

For layer one, 784 / 8 = 98 accepted input groups per neuron. Across 64 outputs this is 6,272 groups. Layer two needs 64 / 8 = 8 groups per output, or 80 total. The ideal total is 6,352 accepted compute groups. At an assumed 100 MHz that is 63.52 microseconds, excluding issue overhead, pipeline drain, SRAM delays, requantization and I/O. This is a calculated lower bound, not a benchmark or achievable-frequency claim.

Eight lanes need eight weight bytes and eight activation bytes per active cycle at the local memory interface: 16 bytes/cycle. At 100 MHz, the implied aggregate read bandwidth is 1.6 GB/s. External bandwidth can be much lower with resident weights and activation reuse, but local ports must still sustain the schedule.

Compare two alternatives in a Python cycle model:

Choice Benefit Cost to quantify
Eight lanes across the reduction dimension Simple sequential weight stream and one output accumulator Re-read activations for each output; reduction tree
Eight outputs in parallel Broadcast each activation across eight output accumulators Different weight banking, more accumulator state
Small two-dimensional array More reuse for larger matrices/batches Fill/drain overhead and poor utilization for narrow shapes

For this initial batch-one MLP, prefer the simplest design that meets the deadline. A systolic array is not automatically the best answer for matrix-vector workloads.

Step 5 Write the architecture specification

Define the modules and their contracts before coding:

Module Contract
Command interface Start, dimensions, base addresses, scale-table pointer, completion/error status
Weight and activation stores Explicit capacity, read latency, port widths, address layout and tail masking
Address generator Iterates output index and reduction groups; never reads beyond valid buffers
MAC datapath Eight signed products, widened reduction, accumulator clear/load and update
Requantization Bias placement, wide multiply, signed rounding, shift, clamp and optional ReLU
Output FIFO Valid/ready backpressure; stable payload while blocked
Controller Owns each buffer and prevents overwrite until consumption

A one-entry output buffer is enough for a first design. Reject a new command while busy. Expose counters for accepted MAC groups, memory stalls and output stalls. Define reset during an active command to abort it and clear pending output, and make the testbench expect that behavior.

Do not describe idealized arrays as SRAMs without implementing their port and latency restrictions. A model with unlimited instantaneous reads can make an infeasible architecture look fast.

Step 6 Implement RTL in increments

Implement one lane first, then the reduction tree, then the controller and memories. This illustrative accumulation fragment shows the required signed extension; it is not a complete accelerator:

logic signed [7:0] a, w;
logic signed [15:0] product;
logic signed [31:0] product_ext, acc;

assign product = a * w;
assign product_ext = {{16{product[15]}}, product};

always_ff @(posedge clk) begin
  if (rst)
    acc <= '0;
  else if (clear)
    acc <= '0;
  else if (in_valid && in_ready)
    acc <= acc + product_ext;
end

This interface gives clear priority over acceptance. Its controller must not assert clear on a cycle that claims to accept a product. If simultaneous clear-and-first-product is desired, define and implement that separately. For the eight-lane version, widen the reduction before addition and carry valid/tag state through every pipeline stage.

Write synthesizable RTL without simulator-specific arithmetic shortcuts. Start with behavioral synchronous memories that model the intended ports. Technology mapping and physical closure remain outside this project.

Step 7 Verify against the reference

Use cocotb or an existing C++ harness with Verilator. Compare every output bit against the integer reference, not just the final class.

Required cases:

  • Zero, extrema, alternating signs, one-hot and randomized inputs and weights.
  • Bias and rescale boundaries, negative rounding ties and saturation.
  • Reduction lengths below eight, exactly eight, and nonmultiples of eight.
  • Input gaps and randomized output backpressure.
  • Reset mid-command and consecutive legal commands.
  • Stable output during a stall; no lost, duplicated or reordered results.
  • End-to-end held-out examples using the exported model weights.

A scoreboard should track accepted transactions rather than wall-clock loop iterations. Save failing random seeds and waveforms. Separate numerical failures from protocol failures so an apparent accuracy regression does not hide a dropped result.

Step 8 Compare the model with RTL cycles

Record cycles per layer, useful MACs, idle cycles and stalls. Explain deviations from the 6,352-group ideal. Sweep lane count 1, 2, 4, 8 and 16 with consistent memory bandwidth; do not scale compute and secretly give the larger design unlimited ports.

Use end-to-end Amdahl accounting. If an accelerated kernel originally contributes fraction f and its speedup is s, the no-overhead bound is 1 / ((1-f) + f/s). Add transfers and setup explicitly. Comparing RTL cycles with Mac wall time requires an assumed clock and complete I/O costs; label such comparisons as projections.

Optional synthesis can expose latches, unexpected multipliers or huge muxes. Generic cell counts are not physical area, and functional simulation does not establish timing, power or manufacturability.

Step 9 Extend to a model you care about

Move to an LLM linear projection after the MLP works. Extract actual dimensions from the chosen model and capture activations separately for prefill and decode. Batch-one decode is often a matrix-vector reuse problem; prefill offers more matrix-matrix reuse. Model weight traffic, quantization scales, activation movement and dispatch costs before choosing an array.

Alternative follow-ons are a 1D convolution for sensor streams, FIR/FFT processing, or one repeated matrix operation in a Kalman filter. Keep preprocessing and unsupported operations in software initially. A hardware/software boundary is a deliberate architectural decision.

Eight implementation milestones

  • A workload contract and measured floating baseline exist.
  • A bit-exact integer model meets the accuracy budget.
  • Buffer layout and memory-port demand are calculated.
  • A cycle model compares at least two dataflows.
  • A single arithmetic unit passes corner cases.
  • The complete engine passes randomized protocol and numerical tests.
  • RTL cycles match an explained performance model.
  • A report states accuracy, cycle count, bandwidth, storage and unresolved physical assumptions.

Plan roughly 8–12 part-time weeks, depending on scope and tool familiarity. The deliverable is a defensible IP design with evidence, rather than an impressive peak MAC count.