WMEM, XMEM and the DMA

WMEM, XMEM and the DMA

Sparsr keeps rows in two memories. WMEM is small, on the chip, and the only memory a wide instruction can address. XMEM is large, off the chip, and only the DMA can copy rows in and out of it. A kernel computes on WMEM and uses XMEM as storage behind it.

What runs today. The Sparsr VM in the SDK has WMEM, XMEM and the DMA, with the behaviour this page describes, from the nightly SDK build of 16 September 2026 onward. The VM's memories are smaller than the card's: see the memory map. The FPGA hardware has a small row memory today, and XMEM and the DMA are still being built.


1. Rows

A row is one wide register's worth of data: 8,192 bits, or 1,024 bytes. It is divided into 256 lanes of 32 bits each. Every memory on this page is counted in rows.

A raw row is its 1,024 bytes stored as they are. A LIL row is stored compressed, as a list of only the lanes that hold a 1-bit. See Host Wire Format for the LIL format.


2. WMEM: the rows a kernel computes on

Property Value
Physical memory on-chip URAM
Size 4,096 rows, 4 MiB
Row format raw, always
Load time a fixed two clock cycles
Reached by wide load and store, by row number; RV32I loads and stores, by byte; the DMA; the host
  • A wide load or store addresses WMEM only. WL copies a row into a wide register, and WS copies a wide register into a row. Neither takes an XMEM row number.
  • Every row in WMEM is raw. A wide load never decompresses anything, so it always takes the same time, and no row is too dense to store.
  • The same rows have byte addresses. Byte k of row n is at 0x9000_0000 + n * 1,024 + k for RV32I loads and stores. See the memory map.

3. XMEM: the rows behind WMEM

Property Value
Physical memory high-bandwidth memory beside the FPGA
Size 16 GiB, about 16 million rows
Row format as the host sent it, raw or LIL
Reached by the DMA, and the host
  • No instruction addresses XMEM. It has no row number in a wide instruction and no address in the scalar map. XMEM row numbers appear only in DMA requests and in host calls.
  • XMEM keeps a row in the form the host sent it. A LIL row is kept compressed in XMEM until the DMA copies it into WMEM.
  • A load from XMEM would be slow and would not take a fixed time. That is why the core never loads from it directly. A copy engine lets the core see one memory with one fixed load time.

4. The DMA

The DMA (direct memory access engine) copies a block of rows between XMEM and WMEM. The core keeps running while it copies.

Copy What happens to each row
XMEM into WMEM A LIL row is expanded to raw. A raw row is copied as it is.
WMEM into XMEM The row is copied raw. The card never compresses a row.

Who can start a copy

A running kernel and the host both can, through the same registers. The registers are at 0xFFE0_0000 in the scalar address map. The memory map lists them.

  • A kernel writes a request with ordinary stores, then reads the status register until the copy is done.
  • The host writes the same registers through the host library. A host copy has finished by the time the call returns.

The rules of a copy

  • One copy at a time. A second request while one is running is refused and does nothing. The status register reports busy. The batch keeps running.
  • Do not touch rows that a copy is still writing. A wide load or store of a WMEM row the DMA is filling is a fatal fault. A half-written row is never returned.
  • Other errors are reported, not trapped. A row past the end of either memory, or a request on a device with no XMEM, sets the fault field in the status register. The copy does not happen, and the batch keeps running.
  • Every register access is a 32-bit, word-aligned load or store. A byte access to the block is a fatal fault.

5. Working through a data set larger than WMEM

A kernel works on one tile of rows and asks for the next tile before it needs it. A tile is a block of rows that fits in WMEM.

  1. The host writes the data set into XMEM, raw or LIL.
  2. The kernel starts a DMA copy of tile 1 into WMEM, and waits for it.
  3. The kernel starts a DMA copy of tile 2 into different WMEM rows.
  4. While that copy runs, the kernel computes on tile 1.
  5. The kernel waits for the copy. If the computing took longer than the copy, the kernel does not wait at all.
  6. It repeats with the next tile, and copies result rows out to XMEM.
  7. The host reads the results from XMEM once, at the end.

Because the engine runs one request at a time, a kernel can start a copy at most one tile ahead.

In C, the SDK header wraps the registers:

#include "sparsr_intrinsics.h"

void kernel_main(void) {
    _sparsr_dma_to_wmem(64, 0, 8);   /* XMEM rows 64 to 71 into WMEM rows 0 to 7 */
    /* ... work on rows the kernel already has ... */
    _sparsr_dma_wait();              /* costs nothing if the copy has already finished */

    if (_sparsr_dma_fault() != 0) {
        _sparsr_exit(1);             /* the request was refused */
    }
}
Intrinsic What it does
_sparsr_dma_to_wmem(src, dst, rows) Starts a copy from XMEM row src into WMEM row dst
_sparsr_dma_to_xmem(src, dst, rows) Starts a copy from WMEM row src out to XMEM row dst
_sparsr_dma_wait() Waits until the copy has finished
_sparsr_dma_busy(), _sparsr_dma_done() Tests the status register
_sparsr_dma_fault() The fault code of the last request, or 0 when it was accepted
_sparsr_dma_rows_done() How many rows the last request copied

On the host, sparsr.h declares sparsr_write_data_xmem(), sparsr_read_data_xmem(), sparsr_xmem_rows() and sparsr_dma_copy().


See also