Writing a Kernel in Sparsr Assembly

Writing a Kernel in Sparsr Assembly

Writing Kernels in Sparsr Assembly (.spasm)

Sparsr execution units run custom assembly programs stored in .spasm (Sparsr Assembly) files. Sparsr Assembly combines standard 32-bit scalar instructions with wide vector operations designed for high-throughput memory transfers and bitwise compute units.

Register File & Syntax Conventions

Sparsr Assembly follows standard MIPS register conventions. Scalar values and wide vector operands are managed using the 32 general-purpose registers:

Register Name Number Usage
$zero / $0 $0 Hard-wired to 0
$at $1 Assembler temporary
$v0$v1 $2$3 Function return values
$a0$a3 $4$7 Function arguments
$t0$t7, $t8$t9 $8$15, $24$25 Temporary registers (frequently used for wide vector operands)
$s0$s7 $16$23 Saved registers
$k0$k1 $26$27 Kernel reserved
$gp, $sp, $fp, $ra $28$31 Global pointer, stack pointer, frame pointer, return address
$wA, $wB - Wide 4096-bit registers

Sparse Matrix Compression Mechanics

The Sparsr processor is custom-built to accelerate operations on binary sparse matrices. Instead of wasting memory bandwidth and ALU cycles on sparse zones (areas filled with zeros), Sparsr implements a dedicated hardware-level compression and decompression engine.

  • Hardware Compression: When storing 4096-bit vectors to CMEM, the compression engine evaluates the density of the 1-bits. It natively converts the 4096-bit sequence into a highly optimized block format (similar to compressed sparse row or run-length encoding), aggressively packing zero-dense regions into minimal byte footprints before saving to CMEM.
  • In-Flight Decompression: When a kernel issues a wide-load instruction to read a compressed matrix block from CMEM into the wA or wB register, the Sparsr decompression pipeline restores the full 4096-bit wide vector entirely in hardware. This happens seamlessly at the register boundary, meaning your ALU instructions always operate on the fully uncompressed 4096-bit state.
  • Bitwise Operations: Once decompressed into wA and wB, the processor can execute highly parallelized 4096-bit boolean arithmetic (wAND, wOR, wXOR, wNOT) in a single clock cycle.

Kernel Optimization

To maximize the massive throughput of the 4096-bit registers and avoid wasting compute cycles on the FPGA, adhere to the following best practices:

  • Avoid Pipeline Stalls: The hardware decompression pipeline has a fixed latency when loading from CMEM into wA/wB. To avoid stalls, separate your load instructions from your wide-ALU instructions by interleaving standard 32-bit MIPS instructions (like loop counter decrements, index incrementing, or branch calculations) in between.
  • Maximize 4096-bit Throughput: Always batch your binary sparse matrix processing. Because CMEM loads/stores operate on compressed data but the registers are 4096 bits wide, try to perform as many bitwise operations as possible on wA and wB before executing a wide-store back to CMEM.
  • Data Alignment: Ensure all CMEM offsets used in wide loads/stores are properly aligned to guarantee single-cycle memory interface transfers between the host system and the FPGA CMEM.

Instruction Set Architecture

Wide Memory Operations (CMEM)

Wide Memory instructions interface directly with Sparsr Co-processor Memory (CMEM) addresses:

  • WL $rd, CMEM_addressWide Load: Loads a wide vector from the specified CMEM numerical address into destination register $rd.
  • WS $rs, CMEM_addressWide Store: Stores a wide vector from source register $rs into the specified CMEM numerical address.

Wide Bitwise Operations

Wide Bitwise instructions perform parallel operations across wide vector registers in a single clock cycle:

  • WAND $rd, $rs, $rtWide Bitwise AND: $rd = $rs & $rt
  • WOR $rd, $rs, $rtWide Bitwise OR: $rd = $rs | $rt
  • WXOR $rd, $rs, $rtWide Bitwise XOR: $rd = $rs ^ $rt

Scalar Operations

Sparsr supports standard 32-bit MIPS scalar arithmetic and memory operations:

  • Memory: LW $rd, offset($rt) (Load Word), SW $rd, offset($rt) (Store Word)
  • Arithmetic & Logic: ADD, ADDU, SUB, SUBU, AND, OR, XOR, NOR (using $rd, $rs, $rt register syntax)

Pipeline Control

  • NOP $0No Operation: Inserts a pipeline stall cycle. Used to pad instructions for execution pipeline latency and hazard prevention.

Pipeline Hazards & Latency Optimization

Due to hardware pipeline latencies when loading wide data from CMEM or staging bitwise arithmetic outputs before writing back to CMEM, kernels must account for pipeline hazards.

To avoid pipeline stalls or timing hazards:

  • Interleaving Instructions: Place independent 32-bit scalar instructions (such as loop counter decrements or index calculations) between wide memory loads and subsequent wide bitwise operations.
  • NOP Padding: If independent instructions are unavailable, insert NOP $0 cycles (typically 5 NOP instructions) between consecutive wide loads and wide bitwise ALU instructions, or between bitwise ops and wide stores.

Kernel Examples

1. Minimal Kernel

The simplest kernel loads a wide vector from CMEM address 1 into register $t1 and writes it back to CMEM address 3:

WL $t1,1
WS $t1,3

2. Multi-Operation Bitwise Kernel

This kernel loads vectors from CMEM addresses 1 and 2, pads pipeline cycles using NOP $0 instructions, computes bitwise AND, OR, and XOR operations into temporary registers, and stores the results to CMEM addresses 3, 4, and 5:

WL $t1,1
WL $t2,2
NOP $0
NOP $0
NOP $0
NOP $0
NOP $0
WAND $t3,$t1,$t2
WOR $t4,$t1,$t2
WXOR $t5,$t1,$t2
NOP $0
NOP $0
NOP $0
NOP $0
NOP $0
WS $t3,3
WS $t4,4
WS $t5,5

3. Scalar Memory & Arithmetic Kernel

Kernels can also perform scalar memory accesses and arithmetic operations:

LW $t1,1($zero)
LW $t2,2($zero)
ADD $t3,$t1,$t2
AND $t4,$t1,$t2
XOR $t5,$t1,$t2
OR $t6,$t1,$t2
SW $t3,3($zero)
SW $t4,4($zero)
SW $t5,5($zero)
SW $t6,6($zero)

[Preview] Writing Kernels in C

While writing .spasm assembly provides direct control over hardware instructions and register allocation, high-level kernel compilation tools are under active development. The C-to-Sparsr compiler toolchain will automatically vectorize standard C expressions into 4096-bit .spasm instructions, manage CMEM address mappings, and automatically schedule pipeline hazard padding under the hood.

Feel free to contact us requesting early access.