.spasm)Sparsr execution units run custom assembly programs stored in .spasm (Sparsr Assembly) files. Sparsr Assembly combines standard 32-bit scalar instructions with wide vector operations designed for high-throughput memory transfers and bitwise compute units.
Sparsr Assembly follows standard MIPS register conventions. Scalar values and wide vector operands are managed using the 32 general-purpose registers:
| Register Name | Number | Usage |
|---|---|---|
$zero / $0 |
$0 |
Hard-wired to 0 |
$at |
$1 |
Assembler temporary |
$v0 – $v1 |
$2 – $3 |
Function return values |
$a0 – $a3 |
$4 – $7 |
Function arguments |
$t0 – $t7, $t8 – $t9 |
$8 – $15, $24 – $25 |
Temporary registers (frequently used for wide vector operands) |
$s0 – $s7 |
$16 – $23 |
Saved registers |
$k0 – $k1 |
$26 – $27 |
Kernel reserved |
$gp, $sp, $fp, $ra |
$28 – $31 |
Global pointer, stack pointer, frame pointer, return address |
$wA, $wB |
- | Wide 4096-bit registers |
The Sparsr processor is custom-built to accelerate operations on binary sparse matrices. Instead of wasting memory bandwidth and ALU cycles on sparse zones (areas filled with zeros), Sparsr implements a dedicated hardware-level compression and decompression engine.
wA or wB register, the Sparsr decompression pipeline restores the full 4096-bit wide vector entirely in hardware. This happens seamlessly at the register boundary, meaning your ALU instructions always operate on the fully uncompressed 4096-bit state.wA and wB, the processor can execute highly parallelized 4096-bit boolean arithmetic (wAND, wOR, wXOR, wNOT) in a single clock cycle.To maximize the massive throughput of the 4096-bit registers and avoid wasting compute cycles on the FPGA, adhere to the following best practices:
wA/wB. To avoid stalls, separate your load instructions from your wide-ALU instructions by interleaving standard 32-bit MIPS instructions (like loop counter decrements, index incrementing, or branch calculations) in between.wA and wB before executing a wide-store back to CMEM.Wide Memory instructions interface directly with Sparsr Co-processor Memory (CMEM) addresses:
WL $rd, CMEM_address – Wide Load: Loads a wide vector from the specified CMEM numerical address into destination register $rd.WS $rs, CMEM_address – Wide Store: Stores a wide vector from source register $rs into the specified CMEM numerical address.Wide Bitwise instructions perform parallel operations across wide vector registers in a single clock cycle:
WAND $rd, $rs, $rt – Wide Bitwise AND: $rd = $rs & $rtWOR $rd, $rs, $rt – Wide Bitwise OR: $rd = $rs | $rtWXOR $rd, $rs, $rt – Wide Bitwise XOR: $rd = $rs ^ $rtSparsr supports standard 32-bit MIPS scalar arithmetic and memory operations:
LW $rd, offset($rt) (Load Word), SW $rd, offset($rt) (Store Word)ADD, ADDU, SUB, SUBU, AND, OR, XOR, NOR (using $rd, $rs, $rt register syntax)NOP $0 – No Operation: Inserts a pipeline stall cycle. Used to pad instructions for execution pipeline latency and hazard prevention.Due to hardware pipeline latencies when loading wide data from CMEM or staging bitwise arithmetic outputs before writing back to CMEM, kernels must account for pipeline hazards.
To avoid pipeline stalls or timing hazards:
NOP $0 cycles (typically 5 NOP instructions) between consecutive wide loads and wide bitwise ALU instructions, or between bitwise ops and wide stores.The simplest kernel loads a wide vector from CMEM address 1 into register $t1 and writes it back to CMEM address 3:
WL $t1,1
WS $t1,3
This kernel loads vectors from CMEM addresses 1 and 2, pads pipeline cycles using NOP $0 instructions, computes bitwise AND, OR, and XOR operations into temporary registers, and stores the results to CMEM addresses 3, 4, and 5:
WL $t1,1
WL $t2,2
NOP $0
NOP $0
NOP $0
NOP $0
NOP $0
WAND $t3,$t1,$t2
WOR $t4,$t1,$t2
WXOR $t5,$t1,$t2
NOP $0
NOP $0
NOP $0
NOP $0
NOP $0
WS $t3,3
WS $t4,4
WS $t5,5
Kernels can also perform scalar memory accesses and arithmetic operations:
LW $t1,1($zero)
LW $t2,2($zero)
ADD $t3,$t1,$t2
AND $t4,$t1,$t2
XOR $t5,$t1,$t2
OR $t6,$t1,$t2
SW $t3,3($zero)
SW $t4,4($zero)
SW $t5,5($zero)
SW $t6,6($zero)
While writing .spasm assembly provides direct control over hardware instructions and register allocation, high-level kernel compilation tools are under active development. The C-to-Sparsr compiler toolchain will automatically vectorize standard C expressions into 4096-bit .spasm instructions, manage CMEM address mappings, and automatically schedule pipeline hazard padding under the hood.
Feel free to contact us requesting early access.