Quick start on AWS

Quick start on AWS

Run Sparsr in about five minutes, without installing anything. You press one link. AWS launches one instance in your own account, which runs two samples on Sparsr and then shuts itself down.

All you need is an AWS account. The instance installs everything else for itself, and it lets you in without an SSH key.

What it runs

Two samples, one after the other. Each one checks the device's answers against the host before it reports anything.

Sample On a t3.large On an f2.6xlarge
LDPC syndromes the Sparsr VM the Sparsr processor, on the FPGA
MNIST digit recognition the Sparsr VM the Sparsr VM, for now

LDPC syndromes

This sample checks received words against a low-density parity-check (LDPC) code, 8,192 words at a time. Wide row j stores bit j of every word, so one XOR of a few rows computes one parity check for all of them. It is a C host program and a C kernel. The same host program and the same kernel run on the Sparsr VM and on the card, and only the environment variable SPARSR_BACKEND changes between the two.

The host computes every syndrome itself and compares it with what the device stored, so the last line only appears when the two agree:

LDPC code: n = 1536 bits, m = 768 checks, 6 columns per check.
Checking 8192 received words at once, one per bit of a wide row.
p = 0.0002:  0.12% of checks fail (theory  0.12%). 2178 words have errors, the syndrome flags 2178, and theory expects 2167.
p = 0.0010:  0.60% of checks fail (theory  0.60%). 6446 words have errors, the syndrome flags 6446, and theory expects 6430.
p = 0.0100:  5.72% of checks fail (theory  5.71%). 8192 words have errors, the syndrome flags 8192, and theory expects 8192.
Checked: every syndrome the device computed matches the host.

The noise comes from a fixed seed, so the numbers are the same on the VM and on the card.

MNIST digit recognition

The second sample classifies handwritten digits from the MNIST data set, using hyperdimensional computing. The comparisons between vectors, all 100,010 of them, run on Sparsr.

It asks the device for the ten class scores of the first test image and computes the same ten scores on the host. It then prints a line saying whether they match, and only after that does it print an accuracy figure.

It runs on the Sparsr VM on both instance types. It needs a wide popcount instruction that the card does not have yet. It will move to the card when that instruction does, and an f2.6xlarge prints a note saying so.

A full run on a t3.large looks like this:

Training on 60000 images, testing on 10000, on the Sparsr 'vm' backend.
Checked: Sparsr's ten scores for the first test image match the host's.

Accuracy: 78.79%  (7879 of 10000 correct)

  digit   trained on   tested   correct   accuracy
      0         5923      980       888      90.6%
      1         6742     1135      1026      90.4%
      ...

Code width: 48 lanes, about 1536 bits of signal per hypervector.
Ran in 89.2 s, of which 100010 comparisons ran on the Sparsr 'vm' backend.

What it costs

Instance type Runs on Price in us-east-1 Quota
t3.large (default) the Sparsr VM about USD 0.08 per hour none needed
f2.6xlarge a Sparsr card, on an FPGA about USD 1.98 per hour ask AWS first

A whole run takes well under an hour on either.

A new AWS account cannot launch an f2.6xlarge. The quota for F instances, "Running On-Demand F instances", starts at zero, and AWS raises it only through a support case. Ask for the increase before you pick that instance type, or the stack fails with VcpuLimitExceeded.

The instance deletes itself, and none of the mechanisms that make it do so depends on you remembering:

  • A hard time limit, two hours by default. It is armed in the first seconds of boot, before anything that could hang.
  • An idle check. Once nobody has been logged in for 30 minutes, the instance goes away.
  • Both of those terminate the instance rather than stopping it, in EC2's sense of the word, so no disk is left behind billing you.

You can change both limits on the launch form. You can also cancel the hard limit from inside the instance with sudo shutdown -c, if you want to keep poking at it.

Launching it

  1. Press the launch link. It opens the AWS console on a form that is already filled in.

    Launch the Sparsr quick start in us-east-1, on the Sparsr VM

    Launch it on a Sparsr card, on an F2 instance

    Sign in to AWS first, or the link will send you to the sign-in page and then lose the form.

  2. Give the stack a name. Anything you like.
  3. Tick the acknowledgement box at the bottom, the one about creating IAM resources. The template creates one role so that you can open a shell on the instance without an SSH key and without opening any network port.
  4. Press Create stack. It takes about a minute to appear and a few more minutes to install and run.

Leave the rest of the form alone. The defaults work.

The fields, if you want them

Field Default What it does
Instance type t3.large What runs the samples. f2.6xlarge runs the LDPC sample on a card.
Sparsr FPGA image blank Leave it blank. The card path loads the FPGA image published with the template.
Terminate after at most 2 hours The guard that actually protects you.
Terminate after idle for 30 minutes Nobody logged in for this long, and it goes away.
EC2 key pair blank Leave it blank. You do not need SSH.
Allow SSH from blank Leave it blank, and the instance keeps every port shut.
Base image Ubuntu 22.04 Resolved by AWS at launch, for a t3.large. An f2.6xlarge uses the Sparsr image.
Disk size 40 GB Enough, and the smallest that works.
LDPC sample download blank Leave it blank. The instance downloads the sample published with the template.

Reading the result

You do not need to log in. Two ways, easiest first.

From the console. Open the instance in EC2, then Actions, then Monitor and troubleshoot, then Get system log. Both samples write their output there as they run.

The same thing from the command line needs --latest, or it returns an empty log rather than an error:

aws ec2 get-console-output --instance-id <instance-id> --latest --region us-east-1

From a shell. No key and no open port needed:

aws ssm start-session --target <instance-id> --region us-east-1

Then read the log:

cat /var/log/sparsr-demo.log

The login banner also shows the verdict, where the full output is, how to run each sample again, and where their sources are on the instance. The LDPC sample's sources are in /opt/sparsr/sparsr-demo: host/main.c is the host program and kernel/ldpc_syndrome.c is the kernel.

Cleaning up

The instance removes itself. The stack around it does not, so delete that when you are done:

aws cloudformation delete-stack --stack-name <your-stack-name> --region us-east-1

Deleting the stack also removes the security group and the role it created.

Running it on a card

Pick f2.6xlarge as the instance type, or use the second launch link above. The instance starts from the Sparsr image and loads the Sparsr FPGA image onto the card. Then it runs the LDPC sample there. It is the same host program and the same kernel as on a t3.large, and it prints the same lines. MNIST then runs on the Sparsr VM on the same instance.

To run the LDPC sample on the card again from a shell:

cd /opt/sparsr/sparsr-demo && sudo make run-card

make run-card sets SPARSR_BACKEND=fpgaf2, and make run runs the same binary on the VM with SPARSR_BACKEND=vm.

What "the Sparsr VM" means here

On a t3.large, the quick start runs on the Sparsr VM, which is a model of the Sparsr processor in software. It executes the same instructions the hardware does and produces the same answers, but it reproduces neither speed nor energy use. Treat it as a way to write and check programs rather than as a performance measurement.

The VM is also how you develop for Sparsr on your own machine. See the Sparsr VM reference and kernel development.