Run large language models
locally on devices

Run larger models without increasing memory capacity or power consumption

Our SymbolicLight event-sparse architecture co-designs the model and accelerator to read only the weights activated for each token.

SymbolicLight event-sparse architecture

For each generated token, only some of the model’s intermediate values are nonzero. These nonzero values are events: they determine which weights to read. This sparsity emerges during training. The model retains all its weights and uses a different subset for each token, unlike pruning, which removes weights.

Memory bandwidth is the bottleneck

A dense model reads all its weights from memory for each generated token. With limited memory bandwidth, larger models generate more slowly and consume more energy. Devices must either run smaller models or use higher-bandwidth memory and more powerful chips, increasing cost and power consumption.

Read fewer weights

The event-sparse architecture reduces memory reads. FPGA simulation counters show that the 194M-parameter prototype skips approximately 52% of weight reads per token. Fewer reads allow the same bandwidth to support larger models, or allow the use of less expensive, lower-power memory.

Dedicated hardware is needed

Enabling sparse execution for the same model on a general-purpose ARM processor reduces net energy per token (after subtracting idle power) by only 11%–18%, with unchanged or slightly lower speed. With our accelerator running on an FPGA, sparse execution gains a larger speed advantage as bandwidth narrows.

Prototype measurements

The results below were measured on physical hardware using the 194M-parameter prototype. The platform is a commercial AMD Xilinx Alveo U50C FPGA card running our accelerator design. Card power is measured by onboard sensors and excludes the host.

1.87×

Sparse-to-dense generation speed under constrained bandwidth

The accelerator and weights are identical; only sparse execution is toggled. The speed ratio is 1.87× when weight-read bandwidth is capped at approximately 2.8 GB/s (estimated from the read limit), and 1.32× without the bandwidth cap. In a separate uncapped-bandwidth measurement, net energy for sparse execution is approximately 55% of that for dense execution.

≈1/9

Net energy per token compared with an ARM processor

For a complete request with the same model, 32 input tokens and 128 generated tokens, after subtracting each platform’s idle power: our accelerator uses 0.00912 J/token and the ROCK 5T (RK3588, four A76 cores) uses 0.08633 J/token. Including idle power, total energy is approximately 1/2.8 of the ROCK 5T result. ROCK 5T measurements cover the whole board at the AC input; U50C measurements use the card’s onboard sensors.

643.2token/s

Single-stream generation speed

With 32 input tokens and 128 generated tokens, timing covers generation only. The complete request, including prompt processing, achieves 518.4 token/s.

27,148turns

Conversation turns without restarting the accelerator card

2,263 sessions completed in 30 minutes.

A token is a unit of text processed by the model, not necessarily a character or a word. Full test conditions for energy, speed and stability are documented in paper 2609.09772 below.

Technical papers

Full technical details of the event-sparse architecture and FPGA implementation are publicly available on arXiv.

  1. SymbolicLight V2: Hybrid Neuromorphic Architecture and Sparse Execution for Low-Energy Language Inference

    arXiv 2609.09772 · September 2026 · Second-generation architecture and FPGA implementation

  2. SymbolicLight V1: Spike-Gated Dual-Path Language Modeling at High Encoder Spike Sparsity

    arXiv 2605.21333 · May 2026 · First-generation architecture

Contact us

Device manufacturers, chip companies and developers: contact us to discuss collaboration.

We also welcome enquiries from institutional and individual investors.

info@symboliclight.com

Join us

We are looking for engineers working on models and hardware. Email us your résumé or examples of your work.