ESP32-S3 commercial preview

Ship more AI on the silicon you already use.

MicroQuant compiles production neural networks into target-specific firmware—delivering 21.8% lower median latency and 91.0% less product flash than LiteRT Micro with ESP-NN on the validated ESP32-S3 workload.

Built for ESP32-S3 LiteRT models Production firmware
MicroQuant EngineMQ-09
KARQ 14.29 ms / inference
21.8%lower latency than ESP-NN
bit-exactboard output
14.29 msKARQ P1 median latency
92.24%W8 model accuracy
26.2 KBKARQ P1 product flash
30.7 KBKARQ P1 peak SRAM
Model coverage

Supported operators.
Customer-specific delivery.

The current compiler covers the building blocks behind dense networks and compact CNNs. Its architecture is extensible and adaptable to each customer's model, silicon, and production requirements.

Dense Matrix multiplication 2D convolution Depthwise convolution Max pooling Average pooling Global average pooling ReLU and clipping Sigmoid and tanh Flatten and reshape Identity Argmax and top-k
Built to adapt

Your model does not need to fit a generic runtime.

MicroQuant can be extended around customer-specific operators, graph patterns, memory budgets, and target hardware.

Discuss your model
Measured performance

Ahead of LiteRT Micro.
Even with ESP-NN enabled.

Speech Commands v2 DS-CNN on ESP32-S3 at 240 MHz. KARQ P1 delivers 21.8% lower median latency and uses 91.0% less product flash than the vendor-optimized reference path.

Best result1.28×ESP-NN throughput
MicroQuant Conventional pathLower latency is better

Speech Commands v2 DS-CNN

Complete on-device comparison

Download results JSON
ConfigurationExecution pathAccuracyMedian latencyp99 latencyProduct flashPeak SRAM
LiteRT MicroStandard runtime91.99%430.017 ms430.064 ms262.9 KB55.9 KB
LiteRT Micro + ESP-NNVendor optimized91.99%18.281 ms18.361 ms291.4 KB56.4 KB
KARQ P1ESP32-S3 optimized91.40%14.288 ms14.291 ms26.2 KB30.7 KB
W8ESP32-S3 optimized92.24%15.441 ms15.446 ms34.9 KB30.7 KB
PCQ4ESP32-S3 optimized91.36%15.734 ms15.737 ms24.8 KB30.7 KB
KARQ P4ESP32-S3 optimized90.14%17.007 ms17.010 ms35.0 KB30.7 KB
SPQ4 G8ESP32-S3 optimized91.70%37.646 ms37.649 ms34.3 KB30.7 KB
SPQ4 G8Portable runtime91.70%99.128 ms99.130 ms34.1 KB30.7 KB
KARQ P1Portable runtime91.40%115.612 ms115.614 ms26.0 KB30.7 KB
KARQ P4Portable runtime90.14%130.364 ms130.366 ms34.8 KB30.7 KB
PCQ4Portable runtime91.36%232.199 ms232.202 ms24.6 KB30.7 KB
W8Portable runtime92.24%232.601 ms232.605 ms34.7 KB30.7 KB
01

Bit-exact execution

Host and device outputs match exactly across every MicroQuant configuration.

02

Production memory discipline

Static execution arenas and zero runtime heap growth keep behavior predictable.

03

Reproducible delivery

Source, generated assets, build identity, and hardware results are bound into the checkpoint.

Commercial engagement

A short path from
model to measured pilot.

Start with the model you already have. We qualify the graph and target, build the best candidates, and return a board-tested integration with decision-ready evidence.

  1. 01

    Model assessment

    Share the model, target board, and product constraints. We confirm fit and define acceptance goals.

  2. 02

    Optimization sprint

    MicroQuant evaluates precision formats and compiles the strongest candidates for your target.

  3. 03

    Hardware pilot

    You receive integrated firmware, measured results, and a clear route to production licensing.

Start a conversation

What could your model do with more room to run?

Tell us what you are building. We will reply with a focused technical assessment and the next practical step.

Typical first stepModel + target review

Request received.

We will follow up at .