We are a memory and AI compute company, and we build AI inference accelerators with optimized computational blocks with integrated memory — delivering radically better tokens per watt.
Modern LLM inference is dominated by transformer matrix multiplications and the movement of weights and activations through memory. GPU-style processors are constrained by energy, HBM bandwidth, cooling, and chip-to-chip interconnects.
Most of the energy is spent moving data — not performing arithmetic. Conventional scaling alone is unlikely to deliver another 10× improvement in performance per watt.
Our A2I accelerator brings AI calculations next to memory — eliminating the data-movement tax that defines today's inference economics.
Memory islands purpose-built for storing and streaming LLM model weights, not generic system RAM.
Matrix-matrix operations execute next to where weights live — bandwidth requirements collapse.
Mixed-signal compute blocks inspired by how biological systems achieve work-per-watt orders of magnitude better than digital ASICs.
Flash and eRAM tiles laid out to feed compute blocks at the right cadence — no idle bandwidth, no wasted joules.
Three dies — non-volatile memory, eRAM, and processor — fabricated separately and stacked into a single package.
NVM
eRAM
PROCESSOR
Mixed-signal multipliers running at sub-W per array, performing the matrix-matrix operations that dominate transformer inference.
Model weights stored in Flash or eRAM, matrix ops adjacent, and a small CPU handling per-layer control flow — all co-packaged.
Tiled arrays each compute a shard of LLM MoE layers — scaling out across chiplets the way models scale out across experts.
Three groundbreaking advantages that set us apart.
Digital computation has reached the limits in scaling and efficiency.
Computation at the physical device level equal gains of tens.
3nm, 5nm? Available? $25M per mask set?
Our fabrication process is available at 1⁄10th of the cost.
Good luck securing your share of DRAM or HBM.
We built our own memory - optimized for AI.
We can help you solve AI at scale in your data center or cluster.
yes but no power
That is our goal too!
We use what you have, no re-inventing the wheel
Gets us to see the light at the end of the tunnel
More compute, at the same or lower power
We're talking to memory manufacturers, hyperscalers, defense partners, and strategic investors. If you build, deploy, or buy LLM inference at scale — let's talk.