AI inference memory and compute company

VeloxX — the new generation of AI memory and compute.

We are a memory and AI compute company, and we build AI inference accelerators with optimized computational blocks with integrated memory — delivering radically better tokens per watt.

xtremeLess energy
sub-WMultiplier arrays
wideCompute array tiles

The limiting factor in AI is power.

Modern LLM inference is dominated by transformer matrix multiplications and the movement of weights and activations through memory. GPU-style processors are constrained by energy, HBM bandwidth, cooling, and chip-to-chip interconnects.

current chatbot session
700 W
Veloxx session
20 W
There is a path. Not the same old GPU design.

Most of the energy is spent moving data — not performing arithmetic. Conventional scaling alone is unlikely to deliver another 10× improvement in performance per watt.

Compute, redesigned around the memory.

Our A2I accelerator brings AI calculations next to memory — eliminating the data-movement tax that defines today's inference economics.

25×
Lower energy per token (vs. GPU/HBM baseline)
01

Custom memory

Memory islands purpose-built for storing and streaming LLM model weights, not generic system RAM.

02

Near-memory compute

Matrix-matrix operations execute next to where weights live — bandwidth requirements collapse.

03

Brain-logic design

Mixed-signal compute blocks inspired by how biological systems achieve work-per-watt orders of magnitude better than digital ASICs.

04

Efficient storage

Flash and eRAM tiles laid out to feed compute blocks at the right cadence — no idle bandwidth, no wasted joules.

Inside the A2I chiplet.

Three dies — non-volatile memory, eRAM, and processor — fabricated separately and stacked into a single package.

NVM die NVM
eRAM die eRAM
Processor die PROCESSOR
A2I chiplet — 3 dies, 1 package
sub-W Matrix · Matrix Mixed-signal

Optimized multiplier array

Mixed-signal multipliers running at sub-W per array, performing the matrix-matrix operations that dominate transformer inference.

sub-WMatrix · MatrixMixed-signal

Memory + Op + CPU on one chiplet

Model weights stored in Flash or eRAM, matrix ops adjacent, and a small CPU handling per-layer control flow — all co-packaged.

NVM ROMRISC-VCo-packaged

Tiled, sharded MoE

Tiled arrays each compute a shard of LLM MoE layers — scaling out across chiplets the way models scale out across experts.

MoE-nativeTile-scalableProgrammable

Built to win where GPUs can't.

Three groundbreaking advantages that set us apart.

First

Low power computation

Tired

Digital computation has reached the limits in scaling and efficiency.

Wired

Computation at the physical device level equal gains of tens.

Second

Process
advantage

Tired

3nm, 5nm? Available? $25M per mask set?

Wired

Our fabrication process is available at 110th of the cost.

Third

No more DRAM
delays

Tired

Good luck securing your share of DRAM or HBM.

Wired

We built our own memory - optimized for AI.

What are your pain points?

We can help you solve AI at scale in your data center or cluster.

01

Want new data center

yes but no power

02

Want to go back to 40 kW racks

That is our goal too!

03

Co-packaged

We use what you have, no re-inventing the wheel

04

A pilot with us

Gets us to see the light at the end of the tunnel

05

Volume and scale

More compute, at the same or lower power

Tokens per second per watt is the new economics.

We're talking to memory manufacturers, hyperscalers, defense partners, and strategic investors. If you build, deploy, or buy LLM inference at scale — let's talk.