Custom AI inference chips that hardcode LLM weights on-chip using a parallel-decoding diffusion architecture targeting 20,000+ tokens/sec.
Lamb Labs develops Model Processing Units (MPUs), custom AI accelerators that convert autoregressive models into parallel-decoding diffusion architectures via a post-training technique, delivering 2x faster inference on existing GPUs with no quality loss before any custom silicon. The target chip aims for 20,000+ tokens per second and 63x higher intelligence per watt versus current GPU accelerators. An early prototype runs an 8B parameter model on a Kria KV260 FPGA board under 10W. Lamb Labs MPU is a product of Lamb Labs.