Show HN: A tiny LLM running at 21,000 tok/s on a $250 FPGA (Live Demo)
An open-source project runs a tiny 3.16M-parameter INT4 language model entirely in a $250 FPGA’s on-chip SRAM, reaching about 21,000 tok/s in a usable single-stream build. The demo is intentionally impractical as a chatbot, but illustrates how eliminating DRAM traffic can transform inference speed and why larger models remain difficult.