« All posts

Full-fabric VHDL LLM Inference Engine: Runs Qwen3.5-class

The full-fabric VHDL LLM inference engine runs Qwen3.5-class transformer inference on FPGA.

The full-fabric VHDL LLM inference engine executes Qwen3.5-class transformer inference, processing 9B on a single card and targeting 27B across two cards. This system operates entirely within FPGA fabric, utilizing INT4 streaming matvec, Gated DeltaNet, and gated attention. Outputs are validated against llama.cpp, ensuring accuracy through comparisons with a bit-accurate C model of the INT4 datapath.

This synthesis was produced from its source by AI; there is no human editor or manual review step. How we work