« All posts

ExLlamaV3 v1.0.0 - Major Performance Upgrades

ExLlamaV3 v1.0.0 brings major performance upgrades and new features for engineers.

After more than a year of development, ExLlamaV3 has released its first production version. Turboderp and Fable have worked tirelessly to implement significant improvements. Key updates include the removal of flash-attention-2 and xformers dependencies, expanded tensor-parallel support, and a new attention kernel that enhances inference speed. These changes are crucial for engineers looking to optimize performance in their applications.

This synthesis was produced from its source by AI; there is no human editor or manual review step. How we work