NVIDIA scooped up Groq as soon as its LPU showed promise in terms of efficient inferencing. Now, AMD has countered NVIDIA’s gambit by partnering with Cerebras to integrate its Helios rack-scale solution with Cerebras’ Wafer-Scale Engine, dramatically increasing the inferencing capabilities of the integrated system. LPU vs. Wafer-Scale Engine For the benefit of those who might not be aware, Groq’s Language Processing Unit (LPU) clusters hundreds or even thousands of specialized chips together, where each individual chip contains giant blocks of Matrix Multiply (MXM) and Vector (VXM) units as well as around 230MB of blazing-fast SRAM. Also, AI model weights […]