AMD and Cerebras Launch AI Inference Solution

(cerebras.ai)

24 points | by rbanffy 20 hours ago

4 comments

  • techgnosis 18 hours ago
    I'm not sure I understand. It's Helios but with WSE attached? And it uses one or the other depending on some criteria?
    • wtallis 18 hours ago
      LLM inference is a two-phase process. The first phase is prompt processing aka prefill. It's compute-heavy but requires relatively low memory bandwidth. The second phase is token generation aka decode, which doesn't require much in the way of FLOPs but wants as much memory bandwidth as possible.

      This announcement is for a system to do the first phase on Helios and the second phase on WSE.

  • JSR_FDED 11 hours ago
    It’s disaggregated so you Know it’s good
  • ckrapu 20 hours ago
    They're a bit late to the party.
  • sccvcxv 13 hours ago
    lol when is the music gonna stop?