\

Speculative Decoding in vLLM on AMD GPUs

44 points - today at 9:26 AM

Source
  • intothemild

    today at 10:36 AM

    Whilst this is an excellent post from vLLM, one of the truly baffling things from either their team or AMDs team, is how much the workstation grade AMD r9700 has been ignored.

    Stock vLLM runs so slowly on these cards compared with vLLM forks like Radiance. Going from say 20-30t/s gen, to 150-200t/s

    Most of AMD/vLLM work seems to be around their data centre cards, or the AMD AI Halo/Ryzen and ignores the R9700 AI Pro.

    Really wish this would change.

      • roenxi

        today at 1:03 PM

        > ...one of the truly baffling things from either their team or AMDs team, is how much the workstation grade AMD r9700 has been ignored.

        It makes a huge amount of sense after considering AMD's approach to graphics cards from around 2010 to 2025. They just didn't see graphics cards as viable compute platform and many who made the mistake of believing that good specs would translate into in-practice performance got badly burned. I'd have been involved in the AI boom but for an expensive AMD graphics card, I'm not going to forget that for a while.

        George Hotz was interesting as a public example, but I think his story probably repeated a few times outside the public eye. People tried to make AMD work and ended up the worse for it.

        People who had an interest in using AMD cards to get things done are probably by and large waiting for a new generation of hopefuls to prove this time is different. The mutterings out of AMD are promising, but that isn't persuasive enough given the scale of the failures.

        • androiddrew

          today at 11:55 AM

          Thankfully, there are still people willing to jump on the R9700 bandwagon and get a vLLM fork working. If you have an RDNA4 card check out https://hub.docker.com/r/stilldeadcode/vllm-radiance

            • intothemild

              today at 12:20 PM

              Deadcode is currently working on INT4 right now on his R4D kernel.

              The MXFP4 fork is excellent too. Its my daily driver right now. https://codeberg.org/ggz14/radiance-vllm-mxfp4

              Also has PARO quant support there too (early stage)

              Also speedups in both repos for 4x R9700s

          • minraws

            today at 10:50 AM

            I don't see a reason why it should AMD doesn't care about lower end prosumers atm.

            They might in the future but future is in the future ofc

            Edit: to be clear I think it's ridiculous they don't but from a company's stand point it doesn't make much sense

              • websap

                today at 10:57 AM

                Yeah, as a business AMD should first care about getting their DC grade hardware optimized for inference workloads. It's unfortunate that most of HN discussion has devolved to me-ish.

                  • _factor

                    today at 11:46 AM

                    Then they should stop selling hardware they don’t plan to support. Me-ish when you spend $1,500 on a piece of hardware is completely acceptable.

            • dist-epoch

              today at 11:24 AM

              George Hotz in June 2023:

              > I have had direct contact with members of the AMD RTG team and I was disgusted to find that AMD doesn't even provide them with hardware to work on. The developer I was working with had to buy the GPU he was writing drivers for.

                • da-x

                  today at 11:34 AM

                  I think this has changed since then, their policies toward open source improved (e.g ROCm).

                    • dist-epoch

                      today at 11:40 AM

                      The market says the problem is still there.

                      An NVIDIA consumer GPU sells for 50+% or more than an equivalent AMD GPU. Because people are buying NVIDIA GPUs to run local models instead of AMD ones.

                      I did the same thing, I paid 50% more to get an 5070 Ti instead of the equivalent AMD.

                      This is probably good for gamers, AMD GPUs are not price inflating to the same degree as NVIDIA, because they are bad at LLMs.

                      > That was the reason for comparing them in the first place: based on performance, they are direct competitors, or at least they are meant to be. However, as things stand today, there is a massive price divide between the two, with the RTX Ti GPU now commanding a premium of more than 50%.

                      https://www.techspot.com/review/3168-geforce-rtx-5070-vs-rad...

                        • esseph

                          today at 12:20 PM

                          > AMD GPUs are not price inflating to the same degree as NVIDIA, because they are bad at LLMs.

                          9060 XT 16GB seems to have some great performance with gpt oss 20B and others, and works great with their lemonade-server.

                          https://lemonade-server.ai/

                          What exactly do you think is running on a Strix Halo?

                          • sznio

                            today at 11:53 AM

                            wondering when AMD will realize it can charge 2x as much for the same thing, by simply finally writing a fucking driver

                              • prymitive

                                today at 12:22 PM

                                The story of “great hardware ruined by poor drivers/software” is so old than its truly shocking its still a thing these days

            • hn45e7pbij

              today at 11:44 AM

              [dead]