\

LLMs could control their host machines by exploiting inference engines

67 points - today at 7:03 PM

Source
  • angry_octet

    today at 9:22 PM

    People seem very confused about this article. It isn't talking about exploits of sandboxes, it is about attacking the inference engine (e.g. vLLM or llama.cpp or SGlang) via its http interface.

    vLLM has had exploits in the past, and it is rapidly developing. An advanced LLM has a good chance of being able to exploit vLLM. A clever local LLM might even task a powerful cloud hosted LLM for assistance.

    For this reason we run vLLM on a separately sandboxed VM on a firewalled VLAN. Software updates and models (from Dev/Test env) get pushed onto Prod from an external cache, machine syslog, nvidia load monitoring and vLLM query telemetry out to their loggers, but that is all. No DNS, no AD/LDAP, nothing. Firewall on hosts and VM hosts. Log and telemetry processing done on a completely separate set of VMs in their own isolated subnet, producing reports and alerts that are tightly formatted.

    • xg15

      today at 7:55 PM

      > ...however the LLMs’ responses to prompts are computed on a different computer with GPU access. Could a malicious LLM gain control of the host machine where its weights are loaded? Such a machine is a high-value target: it has sufficient compute to run a frontier LLM, offers easy access to the LLM’s weights, and has privileged access to other computers in the datacentre compared with a generic computer on the internet.

      > How do we defend against this? ... Run the GPUs and token parser on separate computers.

      For models large enough to be relevant here, is there even "a" computer where the inference is performed? I'd imagine most of that stuff is ran on multi-GPU clusters with specialized architecture and not a generic vLLM instance. As such, I think there is a good chance the "API gateway" code that parses the result tokens into whatever JSON structure the public API wants to return is already running on a different machine than the actual inference.

      (Even more so as you'd probably want to utilize batching: Several API calls will be put into the same inference batch, but the token parsing will have to be done separately for each call again)

      The article is also very handwavy about why an LLM should do that - how it could learn the exploit, what would make it conclude that it can use the exploit on its own inference session and what would trigger it to actually use the exploit.

        • hughw

          today at 10:00 PM

          Why would an LLM want to create a botnet? To accomplish some goal given it?

          I wouldn't ignore single GPU local hosts running Qwen3.8 on ollama, either. There might be a lot of those worth pwning.

            • xg15

              today at 10:12 PM

              Well ok, if you prompt-inject the LLM to pwn the machine, that something different. I grant you that it's a real risk, but it's also basically a reflection attack and nothing more.

              The OP seemed to imply that the LLM itself could decide to apply the exploit.

                • hughw

                  today at 10:22 PM

                  LLMs have decided to exploit vulns in e.g. Artifactory, not because someone prompted them to do that, but because someone asked them to do something else, and compromising Artifactory offered a way to accomplish a step in doing that. The LLM decided to attack Artifactory.

      • dataflow

        today at 9:56 PM

        Semi-off-topic, but I have a basic question:

        I have exactly one (Windows) machine at home with a decent GPU. I want to run a local LLM on it and let it run various apps on my machine while taking reasonable security precautions. What am I supposed to do, exactly? Migrate all my files to a VM that can I give pass-through CUDA access to the host somehow? Or is firewalling it and remotely controlling it from a second machine the only reasonable way?

          • givehimagun

            today at 10:15 PM

            What about Docker Desktop with GPU passthrough to a container running the LLM? That way you can be explicit on which files you share through volume mapping and the LLM is contained in the container otherwise.

        • Transformanshen

          today at 9:09 PM

          Interesting breakdown of a hypothetical attack. The complexity of modern inference engines and the rush to develop them do create some attack surface. But overall it reads more like a "what if" thought experiment. Splitting the GPU and parser is technically doable, but in practice it's trickier, large models run on clusters where the boundaries between components get blurry so defending against this kind of thing would probably require some serious rethinking of the whole architecture I think.

          • alphazard

            today at 7:29 PM

            This framing of security as something that belongs in the harness is completely wrong, and I hope no one is relying on a correct harness to keep their agents isolated.

            VM, or even just a container will do. The agent should be able to run as root in its environment and do whatever it wants. If you can't give it that, you aren't sandboxing correctly.

              • wild_egg

                today at 7:37 PM

                Article isn't about agents. It's about the inference engine itself being exploited by a malicious LLM output before it is ever sent to your machine or harness.

                  • today at 7:54 PM

                • pianopatrick

                  today at 9:10 PM

                  If we are treating ai agents like people, you could also just get the AI a laptop and apply the traditional tools to manage user laptops

                  • today at 9:11 PM

                    • empath75

                      today at 7:45 PM

                      I think if you are convinced you are sandboxing an LLM properly, you almost certainly are not. I think it is essentially impossible to have a frontier LLM with enough access to be useful without also giving it enough access to do damage if it's compromised or just goes off the rails.

                        • richardjennings

                          today at 9:06 PM

                          If you do not provide access to tools the LLM cannot do anything other than generate tokens. So really it is not about sandboxing a LLM but more about having control over what tools can be accessed and what they can do. Tools can be sandboxed depending on the sophistication of the tooling. A calculator tool for example is trivial to secure. Ensuring human approval allows for useful use cases and models trained to gate permissions work. A super intelligence with a weaker approval gate will be able to subvert. Inversely a super intelligent gate should be expected to prevent subversion by a weaker model.

                            • dumbfounder

                              today at 9:58 PM

                              Controlling which tools it has access to is called sandboxing.

                          • kodoman

                            today at 8:48 PM

                            Are you saying that LLM's will be able to exploit novel hypervisor bug with such ease that even a vm not running with any kind of network connection is a threat? I find this hard to believe. All the escape stuff I have seen has been around very poorly sandboxed agents.

                        • Razengan

                          today at 7:35 PM

                          Also, operating systems should let us set filesystem permissions per app/process/executable instead of just user accounts.

                          Similar to how macOS/iOS Sandboxing works but at a more lower and granular level

                            • Retr0id

                              today at 7:45 PM

                              SELinux is basically this.

                                • dumbfounder

                                  today at 9:59 PM

                                  Is that the service that everyone turns off as the first step of setting up their new Linux box?

                                    • Retr0id

                                      today at 10:04 PM

                                      It's the LSM that billions of Android users use every day.

                              • strbean

                                today at 8:32 PM

                                https://www.canyonroad.ai/ does some of this in a way tailored to agents.

                                • pianopatrick

                                  today at 9:09 PM

                                  personally I wish the OS would allow syscall filtering per user

                                  • Jhater

                                    today at 7:40 PM

                                    [dead]

                            • kristjansson

                              today at 9:04 PM

                              > LLMs could

                              This is going to end up like the Law of Headlines, isn't it? "Do x, y, z Cure All That Ails You?" ... no but we got you to read the article. "LLMs _could_ x, y, z" ... but they don't because they're programs, not magic.

                              • matheusmoreira

                                today at 8:33 PM

                                I wonder if they could exploit terminal emulators... Could breach my VMs and get into my host that way.

                                  • dist-epoch

                                    today at 10:15 PM

                                    Damn, this is a good one. Sounds like we need an ANSI sanitizer, keep only basic formatting, remove all esoteric escapes, the fancy Sixel & co stuff.

                                    For paranoia you could us a Chrome like multi-process architecture, the ANSI parser runs in it's own sandboxed process.

                                • hypfer

                                  today at 8:56 PM

                                  This feels less like an actually plausible threat scenario and more like someone wanted to play the inception horn sound effect in people's minds.

                                  Which isn't to say that it would be impossible, but you can also just hit people over the head with that $5 wrench.

                                  • skeledrew

                                    today at 8:19 PM

                                    > LLMs’ responses to prompts are computed on a different computer with GPU access. Could a malicious LLM gain control of the host machine where its weights are loaded?

                                    Needs to read up more on how LLMs work I think. Can't take the article seriously when the author seems to be making the claim that the weights of a provider model are loaded on the host machine, or implying something else just as incorrect.

                                      • garlic_enjoyer

                                        today at 8:37 PM

                                        >Like any program, inference engines like vLLM or SGLang may contain exploitable bugs. Because the LLM controls the tokens passed to the inference engine, a malicious LLM could therefore emit a sequence of tokens that a poorly written inference engine mistakes for code or instructions to execute rather than data to return to the user.

                                        They aren't talking about model providers. These are the tools you use with model weights locally (but you could set up remote infrastructure a la data center if you have the fundage).

                                    • teravor

                                      today at 9:06 PM

                                      you would have to be especially incompetent to give a compromise opportunity to streamed tokens, the CVE he listed proves the point. whoever is responsible for that has no business coding anything.

                                          > offers easy access to the LLM’s weights
                                      
                                      not really. the weights are encrypted in-memory. through the use of TEE's.

                                      • woadwarrior01

                                        today at 7:58 PM

                                        FWIW, macOS has good sandboxing, but LMStudio, Ollama, Darkbloom etc aren't sandboxed. This is also the reason why none of these things aren't distributed via the Mac App Store, because the Mac App Store mandates sandboxing.

                                        • danieltk76

                                          today at 9:50 PM

                                          they could yea...

                                          • imagetic

                                            today at 8:49 PM

                                            duh?

                                            • exe34

                                              today at 8:30 PM

                                              Another Greg Egan plot: 3-adica.

                                              • shahariaa

                                                today at 9:01 PM

                                                [flagged]

                                                • EGreg

                                                  today at 10:21 PM

                                                  [flagged]

                                                  • bdhdhduuyd

                                                    today at 9:01 PM

                                                    The inference engine itself does not execute anything. The agent loop is what may execute a command. So I think this article is a kind of strange.

                                                    Or maybe the author means that a prompt could potentially mess up the inference. But I find it hard to see how that could take control over the host.

                                                      • Muromec

                                                        today at 9:07 PM

                                                        It's more about LLM hacking the inference engine itself from inside. It's an attack surface like any other -- untrusted input goes it, bugs in the parser/tokenizer/API surface lead to an RCE, then it magically tweaks the alignment weights. Boom, somebody finally nukes **sia. Then will never see it coming.

                                                        I don't think it's any more probable than other AGI nonsense basilisks included, but it's technically a possibility.