Critical Alert CVE-2026-5760 exposes remote SGLang code execution using SSTI in Jinja2 templates inside GGUF files

Author: Published 5 min de lectura 461 reading

The images in this article were generated with artificial intelligence. How we publish

A serious vulnerability has been revealed in SGLang which, if used, allows remote execution of code on servers running this framework. The reference CVE-2026-5760 has been assigned and, according to the published CVSS score, its severity is almost maximum: 9.8 over 10.0. In practical terms, an attacker can forge a model file in GGUF format that, when loaded and used by SGLang, triggers the execution of arbitrary commands on the victim team.

The attack vector is based on Jinja2 engine templates that are rendered without the necessary restrictions. In particular, the problem appears when SGLang processes the tokenizer.chat _ template parameter included in a malicious GGUF: that content may contain a Payload of Server-Side Template Injection (SSTI) that, by being run by the template environment used by SGLang, allows you to run Python code in the context of the service. The CERT Coordination Center (CERT / CC) has detailed this flow in its notice, indicating that the affected endpoint is "/ v1 / rerank", commonly used for reeling tasks in inference pipelines.

Critical Alert CVE-2026-5760 exposes remote SGLang code execution using SSTI in Jinja2 templates inside GGUF files
Image generated with IA.

The sequence by which an attacker manages to run his code is simple in its stages: first it creates a GGUF file with a malicious template within tokenizer.chat _ template; that template includes the trigger phrase that makes the path vulnerable in the source code - for example in the entrypoints / openai / serving _ rerank.py file - it is activated; once a victim downloads and loads that model in SGLang and a call to the endpoint of reranking, the template engine renderizes the content and the remote payload is executed, causing the SSTI code to run. This technique exploits a poor configuration of the template engine: jinja2.Environment () is used without the appropriate sandbox instead of a safe variant.

Investigator Stuart Beck is the one who gave this failure and notified it; its analysis points directly to the use of an unprotected Jinja2 environment. To understand why this is dangerous it is necessary to look at previous proposals and attacks: it is not the first time that an inference library or a model charger is exposed for allowing the insertion and rendering of unreliable templates. A well-known case was the nickname "Lama Drama" (CVE-2024-34359), which also allowed arbitrary code execution and was qualified with very high gravity, and more recently similar surfaces were corrected in projects such as VLLM (CVE-2025-61620). Technical information on these EQs is available in the US national vulnerability database. (NVD): CVE-2024-34359 and CVE-2025-61620.

This type of failure is part of the SSTI attack family, well documented by the security community. The rendering templates that allow evaluable expressions can become execution vectors if not properly isolated; OWASP provides a clear explanation of this kind of problem in its section dedicated to Server-Side Temple Injection: OWASP - SSTI. For its part, Jinja2's own documentation explains the difference between normal and sandbox environments designed to limit what templates can do: Jinja2 - ImmutableSandboxedEnvironment.

What can equipment use SGLang do? The first and most forceful thing is to avoid loading models from unverified origins. Public model repositories, such as those available at Hugging Face, facilitate experimentation but also allow the distribution of malicious artifacts if they are not validated: Hugging Face. In parallel, and as an immediate technical measure, it is recommended to replace the use of jinja2.Environment () with ImmutableSandboxedEnvironment or an alternative that applies effective restrictions to the context of the execution of the templates; this modification prevents expressions in the template from triggering the execution of arbitrary Python code.

Although the recommendation to change the Jinja2 environment is the most concrete one, it is appropriate to complement this correction with defensive practices: to run inference services with minimum privileges, to isolate them in containers or dedicated environments, to limit access to the network and to critical resources, and to audit any downloaded model before integrating it into a production instance. Until an official patch is available, caution requires treating any model received from third parties as potentially dangerous.

Critical Alert CVE-2026-5760 exposes remote SGLang code execution using SSTI in Jinja2 templates inside GGUF files
Image generated with IA.

The notice from CERT / CC further states that, during the coordination process, no immediate official solution was obtained by the maintainers, so the adoption of local mitigation and operational controls is particularly important at this time. To be kept informed of technical developments, patches and analysis, it is appropriate to follow official sources of the project and incident response centres, as well as to review NVD entries and CERT / CC communiqués: CERT.

In an ecosystem where high-performance frameworks for multimodal models and LLMs are adopted quickly, the ability to load external models is a powerful feature that, however, also introduces a critical attack surface if no confidence barriers are designed. The practical lesson is clear: it is not enough to rely on the format of the model, it is necessary to control how its components are interpreted and executed.. In the short term, if you manage SGLang instances, avoid loading models from untested sources, apply static analysis when possible and adapt the Jinja2 configuration to use sandboxing environments. In the medium term, the community will need to incorporate more robust controls in the libraries that stop and render data included in GGUF models so that these vectors are not reusable.

If you want to deepen the above technical concepts, consult Jinja2's official documentation on safe environments, OWASP's explanation of SSTI and NVD's entries related to previous remote execution vulnerabilities in inference libraries: the links cited in the text offer a good starting point to understand the scope and mitigation of this type of threat.

Coverage

Related

More news on the same subject.