{"description":"Trending threats, MITRE ATT\u0026CK coverage, and detection metadata. Fed continuously.","feed_url":"https://feed.craftedsignal.io/products/llama-server-b7492-through-b9060/feed.json","home_page_url":"https://feed.craftedsignal.io/","items":[{"_cs_actors":[],"_cs_cpes":[],"_cs_cves":[{"cvss":8.1,"id":"CVE-2026-43632"},{"cvss":8.1,"id":"CVE-2026-43631"}],"_cs_exploited":false,"_cs_has_poc":false,"_cs_poc_references":[],"_cs_products":["llama-server","llama-server (b7492 through b9060)"],"_cs_severities":["high"],"_cs_tags":[],"_cs_type":"advisory","_cs_vendors":["llama.cpp"],"content_html":"\u003cp\u003eCVE-2026-43632 is a use-after-free vulnerability affecting llama-server builds b7492 through b9060. The vulnerability resides in the handling of six specific tokenization-related endpoints: /tokenize, /detokenize, /infill, /apply-template, /rerank, and /anthropic/count_tokens. These endpoints bypass the standard task queue and access the 'ctx_server.vocab' memory structure directly from HTTP worker threads.\u003c/p\u003e\n\u003cp\u003eThe issue stems from a time-of-check-time-of-use (TOCTOU) race condition. When the application is configured with the '--sleep-idle-seconds' flag, the main thread can destroy and free the vocabulary memory after the synchronization lock is released but before the HTTP worker thread has finished utilizing the reference. This leads to memory corruption, service crashes, or potential remote code execution if an attacker can reliably trigger the race condition while the server is transitioning into an idle state.\u003c/p\u003e\n\u003ch2 id=\"impact\"\u003eImpact\u003c/h2\u003e\n\u003cp\u003eSuccessful exploitation of this vulnerability leads to denial-of-service via application crashes or potential remote code execution on the server hosting the llama-server instance. Given the prevalence of local LLM deployment, this represents a significant risk for organizations using llama-server as a backend component for AI-integrated services or private infrastructure, particularly when exposed to untrusted network traffic.\u003c/p\u003e\n\u003ch2 id=\"recommendation\"\u003eRecommendation\u003c/h2\u003e\n\u003cul\u003e\n\u003cli\u003eUpdate llama.cpp to a build version later than b9060 to mitigate CVE-2026-43632.\u003c/li\u003e\n\u003cli\u003eMonitor access logs for sustained, high-frequency requests to the identified tokenization endpoints, as these may indicate attempts to win the race condition.\u003c/li\u003e\n\u003cli\u003eDisable the '--sleep-idle-seconds' configuration flag in production environments until a patched build is deployed to eliminate the triggering condition for the race.\u003c/li\u003e\n\u003c/ul\u003e\n","date_modified":"2026-08-07T01:30:08Z","date_published":"2026-08-06T23:30:34Z","id":"https://feed.craftedsignal.io/briefs/2026-08-llama-cpp-uaf/","summary":"A use-after-free vulnerability in llama-server allows for potential remote code execution via a TOCTOU race condition in tokenization endpoints when using the --sleep-idle-seconds configuration.","title":"Use-After-Free Vulnerability in llama-server","url":"https://feed.craftedsignal.io/briefs/2026-08-llama-cpp-uaf/"}],"language":"en","title":"CraftedSignal Threat Feed - Llama-Server (B7492 Through B9060)","version":"https://jsonfeed.org/version/1.1"}