{"description":"Trending threats, MITRE ATT\u0026CK coverage, and detection metadata. Fed continuously.","favicon":"https://feed.craftedsignal.io/favicon-32x32.png","feed_url":"https://feed.craftedsignal.io/products/mooncake-connector--0.29.0/feed.json","home_page_url":"https://feed.craftedsignal.io/","icon":"https://feed.craftedsignal.io/apple-touch-icon.png","items":[{"_cs_actors":[],"_cs_cpes":["cpe:2.3:a:vllm:mooncake_connector:*:*:*:*:*:*:*:*"],"_cs_cves":[{"cvss":7.5,"id":"CVE-2026-94627"}],"_cs_exploited":false,"_cs_has_poc":false,"_cs_poc_references":[],"_cs_products":["Mooncake connector (\u003c= 0.29.0)"],"_cs_severities":["low"],"_cs_tags":[],"_cs_type":"advisory","_cs_vendors":["vLLM"],"content_html":"\u003cp\u003eThe vLLM Mooncake connector, used in disaggregated prefill/decode deployments, contains a vulnerability in the management of GPU Key-Value (KV) cache block ownership. When concurrent child requests share a single transfer ID, the connector fails to properly reconcile ownership, leading to the accumulation of orphaned memory blocks. An attacker can deliberately trigger this condition by submitting completion requests containing multiple prompts, forcing the system to allocate memory that is never subsequently freed. This persistent memory leak eventually results in GPU memory exhaustion, preventing the processing of legitimate requests and causing a service-wide denial of service until the affected process is restarted. This vulnerability is identified as CVE-2026-94627.\u003c/p\u003e\n\u003ch2 id=\"impact\"\u003eImpact\u003c/h2\u003e\n\u003cp\u003eSuccessful exploitation results in a denial of service for the vLLM instance. This impacts organizations relying on vLLM for high-throughput LLM serving, particularly those using disaggregated deployment architectures. Memory exhaustion renders the service unable to process legitimate requests, leading to potential operational outages in production machine learning environments.\u003c/p\u003e\n\u003ch2 id=\"recommendation\"\u003eRecommendation\u003c/h2\u003e\n\u003cp\u003ePrioritized actions for engineering and security teams:\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003eUpgrade the vLLM Mooncake connector to a version beyond 0.29.0 to patch CVE-2026-94627.\u003c/li\u003e\n\u003cli\u003eImplement request rate limiting and input validation at the API gateway layer to prevent the submission of excessively complex or concurrent prompts designed to trigger this memory exhaustion.\u003c/li\u003e\n\u003cli\u003eMonitor GPU memory usage metrics for abnormal, linear growth patterns that do not correlate with legitimate traffic volume.\u003c/li\u003e\n\u003c/ul\u003e\n","date_modified":"2026-09-21T22:31:23Z","date_published":"2026-09-21T22:31:23Z","id":"https://feed.craftedsignal.io/briefs/2026-09-vllm-mooncake-dos/","summary":"The vLLM Mooncake connector up to version 0.29.0 is susceptible to a denial-of-service attack where malicious completion requests exhaust GPU memory by orphaned KV cache blocks.","title":"Denial of Service in vLLM Mooncake Connector via KV Cache Exhaustion","url":"https://feed.craftedsignal.io/briefs/2026-09-vllm-mooncake-dos/"}],"language":"en","title":"CraftedSignal Threat Feed - Mooncake Connector (\u003c= 0.29.0)","version":"https://jsonfeed.org/version/1.1"}