GCP Vertex AI High Volume Request Pattern Detection
High volumes of GenerateContent requests from a single IP against Vertex AI models indicate potential model extraction, automated scraping, or LLMjacking attacks.
Security teams should monitor for anomalous request volumes against Google Cloud Vertex AI models, specifically targeting the PredictionService.GenerateContent and PredictionService.StreamGenerateContent methods. Observed patterns of high request volume from a single source.ip within a short timeframe may signify model theft, systematic automated scraping, or LLMjacking, where an attacker leverages compromised credentials to exhaust prediction quotas or exfiltrate model knowledge. This behavior is documented by the MITRE ATLAS framework as Exfiltration via AI Inference API (AML.T0024) and Denial of AI Service (AML.T0029). Defenders should correlate these high-frequency events with associated service account activity, prompt content, and token volume to differentiate between legitimate batch processing and unauthorized interaction.
Attack Chain
- Attacker gains unauthorized access to a Google Cloud principal or service account with permission to query Vertex AI models.
- Attacker enumerates accessible model resources within the target GCP project.
- Attacker initiates high-frequency
GenerateContentorStreamGenerateContentcalls targeting a specific model resource. - Attacker iterates through systematically crafted prompts to probe model responses and extract output patterns.
- Attacker continues high-volume requests to maximize data exfiltration or deliberately exhausts the organization's prediction quota (LLMjacking).
- Attacker potentially modifies model configuration via
SetPublisherModelConfigto obscure future monitoring activity. - Attacker achieves final objective of either model weight reconstruction (theft) or resource denial through quota exhaustion.
Impact
Successful exploitation can lead to intellectual property theft through model extraction, unauthorized costs associated with LLMjacking, or service unavailability for legitimate users of the AI platform. These activities impact data confidentiality, billing integrity, and operational availability for organizations relying on GCP AI services.
Recommendation
- Enable GCP Vertex AI
auditlogsforaiplatform.googleapis.comto ensure visibility into prediction requests. - Establish a baseline for normal request volume per model and project to tune thresholds for the detection logic.
- Review billing and quota usage for Vertex AI models for sudden, unexplained spikes.
- If unauthorized activity is identified, rotate credentials for the calling principal and restrict the source IP via VPC Service Controls.
- Pivot to prompt logs to inspect for repetitive probing behavior or large context data dumps indicative of systematic scraping.
Immediate actions
Enable GCP Vertex AI audit logging for aiplatform.googleapis.com
Threat Hunt
Identify source IPs with event_count >= 100 for prediction methods
Data: GCP audit logs (logs-gcp_vertexai.auditlogs-*)