Vertex AI (Gemini API on Vertex) SLA
Google Cloud · AI models · promises up to 99.9% · SLA of Jun 30, 2026 · checked 24h ago
No major outage in the 208 days of history: 100% observed against the 99.9% Vertex AI Platform: Training, Deployment, Batch Prediction.
Last 90 days: no major outage.
Incidents, last 365 days
No incidents on file for this service.
Major outage time per month
Allowed per month: 44 min
Commitments and credits
TierUptimeCredit of the monthly bill
- Gemini Online Inference API (generateContent / streamGenerateContent)Gemini Online Inference API on Gemini Enterprise Agent Platform (95% for models designated for shorter availability)99.5%below 99.5%: 10% · below 99%: 25% · below 95%: 50%
- Provisioned Throughput latency (Gemini 2.5 Pro/Flash/Flash-lite, global endpoint)Monthly Latency Target Attainment (p50 TPS vs. target: 60/80/110 TPS)99%below 99%: 10% · below 95%: 25% · below 90%: 50%
- Vertex AI Platform: Training, Deployment, Batch PredictionVertex AI Platform SLA (https://cloud.google.com/vertex-ai/sla, last modified February 12, 2026)99.9%below 99.9%: 10% · below 99%: 25% · below 95%: 50%
- Vertex AI Platform: Custom Model Online Prediction (2+ nodes), PipelinesVertex AI Platform SLA99.5%below 99.5%: 10% · below 99%: 25% · below 95%: 50%
More in ai models
ServicePromisedFirst creditObserved 365dDowntime
Amazon BedrockAWS · AI models · Standard99.9%First credit10%Observed 365d99.178%All incidents 99.852%
Claude API (Standard tier)Anthropic · AI modelsNoneFirst creditn/aObserved 365d98.76%All incidents 94.774%
Claude API Priority TierAnthropic · AI models · Priority Tier (existing commitments only)99.5%First creditn/aObserved 365d98.8%All incidents 94.774%
Claude Enterprise (claude.ai)Anthropic · AI modelsNoneFirst creditn/aObserved 365d98.62%All incidents 94.774%
Cohere API (SaaS: Command, Embed, Rerank)Cohere · AI modelsNoneFirst creditn/aObserved 365d100%All incidents 99.986%
Cohere private deployments / Model Vault / NorthCohere · AI modelsNoneFirst creditn/aObserved 365dn/aAll incidents 99.986%
Hugging Face Inference EndpointsHugging Face · AI modelsNoneFirst creditn/aObserved 365dNo feed
Azure OpenAI in Foundry ModelsMicrosoft Azure · AI models · Foundry Models (Sold by Azure), availability99.9%First credit10%Observed 365dNo feedAll incidents 100%
Terms, exclusions and sources
- Measured
- Calendar month per Project; Downtime = more than 5% server-side 5xx Error Rate over 5+ consecutive minutes
- Credit cap
- 50% of the amount due for the Covered Service for the month
- Claim window
- Within 30 days from the time Customer becomes eligible
- How to claim
- Notify Google technical support; provide log files/identifying info showing Downtime Periods and when they occurred
- Not covered
- Features designated pre-general availability; Features excluded from the SLA in the Documentation; Errors caused by factors outside Google's reasonable control; Errors from Customer's or third-party software or hardware; Requests using Grounding with Google Search; Deadlines set shorter than the server default (deadline_exceeded)
- Status history
- Status history since Feb 27, 2026
- Weighted uptime
- All incidents, weighted: 99.488% · Downtime
Training, Deployment, and Batch Prediction >= 99.9% ... 99% to < 99.9% 10% 95% to < 99% 25% < 95% 50%
Product renamed on the SLA page to 'Gemini Online Inference API on Gemini Enterprise Agent Platform'. The latency tier's uptime field holds the 99% latency attainment, not availability.
Summaries of published SLAs; the contract you sign governs. Logos via logo.dev; trademarks belong to their owners.