What is Google Vertex AI

Google’s enterprise model-building and generative-AI service is now presented within the Gemini Enterprise Agent Platform. It provides Gemini and partner models, multimodal generation, grounding, tuning, evaluation, deployment, vector search, MLOps and provisioned throughput.

Generative AI pricing

As of July 30, 2026, standard global Gemini 3.1 Pro Preview usage up to 200,000 input tokens costs US$2 per million multimodal input tokens and US$12 per million text output tokens. Gemini 3.5 Flash costs US$1.50 input and US$9 output, while Gemini 3.5 Flash-Lite costs US$0.30 input and US$2.50 output. Cached input is one tenth of the standard input rate for those models.

Flex and batch processing are cheaper. Examples include Gemini 3.5 Flash at US$0.75 per million input tokens and US$4.50 output, and Flash-Lite at US$0.15 input and US$1.25 output. Regional, long-context, tuned-model, image, audio and partner-model rates differ.

Grounding and capacity

Grounding with Google Search or Maps includes 5,000 searches per month across Gemini 3 models, then costs US$14 per 1,000 searches. Grounding with your data costs US$2.50 per 1,000 prompts. Provisioned Throughput is sold in generative AI scale units from US$1,200 per GSU for one week to US$2,000 per GSU per month on a one-year commitment. New Google Cloud customers can receive US$300 in trial credits.

Official sources: Google generative AI pricing, Vertex AI platform pricing.