GCP News - 2026-08-21

2026-08-21
最終更新: 2026-08-27 21:31:31 JST

Google Cloud Release Notes

August 21, 2026

詳細を表示

AI Hypercomputer

Change

AI Hypercomputer has expanded to cover all GPU machine series. Specifically, the documentation for AI Hypercomputer has been significantly updated as follows:

  • Documentation has been added for the A3 Edge, A2, G4, G2, N1+T4, and N1+V100 GPU machine series.
  • Pages are labeled based on whether they are relevant to the newly documented machine series for general GPU, the pre-existing machine series for clustered GPU, or both.

To learn more about the differences between general GPUs and clustered GPUs, see Choose your accelerator infrastructure. To view workload recommendations, see Choose between general GPUs and clustered GPUs.

AlloyDB for PostgreSQL

Fixed

AlloyDB now provides more accurate memory usage estimation, and it prevents out-of-memory (OOM) errors when you build a ScaNN four-level tree index. This feature is in Preview. This update improves the stability of index builds by enforcing memory limits and improving memory estimation under constrained conditions.

For more information, see Create a ScaNN index.

Application Integration

Security

Missing authorization in QueryEngineTask (CVE-2026-12710)

A missing authorization vulnerability (CVE-2026-12710) in QueryEngineTask in Application Integration was patched on April 4, 2026, and no customer action is needed.

Assured Workloads

Feature

The Switzerland Data Boundary with Access Justifications control package is now generally available.

Feature

The Data Boundary for ITAR supports the following products:

  • Apigee
  • Artifact Analysis
  • Backup and DR Service
  • Cloud Billing API
  • Cloud Deploy
  • Spanner
  • Eventarc
  • GKE Image streaming
  • Google Cloud Managed Service for Apache Kafka
  • Vertex AI Online prediction

Cloud Storage

Feature

Rapid Cache supports prefix-level ingest-on-write filtering, allowing you to selectively ingest objects that match specific prefixes rather than ingesting all objects written to a bucket.

For more information, see Ingesting data on write.

Compute Engine

Feature

Generally available: The network- and memory-optimized M4N machine series is generally available. Powered by 5th generation Intel Xeon Scalable processors (Emerald Rapids), M4N instances are purpose-built for network and block storage-intensive workloads such as:

  • High-performance vector databases
  • Retrieval-augmented generation (RAG) data layers
  • Massive in-memory context caching
  • Real-time semantic search

The M4N machine series delivers the highest I/O performance available in Compute Engine, supporting up to 400 Gbps of network bandwidth. M4N also offers leading block storage performance with Hyperdisk Extreme that scales up to 25 GiB/s of bandwidth and 1M IOPS. M4N instances are available in predefined machine shapes, ranging in size from 16 to 224 vCPUs and up to 5,952 GB of DDR5 memory.

Gemini Enterprise

Feature

Gemini Enterprise: AI developer tools on Standard Emerging Market edition

The AI developer tools feature is available on the Gemini Enterprise Standard Emerging Market edition. Your project must be linked to an invoiced Cloud Billing account that receives an active monthly invoice to access the feature. This edition doesn't include bundled base quota or Antigravity credits.

For more information, see the following:

Feature

Gemini Enterprise: New data stores and support for new actions (Public Preview)

The following data stores are available in Public Preview in Gemini Enterprise:

You can search and read data from these data stores using natural language.

Additionally, the following data stores support new actions in Public Preview:

  • Descript: Import media, prompt project agent, and publish project.
  • Gamma: Generate Gamma.
  • Supabase: Apply migration, deploy edge function, execute SQL, pause project, and restore project.

Gemini Enterprise Agent Platform

Feature

xAI's Grok 4.6

Grok 4.6 is available in Preview in Model Garden.

Google Cloud Contact Center as a Service

Announcement

Google Cloud CCaaS 6.4

We've released version 6.4 of Google Cloud CCaaS.

The timing of the update to your instance depends on the deployment schedule that you have chosen. For more information, see Deployment schedules.

Feature

Automatic SIP parameter mapping in contact lists

Contact lists now support automatic SIP header mapping for outbound call destinations.

Administrators: In the Add Destination dialog at Settings > Call > Contact Lists > Contact list management > CREATE OR EDIT CONTACT LIST> Add Destination, toggle Pass Data Parameters to the on position to see the new Automatically Include and Pass all Inbound SIP Headers checkbox.

For more information, see Add a SIP URI address destination to a contact list.

Feature

Email OAuth profiles and Microsoft 365 client credentials

You can now create email OAuth profiles for Microsoft 365 using the client credentials grant type. This lets shared mailboxes authenticate through an app registration instead of relying on an individual user sign-in. OAuth IMAP setup is more resilient for support inboxes and other application-managed mailboxes.

For more information, see Email OAuth profiles and Microsoft 365 client credentials.

Feature

Improved connection times for predictive dialing

We reduced the connection time for predictive dialing to under two seconds to comply with CRTC and FTC regulations.

Fixed

This release addresses the following issues:

  • Fixed an issue where a long delay occurred when an agent clicked Chat Shortcuts in the chat adapter.

  • Fixed an agent desktop issue where the Session Data Feed pane displayed data from the previous chat and the chat transcript stopped updating.

  • Fixed an issue where conversation history loaded slowly for agents using the CCaaS widget embedded in Salesforce.

  • Fixed an issue where outbound calls with zero duration time were missing from Individual Call History reports that were scoped to agents and teams.

  • Fixed an issue where customers couldn't leave a voicemail during a warm transfer.

  • Fixed an issue where the Queued Calls dashboard displayed the incorrect originating queue and transferring agent for a cold-transferred call.

  • Fixed an issue where the agent desktop experienced significant loading delays.

  • Fixed an issue where calls transferred from a virtual agent directly to a human agent bypassed the Keep Waiting overcapacity setting, causing callers to be incorrectly routed to a fallback queue or to voicemail.

  • Fixed a raw data export issue where file names for the MENU_PATH_ITEMS dataset incorrectly contained path_items.

  • Fixed an issue where a parent queue was incorrectly marked as after-hours even though one or more of its child queues were in operation.

  • Fixed an issue with chats that were escalated from a virtual agent to a human agent and then reached a terminal status (canceled, finished, or failed). These chats mistakenly appeared in the Queued Chats dashboard.

  • Fixed an issue where the agent desktop froze and the Call Adapter pane displayed Call on hold after a call was terminated or dropped.

  • Fixed an issue where agents couldn't change their chat status while a call was in the wrap-up stage.

  • Fixed an issue where the call adapter displayed calls waiting when there were actually no calls waiting.

  • Fixed an issue where outbound calls that failed immediately caused the call adapter to be stuck in the In-call state with the timer running.

Google Kubernetes Engine

Change

Per the June 10, 2026 release note, the configuration option to not enroll your cluster in a release channel is deprecated, and will be removed on June 14, 2027. In alignment with this deprecation, creating new clusters not enrolled in a release channel is now only allowed for existing customers. New customers can use a release channel, where you can achieve the same functionality as not enrolling your cluster in a release channel. For more information, see Clusters not enrolled in a release channel.

Change

The Windows Server 2019 (LTSC) GKE node image doesn't receive updates after the December 2025 version. Windows Server 2019 (LTSC) is in the Extended Support period of the Microsoft fixed lifecycle policy and receives only security updates. To prevent stability issues, the GKE node image for Windows Server 2019 (LTSC) is pinned to the December 2025 version. If you use this node image, switch to Windows Server 2022 (LTSC), which is in the Mainstream Support period and receives updates from Microsoft and GKE. For more information, see Creating a cluster using Windows Server node pools.

Google SecOps

Feature

[Spotlight Feature] Relative time filtering in Google SecOps

This feature is in public preview. Google SecOps has updated how relative time filters calculate data ranges. You can now choose from three distinct, mathematically precise operators: Past, Previous, and Current. This change eliminates ambiguity between rolling windows and calendar-aligned periods, ensuring consistent behavior across all time units (like seconds, minutes, hours, days, weeks, months, years) and aligning SecOps dashboards with Search and other Google tools (such as Looker).

For more information, see the Relative time range section of the Understand search guide.

Service Health

Feature

The Personalized Service Health remote MCP server provides a secure environment that lets you send natural language prompts to your AI application so that it can retrieve incident information, audit incidents, and automate debugging on your behalf.

For more information, see Use the Personal Service Health remote MCP server.

Google Cloud Blog

Google is a Leader in the 2026 Gartner® Magic Quadrant™ for Cloud-Native Application Platforms

詳細を表示

We are thrilled to announce that Google has been recognized as a Leader for the third year in a row in the 2026 Gartner® Magic Quadrant™ for Cloud-Native Application Platforms (CNAP). We believe this placement in the Leaders quadrant validates our commitment to providing an accessible, developer-centric platform that accelerates onboarding and supports rapid prototyping across modern workloads.

<div class="article-module h-c-page">
  <div class="h-c-grid">


<figure class="article-image--large
  
  
    h-c-grid__col
    h-c-grid__col--6 h-c-grid__col--offset-3
    
    
  ">

  
  
    
    <img alt="Figure_1_Magic_Quadrant_for_CloudNative_Application_Platforms" src="https://storage.googleapis.com/gweb-cloudblog-publish/images/Figure_1_Magic_Quadrant_for_CloudNative_Ap.max-1000x1000.png" />
    
    </a>
  
</figure>


  </div>
</div>

Our vision for an application-centric cloud focuses on enabling developers to prioritize writing code and building agents or traditional apps by removing infrastructure complexity. Google Cloud provides a unified execution environment supporting serverless, containerized, and agentic deployment options. We believe our placement highlights Google's unique readiness to power both standard enterprise microservices and the next generation of autonomous AI applications.  

Some key features and capabilities of our platform are highlighted below. 

From idea to implementation

Generative AI has ushered in a wave of vibe coding, allowing anyone to go from an idea to a deployed application in a fraction of the time it used to take. To make this even smoother, Google Cloud integrates its serverless infrastructure with AI vibe-coding and prototyping tools. We also simplify access to Google Cloud resources with tools like managed MCP servers and agentic skills — packaged sets of instructions, scripts, and resources to teach an AI how to complete specialized, multi-step workflow. 

  • One-click prototyping in Google AI Studio: Developers can build and deploy full-stack applications directly within Google AI Studio, making it a great environment for prototyping and experimentation. With a single click, you can instantly package and publish your vibe-coded applications to Cloud Run

  • Google-managed MCP servers: To enable AI agents to interact with cloud resources, we support official, fully managed remote MCP servers. An example is the Cloud Run MCP server (via run.googleapis.com/mcp), which allows developers to easily launch endpoints and deploy server-side logic. The MCP tools are deployed with a simple config, skipping cloud builds to launch code in seconds, saving valuable developer time. These fully managed servers are integrated with IAM and VPC Service Controls, and they leverage Model Armor for content security.

  • Google's official Skills Repository: Level up your agents with additional, condensed expertise on various Google Cloud technologies. Published and available in Agent Registry, the repository includes skills for Cloud Run, the Well-Architected Pillar (security, reliability, and cost optimization), and more.

Ready to try vibe coding yourself? Get hands on with this codelab to build a vibe-coded app and deploy it to Cloud Run.

From implementation to enterprise-ready

Translating prototypes into production-grade, secure, and cost-effective enterprise software is where Google Cloud excels, with a full suite of developer, architect, and platform engineering tools. From designing your application to optimizing day 2 operations, we offer the services and tools to help you build, operate, and deploy applications across their entire lifecycle, and you have the freedom to build with any language, any library, and any framework.

Build

  • Build with Google Antigravity: At Google, we’re simplifying and expanding our development ecosystem behind the Antigravity harness, collapsing developer silos into a unified orchestration layer. By integrating multi-step AI reasoning directly into the developer workflow, Antigravity natively brings local codebase development to our cloud-native application platforms (e.g., Cloud Run).

  • Design and deploy with Application Design Center (ADC): Now, you can bridge the gap between developer velocity and enterprise control, using Application Design Center to eliminate manual Terraform and YAML configuration. This platform engineering component helps teams design, standardize, and deploy template-driven applications on Google Cloud. It is also integrated as part of Gemini Cloud Assist design agent and published as an MCP server. With ADC, you can visually design your architecture using Cloud Run services, databases, and event brokers backed by automated Gemini Cloud Assist security templates. Beyond human-guided design, ADC enables programmatic orchestration at the time of no HITL (Human-in-the-Loop), allowing automated pipelines to provision policy-governed Terraform configurations directly and autonomously.

Operate 

  • Intelligent investigations: Integrating Gemini Cloud Assist with native telemetry creates an AI-driven framework for Day-2 incidents. When alerts fire, operators engage Gemini Cloud Assist to instantly synthesize logs and metrics, pinpoint root causes, and generate remediations — context that can be handed off to accelerate support escalations. Crucially, IAM permissions strictly govern all AI recommendations, and help to ensure explicit human-in-the-loop approval are required before any infrastructure changes occur.

  • Cost analysis and optimizations: Machine learning algorithms learn natural seasonal traffic cycles to detect cost anomalies within minutes, triggering notifications to protect your bottom line without risking destructive infrastructure shutdown.  

Deploy

  • Reliability and high availability: Cloud Run is a regional service by default, but you can deploy an app to multiple regions via a single gcloud command. Integrated with service health, Cloud Run automates cross-region failover and failback. If a service in one region becomes unhealthy, traffic is automatically routed to the next-closest healthy region, failing back once the issue is resolved.

  • An open platform: As a long-time and top contributor to the Cloud Native Computing Foundation (CNCF), we operate with an open-source-first strategy. By integrating foundational, community-driven technologies, we help enable application portability for enterprise customers who are increasingly demanding multi-cloud flexibility

Ready to start deploying your apps to Google Cloud? Get hands on with these codelabs:  

From enterprise-ready to autonomous

AI agents are software’s next frontier. They offer more than just increased productivity and efficiency; they can unlock exponential growth. To provide enterprises with robust agentic deployment options, Google Cloud provides a dedicated infrastructure stack tailored specifically to host, govern, and secure autonomous agent fleets. This stack seamlessly integrates with our Agent Development Kit (ADK) as well as other leading agentic frameworks to give developers maximum flexibility.

Gemini Enterprise Agent Runtime
At the core of this stack is Gemini Enterprise Agent Platform and its dedicated Agent Runtime, which delivers the serverless and containerized deployment options you need for enterprise-scale agent development, including the following capabilities:

  • Native personalization: Built-in sessions and memory banks manage context and long-term state, preventing costs from ballooning.

  • Agent observability and tracing: Built on OpenTelemetry (OTel) standards and agentic schemas, turnkey dashboards feature agent topology graphs and interactive trace logs that detail sessions, tool calls, and reasoning paths.

  • Agent evaluation and simulation: Automated simulation tools allow developers to test agents against golden sets with side-by-side comparisons and simulate thousands of interactions to test edge cases.

Hosting agents on Cloud Run
For customers requiring additional flexibility, granular control, or specific regulatory compliance, Cloud Run serves as an excellent serverless alternative to host your agents. Some of its latest features include:

  • Cloud Run instances (coming soon): This primitive manages individual, addressable, long-running singleton resources with integrated Cloud Storage volume mounts, allowing persistent background agents to be deployed cost-effectively.

  • Cloud Run sandboxes: Hard-isolated environments spin up in under 500 milliseconds to safely execute untrusted, model-generated code, protecting the host system from unauthorized access.

Agent security, governance, and auditability
To securely deploy AI agents and prevent unmanaged shadow AI, enterprises need an ironclad governance framework. Google Cloud delivers this through Agent Identity (non-human IAM with cryptographic IDs) to provide an auditable trail of all actions and reasoning; a centralized Agent Registry to manage approved agents, skills, tools and application artifacts,  and prevent unauthorized tool integrations; and an Agent Gateway to proxy traffic, enforce Model Armor policies, and actively block destructive actions. These features are available on Agent Runtime today and will be available soon on Cloud Run and Google Kubernetes Engine (GKE).

Ready to start deploying agents? Check out various codelabs featuring Gemini Enterprise Agent Platform here.

Build the future of cloud-native applications

Whether you’re a vibe coder deploying your first full-stack application, a software architect standardizing production microservices, or an enterprise team scaling a fleet of secure AI agents, Google Cloud delivers the simplicity, elasticity, and security you need. Read the full report: Download your complimentary copy of the 2026 Gartner® Magic Quadrant™ for Cloud-Native Application Platforms (CNAP).


Magic Quadrant for Cloud-Native Application Platforms, By Mukul Saha, Alex Coqueiro, Prasanna Lakshmi Narasimha, Richard Watson, 3 August 2026

Gartner and Magic Quadrant are trademarks of Gartner, Inc. and/or its affiliates. This graphic was published by Gartner, Inc. as part of a larger research document and should be evaluated in the context of the entire document. The Gartner document is available upon request from Google. 

Gartner does not endorse any company, vendor, product or service depicted in its publications, and does not advise technology users to select only those vendors with the highest ratings or other designation. Gartner publications consist of the opinions of Gartner’s business and technology insights organization and should not be construed as statements of fact. Gartner disclaims all warranties, expressed or implied, with respect to this publication, including any warranties of merchantability or fitness for a particular purpose.

How AlloyDB ScaNN scales vector search to 10 billion vectors

詳細を表示

To satisfy the demands of enterprise-grade agentic AI applications, underlying vector databases often struggle to scale effectively as modern use cases can scale to billions of vectors.

As a fully managed PostgreSQL-compatible database service, AlloyDB is engineered to handle demanding enterprise workloads. Combining Google's infrastructure with the reliability of commercial databases, it delivers high availability, scalability, and includes a cutting-edge analytical engine, optimal for agentic AI use cases. A key part of this is its ScaNN index, which now operates efficiently at a scale of 10 billion vectors. This was achieved through a major architectural enhancement: an innovative four-level tree (preview) paired with efficient memory usage.

The 10 billion vector scale challenge

Scaling to a 10 billion vector workload presents significant memory and computational challenges. Previous AlloyDB ScaNN tree-based index was limited to two- or three-level tree configurations, and attempting to scale those structures led to several bottlenecks:

  • Increased compute intensity: Larger tree structures demand significantly more operations for both index construction and query traversal.

  • Memory constraints: The sampling processes required for 10 billion vectors can easily exceed the system's available memory capacity.

Solution: Four-level architecture

The introduction of a four-level tree (preview) is the primary innovation in the recent AlloyDB ScaNN release. This architecture, illustrated in Figure 1, employs a top-down strategy to optimize the balance between accuracy and build efficiency. To maintain high performance and mitigate recall loss, the system integrates key enhancements such as Top-K branch, SOAR, centroid adjustment and balanced tree shape.

<div class="article-module h-c-page">
  <div class="h-c-grid">


<figure class="article-image--large
  
  
    h-c-grid__col
    h-c-grid__col--6 h-c-grid__col--offset-3
    
    
  ">

  
  
    
    <img alt="1" src="https://storage.googleapis.com/gweb-cloudblog-publish/images/1_fpfICUj.max-1000x1000.jpg" />
    
    </a>
  
    <figcaption class="article-image__caption "><p>Figure 1. AlloyDB ScaNN four-level tree architecture</p></figcaption>
  
</figure>


  </div>
</div>

This design has two primary benefits:

1. Reduced compute intensity via hierarchical partitioning

The four-level architecture drastically reduces compute intensity by using hierarchical partitioning to restrict the volume of vectors scanned during a query. Instead of traversing a flat or poorly segmented space, the multi-layered hierarchy narrows down the search path exponentially. Figure 2 illustrates the search spaces across different tree levels, demonstrating how structural layering optimizes traversal efficiency:

<div class="article-module h-c-page">
  <div class="h-c-grid">


<figure class="article-image--large
  
  
    h-c-grid__col
    h-c-grid__col--6 h-c-grid__col--offset-3
    
    
  ">

  
  
    
    <img alt="2" src="https://storage.googleapis.com/gweb-cloudblog-publish/images/2_LWwXC70.max-1000x1000.jpg" />
    
    </a>
  
    <figcaption class="article-image__caption "><p>Figure 2. Search space for two-, three- and four-level trees</p></figcaption>
  
</figure>


  </div>
</div>
  • Two-level: Utilizes coarse partitioning to guide queries, resulting in a basic search complexity of O(N1/2).

  • Three-level: Introduces an intermediate layer to further subdivide clusters, narrowing exploration to O(N1/3).

  • Four-level: Implements refined, highly granular partitions that optimize traversal efficiency down to O(N1/4), sufficiently allowing for more than 10-billion vectors.

By dynamically expanding hierarchical layers as the dataset expands, AlloyDB ScaNN maintains ultra-low query latency and avoids computational scale walls from impacting performance.

2. Efficient memory usage

Achieving a 10 billion vector scale requires high memory efficiency. AlloyDB ScaNN uses these strategies to maximize memory management performance:

  • Balanced tree shape construction: The four-level tree utilizes a balanced configuration to circumvent memory limitations that restrict the size of training datasets. This balanced architecture effectively leverages reduced sampling sizes to construct high-fidelity tree partitions.

  • Sampling optimization: When the system encounters memory limitations, it generates a condensed sampling set that considers performance and accuracy. 

Performance test results

By leveraging the innovative four-level tree architecture in our internal tests, we are able to achieve the following performance results:

  • AlloyDB can scale to over 10 billion vectors with its ScaNN index.

  • AlloyDB can deliver <= 51 ms p95 latency and 95% recall at 10 billion vectors with its ScaNN index.

Get started today

Experience AlloyDB ScaNN's four-level tree (preview) architecture today. You can deploy ScaNN for AlloyDB by following our quickstart guide to set up an instance. For optimized, high-speed vector search, refer to the official ScaNN documentation. New users can also explore AlloyDB through our 30-day free trial program. We can’t wait to hear about what you build!

10 questions every startup should answer before moving to production with their AI prototype

詳細を表示

It’s never been easier to start an AI-powered startup on Google Cloud. 

You grab an API key from Google AI Studio at breakfast, paste it into Antigravity, and by lunch you’ll have a nascent prototype of your product.

But it’s not all one straight line to progress. It's common to bump into these three challenges as you build out your stack:

  • A leaked API key racks up a large bill in 48 hours.

  • A "quick" migration from AI Studio to Gemini Enterprise Agent Platform stalls the roadmap for weeks because nobody on the team owns Identity and Access Management (IAM).

  • The launch works, until the app starts returning HTTP 429 Too Many Requests because of default per-project quotas, and there's no clean path to more capacity without paying a premium.

None of these are unique edge cases. . They're  default failure modes of moving fast without a plan, and we've all done it at least once.

Below are the 10 questions every startup should be ready to answer before they scale,  grouped into the three phases where decisions can shape your future: 

  1. Onboard (setting up your own projects and identities right)

  2. Scale (getting more throughput without breaking the bank) 

  3. Govern (keeping costs, keys, and agents from running away).

These ten are scoped to the prototype-to-production transition itself. Each question ends with a short, runnable snippet you can copy into your own project today. Adjacent decisions that matter just as much but aren't specific to that move, your data layer and RAG architecture, CI/CD, network design, are deliberately out of frame here.

Onboard: get the foundation right (in the first hour).

#1 Where should I start: Google AI Studio or Gemini Enterprise Agent Platform?

Both surfaces expose the same Gemini family of models, but they solve different problems.

  • Google AI Studio (with the Gemini Developer API) is the fastest path from an idea to working code. A browser IDE, an API key, a generous free tier, and no cloud project to configure. It's where most ideas should start, and Google's own guidance says as much.

  • Gemini Enterprise Agent Platform (formerly Vertex AI) has the same Gemini models (plus 3rd party and OSS ones)  with enterprise controls around them: IAM and service-account auth instead of raw keys, VPC Service Controls, Cloud Logging and Monitoring, reserved capacity, regional endpoints, and the compliance surface your first enterprise customer's security review will ask about.

The right answer for most startups is both, sequenced deliberately: first prototype in AI Studio, then migrate before you have real users. The danger for startups is treating them as interchangeable solutions, AI Studio's simple key model does not translate to enterprise controls, and Agent Platform's IAM model might look like overkill until the day it saves you from a stolen-credential incident.

It's less work than it sounds like.

The unified google-genai SDK targets both:

code_block
<ListValue: [StructValue([('code', '# Prototype: Google AI Studio, raw API key\r\nfrom google import genai\r\nclient = genai.Client(api_key="YOUR_AI_STUDIO_KEY")\r\n\r\n# Production: GEAP, no key — uses Application Default Credentials (ADC)\r\nfrom google import genai\r\nclient = genai.Client(\r\n vertexai=True,\r\n project="my-startup-prod",\r\n location="us-central1",\r\n)\r\n\r\nresp = client.models.generate_content(\r\n model="gemini-2.5-pro",\r\n contents="Summarize this contract in three bullets.",\r\n)\r\nprint(resp.text)'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7f5933ee42d0>)])]>

#2  How do I set up a Google Cloud project without becoming an IAM expert?

The biggest reason startups stall on the migration to Agent Platform isn't the code, it's the operational leap from "here's an API key" to a cloud project with folders, service accounts, org policies, logging, and IAM bindings. If your team doesn't have a dedicated cloud admin, that first project setup can eat a week of engineering time. 

Three moves cut that dramatically:

  1. Use an opinionated project template instead of clicking through the console. The Cloud Setup checklist and the Google Cloud Architecture Framework give you a production-grade folder hierarchy (prod / non-prod / dev), a central logging + monitoring project, Security Command Center turned on, and baseline org policies, without you having to design them from scratch.

  2. Enable the APIs you'll actually use, once. Batch it so you're not doing it project-by-project when you need it. The billing-link step is not optional. Every paid API you're about to enable will refuse to activate on a project with no billing account attached, so we handle that first.

  3. Let Gemini pick the roles, but ask it for the narrow ones. You don't have to memorize the roles reference. In the Grant access dialog, Help me choose roles lets you describe the task in plain language, "this service account needs to call Gemini models and read one Cloud Storage bucket", and get predefined roles back with the reasoning shown. One catch worth knowing on day one: by default it suggests roles that cover common journeys, which usually means a service's Admin, Editor, or Viewer. Those are broader than you want. Say "least privileged" or "narrowest access" in the prompt and it returns granular roles instead. Same amount of typing, considerably smaller blast radius when a credential leaks.

    Sources: Get predefined role suggestions with Gemini assistance

code_block
<ListValue: [StructValue([('code', '# One-shot: create a Vertex-ready project and turn on the services a\r\n# typical AI startup uses.\r\ngcloud projects create my-startup-prod --name="My Startup (prod)"\r\ngcloud config set project my-startup-prod\r\n\r\n# REQUIRED before enabling billing-dependent APIs (aiplatform, run, etc.).\r\n# Use `gcloud billing accounts list` to find your billing account ID.\r\ngcloud billing projects link my-startup-prod --billing-account=012345-6789AB-CDEF01\r\n\r\ngcloud services enable \\\r\n aiplatform.googleapis.com \\\r\n run.googleapis.com \\\r\n artifactregistry.googleapis.com \\\r\n logging.googleapis.com \\\r\n monitoring.googleapis.com \\\r\n secretmanager.googleapis.com \\\r\n cloudbilling.googleapis.com'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7f593025d4d0>)])]>

Sources: gcloud services enable reference, · gcloud billing projects link (GA)GE Agent Platform environment setup.

If you're a solo founder, resist the urge to build in your personal GCP account. Create a proper organization or self-owned org first, then create the project inside it. That single decision can make everything else, fromIAM to billing and audit, dramatically easier.

#3 I'm on Google Cloud, how should my code actually authenticate: API keys, service accounts, or user credentials?

There's a hierarchy of safety here, and the easiest option is rarely the right one in production.

  • Raw API keys are fine for local prototyping. They are dangerous in production because they are long-lived, easy to leak into a client bundle or a public repo, and grant unbounded access until you notice.

  • User credentials via OAuth (application default credentials) are best for interactive tools, CLIs, and any code that runs on a developer's laptop.

  • Service accounts with least-privilege IAM roles are the right answer for anything running on a server, in a container, or in a scheduled job.

The pattern you're aiming for is one where your code never sees a key at all. It just calls the Google Auth library, which quietly reads Application Default Credentials (ADC) from the environment,  a short-lived token minted for whichever service account is attached to your Cloud Run service, GKE workload, or Compute Engine VM. You get enterprise-grade auth without writing any auth code.

code_block
<ListValue: [StructValue([('code', '# On a developer laptop\r\ngcloud auth application-default login\r\n\r\n# On a server (Cloud Run, GKE, etc.) — no login, no key file.\r\n# Attach a service account with just the roles the app needs.\r\ngcloud run deploy my-agent \\\r\n --image=us-docker.pkg.dev/my-startup-prod/agents/api:v1 \\\r\n --service-account=agent-runtime@my-startup-prod.iam.gserviceaccount.com \\\r\n --region=us-central1'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7f593025cf50>)])]>
code_block
<ListValue: [StructValue([('code', '# Application code — notice: no keys, no secrets.\r\nfrom google import genai\r\n\r\nclient = genai.Client(\r\n vertexai=True,\r\n project="my-startup-prod",\r\n location="us-central1",\r\n)'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7f593025f990>)])]>

Do one last favor to your future self: give that service account the minimum IAM role your workload actually needs,  usually roles/aiplatform.user for calling models, not the broader admin roles. It takes an extra 30 seconds and prevents the credential from becoming a master key if it leaks.

#4 When should I actually stop procrastinating and migrate from AI Studio's API key to Agent Platform's IAM model?

Sooner than you'd like,  and the correct trigger is not when it breaks. It's when any of these is true:

  • Your key has left your laptop (checked into a repo, pasted into a Slack, shipped in a mobile app).

  • You have more than one person on the team who needs to call the API.

  • You're spending more than a few hundred dollars a month.

  • You're about to onboard paying customers.

A potential pitfall that can catch growing startups off guard is simple: a leaked Gemini API key on an account that normally spends $180 a month gets scraped from a public repo and used to run distillation attacks,  accumulating tens of thousands of dollars in charges before the owner even sees the first billing alert. The Google Cloud Shared Responsibility Model is unambiguous: the customer is liable for charges incurred with their own valid credentials.

The migration itself is genuinely smaller than the anxiety around it. In google-genai it's the two-line change shown in #1. What takes real time is the project setup around it, which is exactly why #2 exists.

Practical checklist for cutover day:

code_block
<ListValue: [StructValue([('code', '# 1. Revoke every existing AI Studio key that has ever left a laptop.\r\n# (Go to https://aistudio.google.com/apikey and delete them.)\r\n\r\n# 2. Confirm your production code has no api_key= arguments.\r\ngrep -rn "api_key" src/\r\n\r\n# 3. Enable GEAP and confirm ADC works locally.\r\ngcloud services enable aiplatform.googleapis.com\r\ngcloud auth application-default login\r\npython -c "\r\nfrom google import genai\r\nc = genai.Client(vertexai=True, project=\'my-startup-prod\', location=\'us-central1\')\r\nprint(c.models.generate_content(model=\'gemini-2.5-flash\', contents=\'ping\').text)\r\n"'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7f593025cb10>)])]>

If step 3 prints a response, you're on Agent Platform.

Scale: get more capacity without paying a premium.

#5 Now that I'm shipping, why on earth am I getting all these HTTP 429 errors, and how do I make them stop?

429 Too Many Requests from Agent Platform almost always means one of two things:

  1. You've hit the Dynamic Shared Quota (DSQ) ceiling for your project's tier. DSQ is a shared pool sized against your project's history,  new projects start with modest limits by design, to prevent abuse across the platform.

  2. You're calling a global endpoint during a global demand spike, competing with worldwide traffic for shared capacity.

The instinctive reaction is to file a quota-increase ticket. You can do that if you must,  but two architectural moves usually solve the problem faster and cheaper.

Pin to a regional endpoint. Over half of startup traffic on Agent Platform defaults to global routing. Pinning to a specific region (say us-central1) sidesteps global contention and typically improves latency at the same time. (One narrow exception, which we'll get to in the next question: if you specifically want Priority PayGo, that feature currently only ships on the `global` endpoint. For everything else, pin regionally.):

code_block
<ListValue: [StructValue([('code', 'from google import genai\r\n\r\n# Global (default): competes against worldwide demand.\r\n# Regional: routes only to the regional cluster, less contention.\r\nclient = genai.Client(\r\n vertexai=True,\r\n project="my-startup-prod",\r\n location="us-central1", # <-- this is the one-line fix\r\n)'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7f593025db90>)])]>

Add real retry and backoff. A 429 is a retryable signal, not a fatal error. Any production client should have exponential backoff with jitter. The modern google-genai SDK ships this behavior built in, but only if you actually enable it. This is easy to overlook. Don't reach for the classic `google.api_core.retry.if_transient_error` decorator you may have seen on older Vertex code. It's designed for the legacy exception classes and does not recognize the new `google.genai.errors.APIError,  so it will silently pass 429s through without retrying. Use the SDK's built-in retry options instead:

code_block
<ListValue: [StructValue([('code', 'from google import genai\r\nfrom google.genai import types\r\n\r\nclient = genai.Client(\r\n vertexai=True, project="my-startup-prod", location="us-central1",\r\n http_options=types.HttpOptions(retry_options=types.HttpRetryOptions(\r\n attempts=5, initial_delay=1.0, max_delay=60.0, exp_base=2.0, jitter=1.0,\r\n http_status_codes=[408, 429, 500, 502, 503, 504],\r\n ))\r\n)'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7f593025f790>)])]>

How do you see this coming?  Preferably not from a user telling you. Agent Platform publishes serving metrics to Cloud Monitoring, and there is a prebuilt dashboard you don't have to assemble: Console → Agent Platform → Dashboard → Model observability. It gives you requests per second, token throughput, first-token latency, and error rates out of the box.

The metric to actually alert on is aiplatform.googleapis.com/publisher/online_serving/model_invocation_count. It carries an error_category label with values of user, system, or capacity. Alerting on capacity isolates genuine throttling from your own bad requests, which a raw 429 count won't do.

One thing worth internalizing, because it trips people up: you cannot build a "warn me at 80% of my quota" alert for Standard PayGo. Under Dynamic Shared Quota there is no fixed per-project number to be at 80% of. A 429 means transient contention for shared capacity, not that you crossed a line. Percent-of-limit alerting only becomes meaningful once you're on Provisioned Throughput, which does expose real limit metrics.

code_block
<ListValue: [StructValue([('code', 'gcloud monitoring policies create --policy-from-file=capacity-alert.yaml'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7f593025c4d0>)])]>

Sources: Agent Platform metrics list, Model observability dashboard, RetryOptions sourcecore retry_base.py, genai errors.pyreduce 429 errors, gcloud monitoring policies create, Dynamic Shared Quota.

Follow the Agent Platform rate limits documentation to understand what your project's current ceiling actually is before you assume you've outgrown it. 

#6 Which consumption mode do I pay for: Standard PayGo, Priority PayGo, or Provisioned Throughput? 

Three consumption models, three completely different workload shapes, and three completely different ways to proceed. Picking the right one can help startups see meaningful savings on AI bills. First let’s define them and then see when they are, or aren’t, a good fit:

Standard PayGo (DSQ): Pay per token from a shared pool; cheap, no guarantees.
Priority PayGo: Pay per token at a premium to jump the queue.
Provisioned Throughput (PT): Prepay for reserved capacity; predictable, use it or lose it.

Consumption type

Best for

Watch out for

Standard PayGo (DSQ)

Early-stage, low-QPS, spiky prototype traffic

429s during spikes; no reliability SLO

Priority PayGo

Bursty, revenue-critical traffic that can't tolerate 429s

Roughly 1.8x the standard token price

Provisioned Throughput (PT)

Steady, predictable, high-volume production traffic

Wasted spend if utilization is under ~40%; overflow to PayGo on spikes

The dominant startup mistake is buying PT too early. Usually  this happens the  week after a big launch when it feels like traffic will only ever go up. PT is reserved capacity. You  pay whether you use it or not, and it only starts paying you back once your baseline is genuinely predictable, not just aspirational.

Here’s a pragmatic sequence:

  1. Weeks one through four on Standard PayGo. Use it to measure your real request shape (tokens per minute at p50 and p99, request bursts, batchable vs. real-time split).

  2. When you get your first bad 429 storm, flip on Priority PayGo for the traffic that actually matters. It's a config change, not a purchase order,  nobody in procurement needs to be involved:

code_block
<ListValue: [StructValue([('code', '# Priority PayGo request: use the global endpoint + two extra headers.\r\nfrom google import genai\r\nfrom google.genai import types\r\n\r\nclient = genai.Client(vertexai=True, project="my-startup-prod", location="global")\r\nresp = client.models.generate_content(\r\n model="gemini-2.5-pro",\r\n contents="Rank these support tickets by urgency: ...",\r\n config=types.GenerateContentConfig(\r\n # Priority PayGo headers, per current GEAP docs.\r\n http_options=types.HttpOptions(headers={"X-Vertex-AI-LLM-Request-Type": "shared", "X-Vertex-AI-LLM-Shared-Request-Type": "priority"}),\r\n ),\r\n)'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7f593025f110>)])]>

3. Once you can predict your baseline TPM, buy PT to cover the flat baseline and let anything above it overflow to PayGo. That's the combined pattern Google recommends for exactly this reason. Best of both worlds, not marketing spin.

 Sources: Priority PayGo docs, google-genai HttpOptions source, GEAP REST reference

#7 Which of my requests actually need to be live, and which should be batch jobs?

Most startup workloads are secretly batch jobs pretending to be real-time. Every one you move off the interactive path frees up DSQ headroom for the traffic that genuinely needs to be fast,  the traffic where a user is actually watching a spinner.

Three questions to help you sort your traffic:

  • Does a human have to see the result within a second? That means:  Live inference.

  • Can the user wait a few seconds and see a spinner? That means:  Still live, but a candidate for streaming.

  • Would the user tolerate "we'll email you when it's ready" or "check back in a bit"?  That means: Batch prediction.

Batch prediction on Agent Platform runs in a completely separate queue, does not consume your interactive DSQ, and is typically about half the price of on-demand inference. That's a rare double win: faster live traffic and a lower bill.

code_block
<ListValue: [StructValue([('code', '# Kick off a batch prediction job from a JSONL file in Cloud Storage.\r\n# Each line is one prompt; results land in another Cloud Storage prefix.\r\nfrom google import genai\r\nfrom google.genai import types\r\n\r\nclient = genai.Client(vertexai=True, project="my-startup-prod", location="us-central1")\r\n\r\njob = client.batches.create(\r\n model="gemini-2.5-flash",\r\n src="gs://my-startup-prod-batch/inputs/nightly-summaries.jsonl",\r\n config=types.CreateBatchJobConfig(\r\n dest="gs://my-startup-prod-batch/outputs/",\r\n ),\r\n)\r\nprint(job.name, job.state)'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7f593025e1d0>)])]>

Common candidates: nightly document summarization, background classification of new signups, bulk translation, embedding backfills, evaluation runs against your test set. If any of those are on your live path today, moving them is often the single highest-leverage change you can make this week.

Govern: Keep costs, keys, and agents under control.

#8 How do I set spend caps that actually reduce cost, and not just send me polite emails while my bill triples?

Until recently the honest answer was that budgets only notify, and you had to build your own brake pedal. That changed in July. There are now three mechanisms, and you should think of them as layers.

  1. A spend cap budget (Preview). Cloud Billing budgets can now enforce rather than just email. Set a spend cap on a project and, when usage costs cross 100% of the budget, Google pauses the service until you manually lift it. Agent Platform is explicitly on the eligible list, alongside the Gemini API, Cloud Run, and Cloud Run functions. Alerts still fire at 50% and 80%, so the pause isn't a surprise.

Three things to know before you rely on it:

  • Each cap covers one project and one eligible service. It is not account-wide protection. If you want Agent Platform and Cloud Run both capped, that's two caps. 

  • Enforcement is not instant and is based on estimated costs. Overages past the cap are billed as normal, so set the number below your real ceiling. Lifting it is manual, and service resumption can take up to an hour. It also pauses Provisioned Throughput usage, so if you've prepaid for capacity, a cap hit stops that too.

  • It's in Preview as of publication, and the eligible-service list is documented as growing. Check the current list before you design around it.

2. A billing budget with a Pub/Sub trigger that disables billing. Still the right tool when you need blast radius the spend cap can't give you: multiple services at once, an entire project, or a service that isn't eligible yet. When the budget hits a threshold, Pub/Sub fires a Cloud Function that detaches the billing account, which stops all billable activity within minutes. Blunter and more dangerous than the native cap — it can leave resources unrecoverable — so reach for it second, not first. Full walkthrough: Automatically respond to budget notifications.

code_block
<ListValue: [StructValue([('code', '# Sketch: create a budget SCOPED TO ONE PROJECT that publishes to Pub/Sub at 50%, 90%, 100%.\r\ngcloud billing budgets create \\\r\n --billing-account=012345-6789AB-CDEF01 \\\r\n --display-name="my-startup-prod hard stop" \\\r\n --budget-amount=2000USD \\\r\n --filter-projects=projects/my-startup-prod \\\r\n --threshold-rule=percent=0.5 \\\r\n --threshold-rule=percent=0.9 \\\r\n --threshold-rule=percent=1.0,basis=current-spend \\\r\n --notifications-rule-pubsub-topic=projects/my-startup-prod/topics/budget-alerts'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7f593025f910>)])]>

Sources:  Manage spend cap budgets, Set up programmatic notifications gcloud billing budgets create reference, Cloud Billing budgets concepts, Disable billing with notifications walkthrough, Programmatic notification payload schema.

Two things to get ahead of  for, as the defaults can cause unexpected issues: 

  1. Limit your budget scope: Without --filter-projects, your budget applies to your entire billing account. A spike in any project will trigger the kill switch for everything. 

  2. Deploy locally: The budget notification doesn't specify which project is affected. To ensure the kill switch only affects the intended project, deploy your Cloud Function in the same project you're protecting (e.g., my-startup-prod).

Then wire up a tiny Cloud Function to that topic that calls projects.updateBillingInfo to unlink the billing account when the 100% threshold fires. That is your circuit breaker.

Mechanical ceilings via quota overrides. Even if you never set up the above kill switch, you can cap the rate at which cost can accumulate by setting explicit per-model, per-region quotas below the platform default. If your app never legitimately needs more than 500 requests per minute for gemini-2.5-pro, cap it there in the Cloud Quotas console; a leaked key can't burn what the quota flatly refuses to serve.

#9 Where should I actually keep secrets? (Not in .env files!)

The short answer is: Secret Manager. Not  in environment variables, not in .env files, and never in your repo. Grant read access via IAM only to the service account that needs it.

code_block
<ListValue: [StructValue([('code', '# Store a third-party API key (Stripe, OpenAI, whatever).\r\necho -n "sk_live_xxx" | gcloud secrets create stripe-live-key --data-file=-\r\n\r\n# Grant only the runtime service account access to read it.\r\ngcloud secrets add-iam-policy-binding stripe-live-key \\\r\n --member=serviceAccount:agent-runtime@my-startup-prod.iam.gserviceaccount.com \\\r\n --role=roles/secretmanager.secretAccessor'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7f593025f9d0>)])]>
code_block
<ListValue: [StructValue([('code', '# Application code fetches it at startup; nothing lives on disk.\r\nfrom google.cloud import secretmanager\r\nsm = secretmanager.SecretManagerServiceClient()\r\nresp = sm.access_secret_version(\r\n name="projects/my-startup-prod/secrets/stripe-live-key/versions/latest"\r\n)\r\nstripe_key = resp.payload.data.decode("utf-8")'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7f593025e9d0>)])]>

Then two little disciplines that pay for themselves the first time you need them:

  • Rotation on a schedule and on suspicion. Secret Manager versions are cheap; treat them as immutable and roll forward. 

  • Detection when a secret leaks. Secret Manager notifications and Google Cloud's Sensitive Data Protection can catch keys checked into a repo or pasted into a log stream,  before an attacker does.

For any AI application that acts on a user's behalf, calls Gmail on their behalf, reads a Drive folder, hits a third-party SaaS with the user's credentials, do not store a long-lived token. Use OAuth 2.0 with short-lived access tokens and a refresh flow, so that when a user rage-quits or a compromised account gets revoked, the agent loses access at the same time. 

#10  How do I stop my brand new AI agent from doing something it absolutely shouldn't?

An agent that can call tools, browse the web, or execute code needs the same defense-in-depth thinking as any other production service, arguably more, because it makes decisions that neither you nor the model can fully predict in advance.

Four layers, none optional once you have real users:

1. Identity for the agent itself. Give the agent its own service account, scoped only to the resources and tools it genuinely needs,  the exact same least-privilege principle as any other workload. Agent Engine supports first-class agent identity so every action can be attributed to a specific agent instance in your audit logs.

2. Sandboxed code execution. If your agent runs generated code,  a common pattern for data-analysis or "run this Python for me" flows, do not run it in your application process. Use an isolated sandbox so a bad combination can't touch your production data.

code_block
<ListValue: [StructValue([('code', '# Enable server-side code execution inside a sandbox for a request.\r\nfrom google import genai\r\nfrom google.genai import types\r\n\r\nclient = genai.Client(vertexai=True, project="my-startup-prod", location="us-central1")\r\nresp = client.models.generate_content(\r\n model="gemini-2.5-pro",\r\n contents="Compute the correlation between these two columns: ...",\r\n config=types.GenerateContentConfig(\r\n tools=[types.Tool(code_execution=types.ToolCodeExecution())],\r\n ),\r\n)'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7f593025e890>)])]>

3. Prompt and response filtering. Model Armor sits in front of your model calls and screens for prompt injection, jailbreaks, sensitive-data exfiltration, and off-brand output,  all of which are essentially guaranteed the moment you have real users being real users.

4. Behavioral monitoring. Security Command Center with threat detection flags anomalies in agent behavior,  a service account suddenly calling an API it's never touched before, an agent reaching out to an unfamiliar external host, an unexpected spike in privileged operations. In near-real-time.

None of these are optional once your agent is acting on behalf of a real user or handling real money.

Your homework, so to speak:

  1. Audit for raw API keys in your repo, your notebooks, and your production runtime. Rotate anything that shouldn't be there.

  2. Move any workload that doesn't need a synchronous response to the Batch API.

  3. Turn on the Model observability dashboard and put one alert on capacity errors, so the next 429 reaches you before it reaches a customer.

  4. Set a spend cap on the project, and keep an eye out for 50% and 80% alerts. If usage crosses 100% of the budget, Google will pause the service until you manually lift it.

Do those four things this week and you're already ahead of most startups shipping AI features. 

Have a scenario you'd like us to cover next? Reach us at Google Cloud for Startups.

Announcing quantum-safe key import in Cloud KMS

詳細を表示

As enterprises increasingly adopt multicloud architectures, bring your own key (BYOK) has become a fundamental pillar for maintaining data sovereignty and helping protect critical cloud workloads. At the same time, quantum computing has rapidly advanced, and security teams need to re-evaluate how they securely transfer encryption keys across networks.

Following our previous announcements of quantum-safe digital signatures and quantum-safe key encapsulation mechanisms (KEMs) in Cloud Key Management Service (Cloud KMS), we are excited to announce the preview of quantum-safe key import in Cloud KMS for software-based cryptographic keys.

Our updated quantum-safe BYOK capability, the first step of the next phase of our post-quantum cryptography (PQC) migration timeline, can help you protect your sensitive keys before a cryptographically-relevant quantum computer (CRQC) emerges.

As you adopt quantum-safe key import to help protect your keys in transit, you can also monitor your overall post-quantum posture with Cloud KMS PQC insights, now generally available. This high-level visual illustrates your asymmetric keys based on the categorization of the algorithms they use, and can help you plan for future modernization and support long-term resilience.

The threat: Store Now, Decrypt Later attacks

Traditional key import methods rely on classical asymmetric encryption standards to wrap keys during transit. While these algorithms successfully defend against today’s threats, they will become fundamentally insecure when a viable quantum computer emerges that can potentially decrypt keys that adversaries have intercepted and stored. 

Quantum-safe key import helps mitigate these store now, decrypt later (SNDL) attacks by wrapping your keys in a quantum-resistant envelope from day one.

Building a quantum-resistant envelope for keys

The post-quantum transit mechanism now available in Cloud KMS uses hybrid public key encryption (HPKE). Our new import method wraps your sensitive software key material in a quantum-resistant transit envelope. The process integrates into the existing Cloud KMS API workflow to minimize your work:

  • Initiating the job: The client creates a new import job through the Cloud KMS API, requesting a post-quantum HPKE import method.

  • Key generation: The Cloud KMS server generates a post-quantum KEM private key and exposes the corresponding public key to the client.

  • Client-side wrapping: Using a supported cryptographic library (such as Tink or OpenSSL), the client executes an HPKE Seal() operation. This encapsulates the public key to establish a shared secret, derives an ephemeral AES key using HKDF-SHA256, and encrypts the target key material.

  • Submission: The client transmits the encapsulated ciphertext concatenated directly with the encrypted key material back to the Cloud KMS endpoint, which already has quantum-safe data-in-transit protection built-in.

  • Unwrapping: The Cloud KMS server executes an HPKE Open() operation using its private portion of the wrapping key to safely decrypt and help protect the key material within the Cloud KMS boundary.

For the KEM layer, you can choose between X-Wing, ML-KEM-768, or ML-KEM-1024. The key derivation layer utilizes HKDF-SHA-256, and the final symmetric wrapper employs AES-256-GCM with standard 12-byte nonces.

To learn more about setting up your import jobs, preparing your local key material using external cryptographic libraries, and managing quantum-safe solutions, check out our Cloud KMS quantum safe key import documentation.

A critical milestone in Google Cloud's PQC journey

The global migration to post-quantum cryptography is a marathon that you take one milestone at a time. Today, you can create your first quantum safe key import job and begin the process of helping make your applications quantum-safe. 

We welcome your feedback and invite you to reach out to explore how we can support your organization's post-quantum strategy.

Expanding Google Antigravity for enterprise customers

詳細を表示

Since announcing Google Antigravity in Gemini Enterprise Agent Platform at I/O in May, we’ve heard helpful feedback from our customers. Your developers want easy access to coding agents across surfaces. Your enterprise governance team wants security controls and license management. And your finance team wants pooled usage so that no prepaid token ever goes unused. Now, everybody finally gets what they want:

  • Antigravity is available now as part of eligible Gemini Enterprise app subscriptions, including out-of-the-box administrative and spend controls.

  • New IDE extensions let developers use Antigravity in the IDEs of their choice, including VS Code.

Unify AI developer tools and enterprise-grade controls in one subscription

Equipping your developers with advanced agentic tools shouldn't mean managing separate add-on licenses, invoices, billing consoles or security settings. With AI developer tools included in Gemini Enterprise subscriptions, administrators can easily enable Antigravity and Android Studio for users with eligible Gemini Enterprise Standard, Plus, and Standard Emerging Market licenses, and maintain full governance with spend, security, observability and usage metrics consolidated in the Gemini Enterprise admin console.

Unblock your developers while controlling spend

With billing flexibility and cost management tools in Gemini Enterprise, you can ensure your developers have the resources they need while managing costs:

  • Granular spend thresholds: Administrators can set monthly project-level budget caps directly in the Billing console, with additional per-user and team controls rolling out later this year.

  • Pooled quotas: Shared token pools provide flexibility to high-demand teams, preventing purchased quota from sitting idle across the organization.

  • Overage enablement: To maintain continuous developer workflows when pooled quotas are met, administrators can opt into overages with monthly spend caps, smoothly transitioning excess usage to standard consumption-based rates. 

  • Usage metrics: Centralized usage tracking provides visibility into token consumption, API calls, and developer activity, enabling organizations to continuously optimize their AI investments.

<div class="article-module h-c-page">
  <div class="h-c-grid">


<figure class="article-image--large
  
  
    h-c-grid__col
    h-c-grid__col--6 h-c-grid__col--offset-3
    
    
  ">

  
  
    
    <img alt="Gif 1" src="https://storage.googleapis.com/gweb-cloudblog-publish/original_images/Gif_1_vK8C2Nf.gif" />
    
    </a>
  
</figure>


  </div>
</div>

Safeguard your organization’s code and data with built-in privacy and security

Gemini Enterprise subscriptions bring Google Antigravity under Google Cloud’s standard security and compliance protections. Administrators and IT teams can set clear boundaries around workspace access, enable full audit logging, and enforce data privacy from a single console:

  • Configurable security policies: Enforce security and compliance controls, such as workspace sandboxing, and browser and MCP server access, to help ensure AI agents operate safely within authorized enterprise environments.

  • Central audit logging: Enable comprehensive audit logging with a single toggle, capturing prompts, agent responses, and metadata for compliance reporting.

  • Data privacy: Maintain data ownership under Google Cloud’s Terms of Service, ensuring all agent activity executes strictly within your secure cloud boundary.

Bring agentic coding directly into your team’s preferred development environments 

Starting today your developers can use Antigravity across the surfaces they already know and use — including Visual Studio Code, Visual Studio (preview), Jetbrains (preview) and Zed IDEs (preview)  via the new IDE extensions as well as the Antigravity 2.0 desktop app, and the Antigravity CLI.

Throughout all surfaces, administrators can enforce corporate identity standards while removing setup friction for technical teams via native support for Workforce Identity Federation (WIF) and Application Default Credentials (ADC).

<div class="article-module h-c-page">
  <div class="h-c-grid">


<figure class="article-image--large
  
  
    h-c-grid__col
    h-c-grid__col--6 h-c-grid__col--offset-3
    
    
  ">

  
  
    
    <img alt="VSCode IDE Plugin" src="https://storage.googleapis.com/gweb-cloudblog-publish/original_images/VSCode_IDE_Plugin_gif_v3_720p.gif" />
    
    </a>
  
    <figcaption class="article-image__caption "><p>AGY IDE extension demo</p></figcaption>
  
</figure>


  </div>
</div>

What our customers are saying

From rapid code generation to end-to-end task automation, Google Antigravity is giving engineering teams the momentum of cutting edge AI development backed by the stability, governance, and scale of Google Cloud. Here is how leading enterprise customers and partners are driving measurable outcomes in production:

“Deploying Antigravity in Gemini Enterprise allows Accenture to arm our engineers with Google DeepMind’s premier technology on the secure, trusted foundation of Google Cloud. Abstracting away operational complexity ensures our teams don't have to choose between developer speed and enterprise-grade governance — freeing them to deliver high-velocity engineering and transformative value for our clients.” — Chetna Sehgal, Global Practice Lead, Accenture Google Business Group

“At AirAsia and across the Group, we’re all about empowering our people. Bringing highly capable Gemini models directly into our daily workflows with Antigravity 2.0 does exactly that. We are putting the most advanced AI capabilities into the hands of our entire workforce, from software engineering to finance, marketing, legal, HR and much more. This empowers both our developers and critical back-office teams to innovate at an unprecedented pace and drive proven time-savings across the board.” — Nikunj Shanti, CTO, AirAsia Next

“Enterprises are moving beyond AI experimentation and expecting measurable business outcomes. With Gemini Enterprise and next-generation developer tools like Antigravity 2.0 and Antigravity CLI, we see significant opportunities to further embed agentic AI, particularly the advanced reasoning capabilities of Gemini models, directly into software delivery workflows. This goes beyond productivity as it enables faster decision-making, higher code quality, and reduced technical debt at scale. What stands out is how these capabilities are helping our teams evolve from writing code to orchestrating outcomes, strengthening every phase of the software development lifecycle while scaling innovation securely and responsibly.” — Rakesh Aerath, President, Asia Pacific Global Delivery Centers of Excellence, CGI

“Every developer workflow is unique, and agentic AI should adapt to the engineer, not the other way around. With Google Antigravity supported across developers' preferred IDEs, the desktop app, and the CLI, Cognizant can seamlessly embed agentic engineering across our global delivery centers. It gives our teams the freedom to choose their preferred surface while delivering high-velocity, secure software for our clients.”— Rajesh Varrier, President, Global Operations and Chairman & Managing Director, Cognizant India

“Antigravity has played a key role in advancing our AI-first strategy at Datamatics. Over the last few months, our teams have used it to rapidly build and deploy multiple applications, accelerate solution development, and embed AI into core business processes. From AI Impact Hub to analytics and sales enablement solutions, it has helped us move beyond experimentation to real execution, delivering measurable business outcomes while enabling teams to innovate faster and at scale.”— Vijay Venkatachalam, Vice President, Information Systems Group (ISG), Datamatics

"Embedding Google Antigravity's autonomous capabilities directly into Gemini Enterprise development environments allows our internal and Forward Deployed Engineering teams to automate complex tasks. Supported by the platform’s new FinOps and governance controls, our teams can confidently focus on orchestrating high-value, secure and cost-efficient outcomes for our clients at scale."  — Faruk Muratovic, US AI & Engineering Strategy and Services Leader, Deloitte

“Adopting Antigravity places Wipro at the leading edge of the AI-driven software development lifecycle. Combining Antigravity 2.0, CLI, and the new IDE extensions with a seamless developer experience and superior code accuracy fits naturally into our AI-first engineering strategy. Working with Google Cloud allows us to accelerate software delivery and bring next-generation value to our global enterprise clients.”  — Debashish Ghosh, Vice President and Global Head, Google Partnership, Wipro

Get started with Antigravity in Gemini Enterprise

Google Antigravity in Gemini Enterprise is available today for eligible Gemini Enterprise Standard, Plus, and Standard Emerging Market licenses, with broader support coming soon.

  • For administrators: Visit the enterprise setup guide to enable AI Developer tools for Google Antigravity and Android Studio. 

  • For developers: Start building with Antigravity 2.0, the Antigravity CLI, or your preferred IDEs via Antigravity IDE extensions.

Google Cloud Japan Blog

プライバシー重視の AI で脳腫瘍研究を後押し

詳細を表示

※この投稿は米国時間 2026 年 8 月 7 日に、Google Cloud blog に投稿されたものの抄訳です。

医療と AI の融合は、目覚ましいイノベーションをもたらしています。しかし現在、開発者は患者のプライバシー保護と、実世界の多様なデータに基づく検証と評価を両立させた、信頼性の高い医療 AI ツールを構築するという難題に直面しています。Google Cloud は、戦略的なコラボレーションと Confidential Computing を組み合わせてこの難題にアプローチしています。

検証プロセスにおいて患者のプライバシーと AI モデルの両方を保護するために、Google は MedPerf イニシアチブを通じて MLCommons と連携しています。今年の Google Cloud Next で初めて発表されたこのパートナーシップは、Confidential Computing を活用して、実際の環境で AI モデルをベンチマークするための安全なデータ クリーンルームを確立するものです。

課題: データを見ることなく AI を評価する

テクノロジー業界と学術界に 125 以上のメンバーを擁するグローバル コミュニティである MLCommons は、医療 AI の評価を標準化するため、2023 年に MedPerf を立ち上げました。AI モデルのベンチマーク用オープンソース プラットフォームである MedPerf は、連携評価によるモデルテストを通じて臨床研究を推進しています。

Google Cloud Confidential Space を使用すると、ハードウェアで分離された高信頼実行環境(TEE)内で独自の AI モデルを評価できます。この特別な仮想マシンは、実行中のメモリを暗号化し、オペレーティング システムを強化するため、病院や研究機関、その他の参加者、さらには Google であっても、評価中にモデルコードや患者データを閲覧することはできません。

医療 AI のベンチマークには高いコンピューティング負荷がかかるため、Confidential VM は CPU だけでなく GPU にも拡張されています。GPU 加速推論時においてもモデルの重みと患者データを保護するため、MedPerf は NVIDIA H100 GPU を搭載した Google Cloud の A3 マシンシリーズ上で稼働します。このシステムは、CPU の Intel TDX テクノロジーと GPU の NVIDIA Confidential Computing を組み合わせた構成となっています。

患者データがワークロードに投入される前に、システムは、承認されたコードのみが正規の Confidential Computing ハードウェアで実行されていること、および環境が適切にセキュリティ強化されていることを暗号によって証明します。

ML Commons MedPerf が Google Cloud Confidential Computing と連携し、医療 AI モデルの安全な実環境評価を実現する仕組みをご紹介しています。

理論を大きな成果へ: 脳腫瘍研究のさらなる進展

このテクノロジーは、Federated Tumor Segmentation(FeTS)イニシアチブを通じて、すでに極めて重要な研究を推進する原動力となっています。膠芽腫などの脳腫瘍は症例数が少ないため、単一の病院だけで AI トレーニングの精度を高めるのに十分なデータを収集することは困難です。

さらに厄介なことに、ある病院では完璧に機能するモデルが、患者の属性、データの取得手法、さらには機器の差異により、別の病院では十分な性能を発揮できないケースがあります。

インディアナ大学の Spyridon Bakas 博士、ノースウェスタン大学の Yury Velichko 博士、カナダのアルバータ大学の Amber Simpson 博士といった先見の明のある研究者の協力を得て、Google Cloud 上の MedPerf は世界各地の脳 MRI の非公開データで AI モデルを検証し、潜在的なパフォーマンス ギャップを特定しています。

たとえば、あるサイトでは 95% の精度を発揮するモデルが、別のサイトでは 63% の精度しか達成できない場合があります。Google の協働的アプローチは、臨床現場に届く AI ツールが、特定の層に偏らない真に代表的な患者集団で機能することを保証するものです。

臨床現場での信頼確立と検証

このコラボレーションによる成果は、臨床研究の最前線にいる方々の言葉に最もよく表れています。

ノースウェスタン大学放射線科准教授の Yury Velichko 博士は次のように述べています。「Google Cloud でのフェデレーション ラーニングのテストを通じて、医療 AI の未来は、安全かつスケーラブルな、共同研究に適したクラウド環境にあることがわかりました。管理された研究室レベルの環境から一歩踏み出し、本番環境対応のインフラストラクチャでこれらのワークフローを検証できたことは、実際の臨床アプリケーションにおけるフェデレーション ラーニングのパフォーマンスとセキュリティを評価する、またとない機会となりました。」

MLCommons の MedPerf リードである Alexandros Karargyris 氏は次のように語っています。「医療 AI は世界中の患者に多大な恩恵をもたらす可能性を秘めています。しかしそれを実現するには、評価に用いるベンチマークが臨床医、研究者、規制当局から信頼されるものでなければなりません。MedPerf を Google Cloud の Confidential Computing インフラストラクチャ上に構築することで、プライバシー、知的財産、ベンチマークの整合性を損なうことなく、実際の患者データで AI モデルを厳密にテストできる未来に向けた大きな一歩を踏み出せました。」

未来:安全な医療のブレークスルーを普及させる

MLCommons と Google Cloud のコラボレーションは、医療 AI における「プライバシー バイ デザイン」への根本的な転換を示すものです。Google は、データとモデルを安全に共有、評価しやすくすることで、より迅速かつ安全で公平性の高い医療イノベーションの実現を後押ししています。

Google Cloud 上での MedPerf プラットフォームの活用にご関心のある研究機関や医療モデル開発者の方は、medical@mlcommons.org または担当の Google Cloud アカウント チームまでお問い合わせください。

- Google Cloud、シニア クラウド セキュリティ プロダクト マネージャー、Rene Kolga

- Google、ソフトウェア エンジニア、Peter Mattson

AI 時代のデジタル主権: 制御とイノベーションの二者択一が不要に

詳細を表示

※この投稿は米国時間 2026 年 8 月 7 日に、Google Cloud blog に投稿されたものの抄訳です。

コンプライアンスと主権に関する厳格な要件を持つ企業や政府機関は、機密データをオンプレミスに保持しているがために最新の AI を利用できないことがよくあります。こうした組織は、次の 3 つの大きなリスクを管理しています。

  1. 管轄権のリスク: 現地の規制の変化や、知的財産(権)保護の必要性、国外からのデータアクセス要求といった要因により国内でのデータ管理が不可欠となっています。

  2. 経済的自立: 海外のインフラストラクチャ プロバイダに依存すると、重要なサービスの脆弱化を招くおそれがあります。

  3. 地政学的リスク: 予測不可能な世界情勢の混乱から重要な国内サービスを保護する必要があります。

Google が「AI インフラストラクチャの現状」レポート作成のため最近実施した調査では、対象となった 1,400 人を超えるシニア IT リーダーの 48% が、現地のデータ セキュリティ法を遵守するため、データ所在地の管理機能を備えたインフラストラクチャを優先していると回答しています。

<figure class="article-image--wrap-small
  
  ">

  
  
    
    <img alt="image1" src="https://storage.googleapis.com/gweb-cloudblog-publish/images/image1_0SuEihC.max-1000x1000.png" />
    
    </a>
  
</figure>







  <p>しかし、オンプレミス環境にとどまることは、もはや最新のイノベーションから取り残されることを意味しません。このギャップを埋める手段として、オンプレミスとマルチクラウドを組み合わせたハイブリッド ソリューションを導入する組織が増えてきています。Google の調査によると、現在 52% の組織が AI 活用のためにハイブリッド クラウドのアプローチを採用しています。</p><p>このアプローチでは、パブリック クラウドの圧倒的な処理能力を活用しながら、主権やコンプライアンス維持といったローカル環境のメリットも享受でき、データの所在地やアクセス権を自ら完全に制御することが可能です。これまで、こうした厳格なデータルール規制下にある組織は、高度な AI を簡単に取り入れることはできませんでした。独自の AI システムの構築も、スピードとコストの両面で現実的ではありませんでした。</p><p>そこで登場したのが、Google の <a href="https://cloud.google.com/distributed-cloud">Google Distributed Cloud(GDC)</a>です。GDC を使用すると、データセンターやエッジなど、お客様が必要とする場所で Google Cloud を利用できるようになります。AI ワークロードの主権要件を満たすために、次の 2 つのデプロイモデルが用意されています。</p><ul><li><b>エアギャップ:</b> Google Cloud やパブリック インターネットから完全に遮断された環境で動作する、外部接続を必要としないソリューション。Google がリモートでシャットダウンすることはできません。</li><li><b>コネクテッド:</b> 既存のハードウェア上で直接実行される、Google がソフトウェア ライフサイクルを統合的に管理するモデル。</li></ul><p>GDC は、AI ワークロード向けに最適化されたインフラストラクチャ、Gemini またはオープンモデルの選択肢、費用対効果の高い推論サービスを備えた、オンプレミスの完全な AI ソリューションを提供します。この基盤を活用すれば、データを完全に制御しながらセキュアな AI エージェントを構築して実行できます。</p><p><b>オンプレミスでソブリン AI のニーズを満たす</b></p><p>データ管理か、AI イノベーションかの二者択一を迫られることはもうありません。Google Distributed Cloud なら、データの主権を完全に維持しながら世界をリードする AI を自社の環境に直接導入できます。</p><p>大手企業のハイブリッド戦略の詳細を紹介した「<a href="https://cloud.google.com/resources/content/state-of-infrastructure-in-the-agentic-ai-era?e=48754805">AI インフラストラクチャの現状</a>」レポートを、ぜひご一読ください。</p><p><b><i>-</i></b> <i>Google Cloud、Distributed &amp; Sovereign Cloud 担当バイス プレジデント兼ゼネラル マネージャー</i><b><i>、Ankur Mehrotra</i></b></p>
</div>

Gemini in Database Migration Service を使用した PostgreSQL への移行の迅速化

詳細を表示

※この投稿は米国時間 2026 年 8 月 12 日に、Google Cloud blog に投稿されたものの抄訳です。

次のようなシナリオを考えてみましょう。Oracle や SQL Server などの既存の商用データベースから、オープンソースの PostgreSQL や AlloyDB for PostgreSQL などのフルマネージド サービスにコア アプリケーションを移行するとします。

最初のフェーズは順調に進みます。スキーマが変換され、テーブルにデータが入力され、データ移行パイプラインが数テラバイトのデータを数時間で転送します。プロジェクトは予定より早く進んでいるようです。

そのとき、チームはボトルネックに直面します。

既存のデータベースには、何百ものストアド プロシージャ、複雑なトリガー、そして PL/SQL や T-SQL など独自の SQL 言語で記述されたカスタム関数が埋め込まれています。これらのルーチンには、トランザクションの検証、注文処理、カスタム レポートの処理など、長年にわたる重要なビジネス ロジックが含まれています。

モダナイゼーション プロジェクトは急に行き詰まります。数千行に及ぶ手続き型ロジックの変換には、2 つの言語に関する専門知識、数か月にわたる手作業での書き換えが必要となり、高確率で変換エラーが伴います。このコード変換は、データベース移行の「ラスト ワンマイル」のボトルネックであり、移行における最も複雑な部分です。

幸いなことに、このラスト ワンマイル問題は、最近の AI の進歩によって解決できます。Database Migration Service(DMS)には、Gemini を活用した AI によるコード変換が含まれています。移行ワークフローに生成 AI を直接組み込むことで、ストアド プロシージャ、トリガー、カスタム関数を、より高速かつ正確に PostgreSQL PL/pgSQL コードに変換できます。

ストアド プロシージャの変換に関する課題

商用データベース エンジンでは、ストアド プロシージャ、ユーザー定義関数、パッケージ本体、条件付きロジックに、ベンダー固有の構文が使用されています。このロジックを PostgreSQL PL/pgSQL に変換するには、変数の定義、例外処理ブロック、カーソルループ、組み込み関数のマッピングが必要になります。

何百ものストアド プロシージャを含む複雑なエンタープライズ スキーマを移行する場合、たいていは手動でのコード変換に数か月のエンジニアリング作業が必要になります。データベース チームは、以前のロジックを 1 行ずつ解析し、条件分岐を実装し直して、エンジン間のデータ型変換を検証する必要があります。

DMS での AI によるコード変換

Gemini in Database Migration Service は、Google Cloud コンソール内で直接、この変換作業を迅速化します。DMS は、スキーマを自動的に変換するとともに、AI 生成のコードを提案して、ソース言語と PostgreSQL の構造上の違いを説明します。

PL/pgSQL に変換されたコードは、元のソースコードと並べて表示されます。データベース チームは提案内容をリアルタイムで確認、編集、検証できます。

<div class="article-module h-c-page">
  <div class="h-c-grid">


<figure class="article-image--large
  
  
    h-c-grid__col
    h-c-grid__col--6 h-c-grid__col--offset-3
    
    
  ">

  
  
    
    <img alt="1 - DMS_Code_Conversion_Console" src="https://storage.googleapis.com/gweb-cloudblog-publish/images/1_-_DMS_Code_Conversion_Console.max-1000x1000.jpg" />
    
    </a>
  
    <figcaption class="article-image__caption "><p>図 1: Database Migration Service のインターフェースに、変換前後のコードと Gemini によるインラインの説明が表示されている。</p></figcaption>
  
</figure>


  </div>
</div>

AI の統合が重要である理由

主要ベンダーのほとんどの AI アプリや AI ツールには、SQL コードなどのコードを生成、変換する機能があります。しかし、汎用の AI チャットツールが提供するスニペットの変換機能では、エンタープライズ データベースの変換には不十分です。Gemini in Database Migration Service には、以下の重要な利点があります。

  • スキーマの完全なコンテキスト: Gemini in DMS は、コード スニペットを個別に評価するのではなく、移行プロジェクト全体にわたり、テーブルの関係、データ型、依存ビュー、プロシージャ間参照などのデータベース コンテキストを総合的に分析します。

  • エンタープライズ向けのセキュリティとプライバシー: コード変換は、厳密に Google Cloud プロジェクトの境界と IAM ガバナンスの範囲内で実行されるため、独自のビジネス ロジックと知的財産を保護できます。

  • 統合された実行ワークスペース: DMS を使用すると、何百ものファイルにわたる手動でのコピーと貼り付けが不要になります。単一のコンソール内で、並べて表示されたコードの違いと AI によるインラインの説明を確認し、必要に応じてコードを編集したうえで、検証済みの PL/pgSQL ルーティンをターゲット データベースに直接デプロイできます。

  • 決定論的な精度と AI コンパイル: DMS は決定論的なコンパイラ ルールによる 1 対 1 のマッピング(標準的な DDL 変換、スカラー関数、明確に定義された構文変換など)と、複雑な手続き型ブロックのための Gemini の文脈的統合を組み合わせることで、モデルのドリフトを防ぎ、正確で予測可能な変換を保証します。

Oracle PL/SQL から PostgreSQL への変換

顧客の合計注文額を算出し、独自の NVL 関数と DECODE 関数を使用して段階別の割引を適用する Oracle PL/SQL ストアド プロシージャを考えてみましょう。当初のワークフローでは、NVL を COALESCE に手動でマッピングし、DECODE ステートメントを標準の CASE 式に書き換え、WHEN NO_DATA_FOUND THEN などの例外ブロックを調整する必要がありました。

DMS で移行評価を実行すると、Gemini がソース プロシージャを分析し、ネイティブの PostgreSQL PL/pgSQL コードを生成します。

ソース: Oracle PL/SQL

code_block
<ListValue: [StructValue([('code', 'CREATE OR REPLACE PROCEDURE calculate_discount (\r\n p_customer_id IN NUMBER,\r\n p_discount OUT NUMBER\r\n) AS\r\n v_total NUMBER := 0;\r\nBEGIN\r\n SELECT NVL(SUM(amount), 0) INTO v_total\r\n FROM orders WHERE customer_id = p_customer_id;\r\n \r\n p_discount := DECODE(TRUE, v_total > 10000, 0.15, v_total > 5000, 0.10, 0.05);\r\nEXCEPTION\r\n WHEN NO_DATA_FOUND THEN\r\n p_discount := 0;\r\nEND;\r\n/'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7f0e92cda190>)])]>

ターゲット: PostgreSQL PL/pgSQL(Gemini in DMS で変換)

code_block
<ListValue: [StructValue([('code', 'CREATE OR REPLACE FUNCTION calculate_discount (\r\n p_customer_id NUMERIC,\r\n OUT p_discount NUMERIC\r\n) RETURNS NUMERIC AS $$\r\nDECLARE\r\n v_total NUMERIC := 0;\r\nBEGIN\r\n SELECT COALESCE(SUM(amount), 0) INTO v_total\r\n FROM orders WHERE customer_id = p_customer_id;\r\n\r\n p_discount := CASE\r\n WHEN v_total > 10000 THEN 0.15\r\n WHEN v_total > 5000 THEN 0.10\r\n ELSE 0.05\r\n END;\r\nEND;\r\n$$ LANGUAGE plpgsql;'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7f0e93f34210>)])]>

Gemini は、生成された SQL とともに、NVL が COALESCE に変換された理由と、Oracle の DECODE 関数が PostgreSQL で明示的な CASE ブロックに変換された方法について、詳しい説明をインラインで表示します。

<div class="article-module h-c-page">
  <div class="h-c-grid">


<figure class="article-image--large
  
  
    h-c-grid__col
    h-c-grid__col--6 h-c-grid__col--offset-3
    
    
  ">

  
  
    
    <img alt="2 - Migration_Workflow_Diagram" src="https://storage.googleapis.com/gweb-cloudblog-publish/images/2_-_Migration_Workflow_Diagram.max-1000x1000.jpg" />
    
    </a>
  
    <figcaption class="article-image__caption "><p>図 2: Gemini in DMS を活用したエンドツーエンドのデータベース コード変換パイプライン。</p></figcaption>
  
</figure>


  </div>
</div>

スキーマの完全な管理と検証の維持

セキュリティ、透明性、コードの精度は、データベースのモダナイゼーションにおいて依然として中心的な要素です。Gemini in DMS は、厳密に Google Cloud の確立されたセキュリティ境界内で動作し、コードをプロジェクト内に限定して公開します。

変換と検証のプロセスは、信頼性を保証するため、構造化されたワークフローに沿って行われます。

  • スキーマ コンテキストの自動取得: DMS 変換ワークスペースを設定すると、テーブル スキーマ、データ型、外部キー制約、プロシージャ間の依存関係など、ソース データベースの全メタデータが自動的に解析されます。Gemini では、このプロジェクト全体のコンテキストを参照してコードが生成されるため、依存オブジェクトの定義を手動で指定する必要はありません。

  • 構文と依存関係の自動検証: コードの生成時、DMS はターゲットとなる PostgreSQL の構文ルールに照らして検証パーサーを実行します。検証ステータス インジケーター(変換済み、警告、要対応など)が各オブジェクトに割り当てられるため、手動での確認が必要なルーティンはすぐにわかります。

  • インタラクティブな評価状態: スキーマの変更はすべて完全に管理できます。変換ワークスペース内で、ターゲット データベースに変更を適用する前に、コードを並べて違いを確認したり、AI によるインラインの説明を確認したり、PL/pgSQL コードを直接編集したりできます。

  • デプロイと検証のステージング: コードがワークスペース検証に合格したら、変換されたスキーマと関数をターゲットのステージング インスタンス(Cloud SQL や AlloyDB など)に適用し、本番環境への切り替え前に機能の実行とパフォーマンス テストを行うことができます。

データベースのモダナイゼーションの効率化

Database Migration Service の AI によるコード変換により、データベース チームは以前のデータベース ロジックを数か月ではなく数日で変換できます。データベース管理者とアプリケーション開発者は、貴重な時間を費やしてコードを最初から書き直す代わりに、新しい機能の追加、パフォーマンスのテスト、アプリケーションのモダナイゼーションに注力できます。

Oracle と SQL Server の一般的な変換シナリオと、DMS がそれらを PostgreSQL に変換する方法の代表的な例については、新しい Gemini に習う PostgreSQL 動画シリーズをご覧ください。

データベース移行に関する悩みを解消し、以前のデータベース ロジックをシームレスに変換する方法を Gemini に教えてもらいましょう。

Database Migration Service を使ってみる

データベースの変換を始めるには、Database Migration Service コンソール(https://console.cloud.google.com/dms)で移行評価を実行するか、異種間移行ガイド(https://cloud.google.com/database-migration)をお読みください。

- Google Cloud、戦略的クラウド エンジニア、Tanya Sharma

「ノイジー ネイバー」問題の解決: シャーディング アーキテクチャがマルチテナント プラットフォームを守る方法

詳細を表示

※この投稿は米国時間 2026 年 8 月 6 日に、Google Cloud blog に投稿されたものの抄訳です。

マルチテナントの SaaS プロバイダ、社内データ プラットフォームを管理する大企業、混合ワークロードのデータ処理を扱う企業のいずれであっても、共有インフラストラクチャ環境を管理するということは、「ノイジー ネイバー」という共通の脅威に直面することを意味します。単一のテナントで大量のデータバーストやデータベース インスタンスの障害が発生した場合、環境内の他のテナントをダウンさせ、複数の重要なデータ パイプラインにわたって大幅なバックログの蓄積やグローバル規模でのサービスレベル契約(SLA)違反を引き起こす可能性があります。

モノリシック アーキテクチャから、シャーディングを取り入れたハブ アンド スポーク パターンに移行して、プラットフォームの復元力を確保する方法をご紹介します。

課題: モノリシックのボトルネック

従来の一般的なアーキテクチャでは、すべてのテナントとビジネス ドメインのデータを 1 つの巨大なストリームで処理します。パイプラインが統合されているため、特定のデータベース テナント インスタンスでパフォーマンスの問題が発生すると、バック プレッシャーが発生し、プラットフォーム上の他のすべてのテナントのパフォーマンスが低下します。

痛手となる影響:

  • 100% の影響範囲: 1 つのデータベース障害で、すべての処理が停止する可能性があります。

  • 非効率的なスケーリング: 多くの場合、「最悪のケース」のテナントに合わせてリソースをスケールする必要があるため、無駄な支出が多くなります。

  • SLA の不安定性: 1 つの大規模なテナントがシステム全体を遅延させる可能性がある場合、グローバル規模の SLA を維持することはほぼ不可能です。

解決策: シャーディングによるハブ アンド スポーク アーキテクチャ

この課題を解決するには、処理をルーティング用の「ハブ」と、分離された実行用の「スポーク」に分けます。

<div class="article-module h-c-page">
  <div class="h-c-grid">


<figure class="article-image--large
  
  
    h-c-grid__col
    h-c-grid__col--6 h-c-grid__col--offset-3
    
    
  ">

  
  
    
    <img alt="1" src="https://storage.googleapis.com/gweb-cloudblog-publish/images/1_AstMUiw.max-1000x1000.png" />
    
    </a>
  
</figure>


  </div>
</div>

1. ハブ: ルーター パイプライン

ハブは、トラフィック コントローラとして機能する軽量の Dataflow ジョブです。統合されたソーストピックから読み取り、テナント ID またはビジネス ドメインを解析して、データを分離されたバッファにファンアウトします。これにより、エントリー ポイントがシンプルかつ堅牢になります。

2. バッファ: 永続的な分離

ハブとスポークの間に Pub/Sub トピックを導入します。これらは永続的な緩衝器として機能し、低速なダウンストリーム シンクが元のソースをバックアップするのを防ぎます。

3. スポーク: 分離された実行

1 本の巨大なパイプラインではなく、複数の小規模な Dataflow インスタンスをデプロイして、ワークロードごとに分類します。

  • ティア 1(優先度高): 重要なテナント向けにリソース割り当てを多くした専用パイプライン。

  • 共有ティア: 小規模なテナント向けにパイプラインをグループ化して費用を最適化。

  • ドメイン固有: コードの複雑さを分離するために、複雑なロジック(異なるビジネス ドメインの分離など)に特化したパイプライン。

メリットの比較

機能

モノリシック(従来)

ハブ アンド スポーク

フォールト トレランス

1 つの DB の障害ですべてが停止

障害は特定のスポークに限定

影響範囲

100%

5% 未満(1 つのスポークに限定)

リソースのスケーリング

「最悪のケース」を想定してスケーリング

テナントの負荷ごとに独立したスケーリング

メンテナンス

グローバルなアップデートが全体に影響

他のドメインに影響を与えずに 1 つのドメインを更新

実装に関する活用のヒント

このアーキテクチャへの移行は、単に構成図の中でボックスを移動するだけではありません。最大限の安定性を確保するためには、追加で「スポーク」レベルを最適化することが推奨されます。

  • デッドレター キュー(DLQ)の実装: 単一の SQL 例外によってパイプラインが停止しないようにします。失敗したレコードをストレージ(BigQuery や Google Cloud Storage など)にルーティングして、後で調査します。

  • 厳格な接続プール: データベースには接続制限があります。スレッドセーフなシングルトン パターンを使用し、ワーカーあたりの MaximumPoolSize を低く(1~2 などに)設定して、自動スケーリング中にデータベースを使い果たさないようにします。

  • 非同期 I/O: GroupIntoBatches 変換を使用して書き込みをバッファリングすることで、データベースに起因するレイテンシのトリガーとなることが多い接続オーバーヘッドを削減します。

まとめ

シャーディングのアプローチを採用することで、プラットフォームは「ノイジー ネイバー」が近隣の脅威にならないことを保証できます。このアーキテクチャは、厳格な SLA を維持するために必要な分離を実現すると同時に、独立したスケーリングと、より安全なデプロイを可能にします。

シャーディングによるハブ アンド スポーク アーキテクチャの実装について詳しくは、Dataflow のドキュメントをご覧ください。

- Google Cloud、クラウド データ エンジニア Sri Harshini Donthineni

- Google Cloud、クラウド データ コンサルタント Abdullateef Abdulsalam

Google Cloud Blog (AI & ML)

How AlloyDB ScaNN scales vector search to 10 billion vectors

詳細を表示

To satisfy the demands of enterprise-grade agentic AI applications, underlying vector databases often struggle to scale effectively as modern use cases can scale to billions of vectors.

As a fully managed PostgreSQL-compatible database service, AlloyDB is engineered to handle demanding enterprise workloads. Combining Google's infrastructure with the reliability of commercial databases, it delivers high availability, scalability, and includes a cutting-edge analytical engine, optimal for agentic AI use cases. A key part of this is its ScaNN index, which now operates efficiently at a scale of 10 billion vectors. This was achieved through a major architectural enhancement: an innovative four-level tree (preview) paired with efficient memory usage.

The 10 billion vector scale challenge

Scaling to a 10 billion vector workload presents significant memory and computational challenges. Previous AlloyDB ScaNN tree-based index was limited to two- or three-level tree configurations, and attempting to scale those structures led to several bottlenecks:

  • Increased compute intensity: Larger tree structures demand significantly more operations for both index construction and query traversal.

  • Memory constraints: The sampling processes required for 10 billion vectors can easily exceed the system's available memory capacity.

Solution: Four-level architecture

The introduction of a four-level tree (preview) is the primary innovation in the recent AlloyDB ScaNN release. This architecture, illustrated in Figure 1, employs a top-down strategy to optimize the balance between accuracy and build efficiency. To maintain high performance and mitigate recall loss, the system integrates key enhancements such as Top-K branch, SOAR, centroid adjustment and balanced tree shape.

<div class="article-module h-c-page">
  <div class="h-c-grid">


<figure class="article-image--large
  
  
    h-c-grid__col
    h-c-grid__col--6 h-c-grid__col--offset-3
    
    
  ">

  
  
    
    <img alt="1" src="https://storage.googleapis.com/gweb-cloudblog-publish/images/1_fpfICUj.max-1000x1000.jpg" />
    
    </a>
  
    <figcaption class="article-image__caption "><p>Figure 1. AlloyDB ScaNN four-level tree architecture</p></figcaption>
  
</figure>


  </div>
</div>

This design has two primary benefits:

1. Reduced compute intensity via hierarchical partitioning

The four-level architecture drastically reduces compute intensity by using hierarchical partitioning to restrict the volume of vectors scanned during a query. Instead of traversing a flat or poorly segmented space, the multi-layered hierarchy narrows down the search path exponentially. Figure 2 illustrates the search spaces across different tree levels, demonstrating how structural layering optimizes traversal efficiency:

<div class="article-module h-c-page">
  <div class="h-c-grid">


<figure class="article-image--large
  
  
    h-c-grid__col
    h-c-grid__col--6 h-c-grid__col--offset-3
    
    
  ">

  
  
    
    <img alt="2" src="https://storage.googleapis.com/gweb-cloudblog-publish/images/2_LWwXC70.max-1000x1000.jpg" />
    
    </a>
  
    <figcaption class="article-image__caption "><p>Figure 2. Search space for two-, three- and four-level trees</p></figcaption>
  
</figure>


  </div>
</div>
  • Two-level: Utilizes coarse partitioning to guide queries, resulting in a basic search complexity of O(N1/2).

  • Three-level: Introduces an intermediate layer to further subdivide clusters, narrowing exploration to O(N1/3).

  • Four-level: Implements refined, highly granular partitions that optimize traversal efficiency down to O(N1/4), sufficiently allowing for more than 10-billion vectors.

By dynamically expanding hierarchical layers as the dataset expands, AlloyDB ScaNN maintains ultra-low query latency and avoids computational scale walls from impacting performance.

2. Efficient memory usage

Achieving a 10 billion vector scale requires high memory efficiency. AlloyDB ScaNN uses these strategies to maximize memory management performance:

  • Balanced tree shape construction: The four-level tree utilizes a balanced configuration to circumvent memory limitations that restrict the size of training datasets. This balanced architecture effectively leverages reduced sampling sizes to construct high-fidelity tree partitions.

  • Sampling optimization: When the system encounters memory limitations, it generates a condensed sampling set that considers performance and accuracy. 

Performance test results

By leveraging the innovative four-level tree architecture in our internal tests, we are able to achieve the following performance results:

  • AlloyDB can scale to over 10 billion vectors with its ScaNN index.

  • AlloyDB can deliver <= 51 ms p95 latency and 95% recall at 10 billion vectors with its ScaNN index.

Get started today

Experience AlloyDB ScaNN's four-level tree (preview) architecture today. You can deploy ScaNN for AlloyDB by following our quickstart guide to set up an instance. For optimized, high-speed vector search, refer to the official ScaNN documentation. New users can also explore AlloyDB through our 30-day free trial program. We can’t wait to hear about what you build!

10 questions every startup should answer before moving to production with their AI prototype

詳細を表示

It’s never been easier to start an AI-powered startup on Google Cloud. 

You grab an API key from Google AI Studio at breakfast, paste it into Antigravity, and by lunch you’ll have a nascent prototype of your product.

But it’s not all one straight line to progress. It's common to bump into these three challenges as you build out your stack:

  • A leaked API key racks up a large bill in 48 hours.

  • A "quick" migration from AI Studio to Gemini Enterprise Agent Platform stalls the roadmap for weeks because nobody on the team owns Identity and Access Management (IAM).

  • The launch works, until the app starts returning HTTP 429 Too Many Requests because of default per-project quotas, and there's no clean path to more capacity without paying a premium.

None of these are unique edge cases. . They're  default failure modes of moving fast without a plan, and we've all done it at least once.

Below are the 10 questions every startup should be ready to answer before they scale,  grouped into the three phases where decisions can shape your future: 

  1. Onboard (setting up your own projects and identities right)

  2. Scale (getting more throughput without breaking the bank) 

  3. Govern (keeping costs, keys, and agents from running away).

These ten are scoped to the prototype-to-production transition itself. Each question ends with a short, runnable snippet you can copy into your own project today. Adjacent decisions that matter just as much but aren't specific to that move, your data layer and RAG architecture, CI/CD, network design, are deliberately out of frame here.

Onboard: get the foundation right (in the first hour).

#1 Where should I start: Google AI Studio or Gemini Enterprise Agent Platform?

Both surfaces expose the same Gemini family of models, but they solve different problems.

  • Google AI Studio (with the Gemini Developer API) is the fastest path from an idea to working code. A browser IDE, an API key, a generous free tier, and no cloud project to configure. It's where most ideas should start, and Google's own guidance says as much.

  • Gemini Enterprise Agent Platform (formerly Vertex AI) has the same Gemini models (plus 3rd party and OSS ones)  with enterprise controls around them: IAM and service-account auth instead of raw keys, VPC Service Controls, Cloud Logging and Monitoring, reserved capacity, regional endpoints, and the compliance surface your first enterprise customer's security review will ask about.

The right answer for most startups is both, sequenced deliberately: first prototype in AI Studio, then migrate before you have real users. The danger for startups is treating them as interchangeable solutions, AI Studio's simple key model does not translate to enterprise controls, and Agent Platform's IAM model might look like overkill until the day it saves you from a stolen-credential incident.

It's less work than it sounds like.

The unified google-genai SDK targets both:

code_block
<ListValue: [StructValue([('code', '# Prototype: Google AI Studio, raw API key\r\nfrom google import genai\r\nclient = genai.Client(api_key="YOUR_AI_STUDIO_KEY")\r\n\r\n# Production: GEAP, no key — uses Application Default Credentials (ADC)\r\nfrom google import genai\r\nclient = genai.Client(\r\n vertexai=True,\r\n project="my-startup-prod",\r\n location="us-central1",\r\n)\r\n\r\nresp = client.models.generate_content(\r\n model="gemini-2.5-pro",\r\n contents="Summarize this contract in three bullets.",\r\n)\r\nprint(resp.text)'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7ff5df22c410>)])]>

#2  How do I set up a Google Cloud project without becoming an IAM expert?

The biggest reason startups stall on the migration to Agent Platform isn't the code, it's the operational leap from "here's an API key" to a cloud project with folders, service accounts, org policies, logging, and IAM bindings. If your team doesn't have a dedicated cloud admin, that first project setup can eat a week of engineering time. 

Three moves cut that dramatically:

  1. Use an opinionated project template instead of clicking through the console. The Cloud Setup checklist and the Google Cloud Architecture Framework give you a production-grade folder hierarchy (prod / non-prod / dev), a central logging + monitoring project, Security Command Center turned on, and baseline org policies, without you having to design them from scratch.

  2. Enable the APIs you'll actually use, once. Batch it so you're not doing it project-by-project when you need it. The billing-link step is not optional. Every paid API you're about to enable will refuse to activate on a project with no billing account attached, so we handle that first.

  3. Let Gemini pick the roles, but ask it for the narrow ones. You don't have to memorize the roles reference. In the Grant access dialog, Help me choose roles lets you describe the task in plain language, "this service account needs to call Gemini models and read one Cloud Storage bucket", and get predefined roles back with the reasoning shown. One catch worth knowing on day one: by default it suggests roles that cover common journeys, which usually means a service's Admin, Editor, or Viewer. Those are broader than you want. Say "least privileged" or "narrowest access" in the prompt and it returns granular roles instead. Same amount of typing, considerably smaller blast radius when a credential leaks.

    Sources: Get predefined role suggestions with Gemini assistance

code_block
<ListValue: [StructValue([('code', '# One-shot: create a Vertex-ready project and turn on the services a\r\n# typical AI startup uses.\r\ngcloud projects create my-startup-prod --name="My Startup (prod)"\r\ngcloud config set project my-startup-prod\r\n\r\n# REQUIRED before enabling billing-dependent APIs (aiplatform, run, etc.).\r\n# Use `gcloud billing accounts list` to find your billing account ID.\r\ngcloud billing projects link my-startup-prod --billing-account=012345-6789AB-CDEF01\r\n\r\ngcloud services enable \\\r\n aiplatform.googleapis.com \\\r\n run.googleapis.com \\\r\n artifactregistry.googleapis.com \\\r\n logging.googleapis.com \\\r\n monitoring.googleapis.com \\\r\n secretmanager.googleapis.com \\\r\n cloudbilling.googleapis.com'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7ff5e41e8d50>)])]>

Sources: gcloud services enable reference, · gcloud billing projects link (GA)GE Agent Platform environment setup.

If you're a solo founder, resist the urge to build in your personal GCP account. Create a proper organization or self-owned org first, then create the project inside it. That single decision can make everything else, fromIAM to billing and audit, dramatically easier.

#3 I'm on Google Cloud, how should my code actually authenticate: API keys, service accounts, or user credentials?

There's a hierarchy of safety here, and the easiest option is rarely the right one in production.

  • Raw API keys are fine for local prototyping. They are dangerous in production because they are long-lived, easy to leak into a client bundle or a public repo, and grant unbounded access until you notice.

  • User credentials via OAuth (application default credentials) are best for interactive tools, CLIs, and any code that runs on a developer's laptop.

  • Service accounts with least-privilege IAM roles are the right answer for anything running on a server, in a container, or in a scheduled job.

The pattern you're aiming for is one where your code never sees a key at all. It just calls the Google Auth library, which quietly reads Application Default Credentials (ADC) from the environment,  a short-lived token minted for whichever service account is attached to your Cloud Run service, GKE workload, or Compute Engine VM. You get enterprise-grade auth without writing any auth code.

code_block
<ListValue: [StructValue([('code', '# On a developer laptop\r\ngcloud auth application-default login\r\n\r\n# On a server (Cloud Run, GKE, etc.) — no login, no key file.\r\n# Attach a service account with just the roles the app needs.\r\ngcloud run deploy my-agent \\\r\n --image=us-docker.pkg.dev/my-startup-prod/agents/api:v1 \\\r\n --service-account=agent-runtime@my-startup-prod.iam.gserviceaccount.com \\\r\n --region=us-central1'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7ff5e41eb1d0>)])]>
code_block
<ListValue: [StructValue([('code', '# Application code — notice: no keys, no secrets.\r\nfrom google import genai\r\n\r\nclient = genai.Client(\r\n vertexai=True,\r\n project="my-startup-prod",\r\n location="us-central1",\r\n)'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7ff5e41eaf10>)])]>

Do one last favor to your future self: give that service account the minimum IAM role your workload actually needs,  usually roles/aiplatform.user for calling models, not the broader admin roles. It takes an extra 30 seconds and prevents the credential from becoming a master key if it leaks.

#4 When should I actually stop procrastinating and migrate from AI Studio's API key to Agent Platform's IAM model?

Sooner than you'd like,  and the correct trigger is not when it breaks. It's when any of these is true:

  • Your key has left your laptop (checked into a repo, pasted into a Slack, shipped in a mobile app).

  • You have more than one person on the team who needs to call the API.

  • You're spending more than a few hundred dollars a month.

  • You're about to onboard paying customers.

A potential pitfall that can catch growing startups off guard is simple: a leaked Gemini API key on an account that normally spends $180 a month gets scraped from a public repo and used to run distillation attacks,  accumulating tens of thousands of dollars in charges before the owner even sees the first billing alert. The Google Cloud Shared Responsibility Model is unambiguous: the customer is liable for charges incurred with their own valid credentials.

The migration itself is genuinely smaller than the anxiety around it. In google-genai it's the two-line change shown in #1. What takes real time is the project setup around it, which is exactly why #2 exists.

Practical checklist for cutover day:

code_block
<ListValue: [StructValue([('code', '# 1. Revoke every existing AI Studio key that has ever left a laptop.\r\n# (Go to https://aistudio.google.com/apikey and delete them.)\r\n\r\n# 2. Confirm your production code has no api_key= arguments.\r\ngrep -rn "api_key" src/\r\n\r\n# 3. Enable GEAP and confirm ADC works locally.\r\ngcloud services enable aiplatform.googleapis.com\r\ngcloud auth application-default login\r\npython -c "\r\nfrom google import genai\r\nc = genai.Client(vertexai=True, project=\'my-startup-prod\', location=\'us-central1\')\r\nprint(c.models.generate_content(model=\'gemini-2.5-flash\', contents=\'ping\').text)\r\n"'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7ff5e41ea1d0>)])]>

If step 3 prints a response, you're on Agent Platform.

Scale: get more capacity without paying a premium.

#5 Now that I'm shipping, why on earth am I getting all these HTTP 429 errors, and how do I make them stop?

429 Too Many Requests from Agent Platform almost always means one of two things:

  1. You've hit the Dynamic Shared Quota (DSQ) ceiling for your project's tier. DSQ is a shared pool sized against your project's history,  new projects start with modest limits by design, to prevent abuse across the platform.

  2. You're calling a global endpoint during a global demand spike, competing with worldwide traffic for shared capacity.

The instinctive reaction is to file a quota-increase ticket. You can do that if you must,  but two architectural moves usually solve the problem faster and cheaper.

Pin to a regional endpoint. Over half of startup traffic on Agent Platform defaults to global routing. Pinning to a specific region (say us-central1) sidesteps global contention and typically improves latency at the same time. (One narrow exception, which we'll get to in the next question: if you specifically want Priority PayGo, that feature currently only ships on the `global` endpoint. For everything else, pin regionally.):

code_block
<ListValue: [StructValue([('code', 'from google import genai\r\n\r\n# Global (default): competes against worldwide demand.\r\n# Regional: routes only to the regional cluster, less contention.\r\nclient = genai.Client(\r\n vertexai=True,\r\n project="my-startup-prod",\r\n location="us-central1", # <-- this is the one-line fix\r\n)'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7ff5e41eb790>)])]>

Add real retry and backoff. A 429 is a retryable signal, not a fatal error. Any production client should have exponential backoff with jitter. The modern google-genai SDK ships this behavior built in, but only if you actually enable it. This is easy to overlook. Don't reach for the classic `google.api_core.retry.if_transient_error` decorator you may have seen on older Vertex code. It's designed for the legacy exception classes and does not recognize the new `google.genai.errors.APIError,  so it will silently pass 429s through without retrying. Use the SDK's built-in retry options instead:

code_block
<ListValue: [StructValue([('code', 'from google import genai\r\nfrom google.genai import types\r\n\r\nclient = genai.Client(\r\n vertexai=True, project="my-startup-prod", location="us-central1",\r\n http_options=types.HttpOptions(retry_options=types.HttpRetryOptions(\r\n attempts=5, initial_delay=1.0, max_delay=60.0, exp_base=2.0, jitter=1.0,\r\n http_status_codes=[408, 429, 500, 502, 503, 504],\r\n ))\r\n)'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7ff5e41ea310>)])]>

How do you see this coming?  Preferably not from a user telling you. Agent Platform publishes serving metrics to Cloud Monitoring, and there is a prebuilt dashboard you don't have to assemble: Console → Agent Platform → Dashboard → Model observability. It gives you requests per second, token throughput, first-token latency, and error rates out of the box.

The metric to actually alert on is aiplatform.googleapis.com/publisher/online_serving/model_invocation_count. It carries an error_category label with values of user, system, or capacity. Alerting on capacity isolates genuine throttling from your own bad requests, which a raw 429 count won't do.

One thing worth internalizing, because it trips people up: you cannot build a "warn me at 80% of my quota" alert for Standard PayGo. Under Dynamic Shared Quota there is no fixed per-project number to be at 80% of. A 429 means transient contention for shared capacity, not that you crossed a line. Percent-of-limit alerting only becomes meaningful once you're on Provisioned Throughput, which does expose real limit metrics.

code_block
<ListValue: [StructValue([('code', 'gcloud monitoring policies create --policy-from-file=capacity-alert.yaml'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7ff5e41e83d0>)])]>

Sources: Agent Platform metrics list, Model observability dashboard, RetryOptions sourcecore retry_base.py, genai errors.pyreduce 429 errors, gcloud monitoring policies create, Dynamic Shared Quota.

Follow the Agent Platform rate limits documentation to understand what your project's current ceiling actually is before you assume you've outgrown it. 

#6 Which consumption mode do I pay for: Standard PayGo, Priority PayGo, or Provisioned Throughput? 

Three consumption models, three completely different workload shapes, and three completely different ways to proceed. Picking the right one can help startups see meaningful savings on AI bills. First let’s define them and then see when they are, or aren’t, a good fit:

Standard PayGo (DSQ): Pay per token from a shared pool; cheap, no guarantees.
Priority PayGo: Pay per token at a premium to jump the queue.
Provisioned Throughput (PT): Prepay for reserved capacity; predictable, use it or lose it.

Consumption type

Best for

Watch out for

Standard PayGo (DSQ)

Early-stage, low-QPS, spiky prototype traffic

429s during spikes; no reliability SLO

Priority PayGo

Bursty, revenue-critical traffic that can't tolerate 429s

Roughly 1.8x the standard token price

Provisioned Throughput (PT)

Steady, predictable, high-volume production traffic

Wasted spend if utilization is under ~40%; overflow to PayGo on spikes

The dominant startup mistake is buying PT too early. Usually  this happens the  week after a big launch when it feels like traffic will only ever go up. PT is reserved capacity. You  pay whether you use it or not, and it only starts paying you back once your baseline is genuinely predictable, not just aspirational.

Here’s a pragmatic sequence:

  1. Weeks one through four on Standard PayGo. Use it to measure your real request shape (tokens per minute at p50 and p99, request bursts, batchable vs. real-time split).

  2. When you get your first bad 429 storm, flip on Priority PayGo for the traffic that actually matters. It's a config change, not a purchase order,  nobody in procurement needs to be involved:

code_block
<ListValue: [StructValue([('code', '# Priority PayGo request: use the global endpoint + two extra headers.\r\nfrom google import genai\r\nfrom google.genai import types\r\n\r\nclient = genai.Client(vertexai=True, project="my-startup-prod", location="global")\r\nresp = client.models.generate_content(\r\n model="gemini-2.5-pro",\r\n contents="Rank these support tickets by urgency: ...",\r\n config=types.GenerateContentConfig(\r\n # Priority PayGo headers, per current GEAP docs.\r\n http_options=types.HttpOptions(headers={"X-Vertex-AI-LLM-Request-Type": "shared", "X-Vertex-AI-LLM-Shared-Request-Type": "priority"}),\r\n ),\r\n)'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7ff5e41e9910>)])]>

3. Once you can predict your baseline TPM, buy PT to cover the flat baseline and let anything above it overflow to PayGo. That's the combined pattern Google recommends for exactly this reason. Best of both worlds, not marketing spin.

 Sources: Priority PayGo docs, google-genai HttpOptions source, GEAP REST reference

#7 Which of my requests actually need to be live, and which should be batch jobs?

Most startup workloads are secretly batch jobs pretending to be real-time. Every one you move off the interactive path frees up DSQ headroom for the traffic that genuinely needs to be fast,  the traffic where a user is actually watching a spinner.

Three questions to help you sort your traffic:

  • Does a human have to see the result within a second? That means:  Live inference.

  • Can the user wait a few seconds and see a spinner? That means:  Still live, but a candidate for streaming.

  • Would the user tolerate "we'll email you when it's ready" or "check back in a bit"?  That means: Batch prediction.

Batch prediction on Agent Platform runs in a completely separate queue, does not consume your interactive DSQ, and is typically about half the price of on-demand inference. That's a rare double win: faster live traffic and a lower bill.

code_block
<ListValue: [StructValue([('code', '# Kick off a batch prediction job from a JSONL file in Cloud Storage.\r\n# Each line is one prompt; results land in another Cloud Storage prefix.\r\nfrom google import genai\r\nfrom google.genai import types\r\n\r\nclient = genai.Client(vertexai=True, project="my-startup-prod", location="us-central1")\r\n\r\njob = client.batches.create(\r\n model="gemini-2.5-flash",\r\n src="gs://my-startup-prod-batch/inputs/nightly-summaries.jsonl",\r\n config=types.CreateBatchJobConfig(\r\n dest="gs://my-startup-prod-batch/outputs/",\r\n ),\r\n)\r\nprint(job.name, job.state)'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7ff5e41e90d0>)])]>

Common candidates: nightly document summarization, background classification of new signups, bulk translation, embedding backfills, evaluation runs against your test set. If any of those are on your live path today, moving them is often the single highest-leverage change you can make this week.

Govern: Keep costs, keys, and agents under control.

#8 How do I set spend caps that actually reduce cost, and not just send me polite emails while my bill triples?

Until recently the honest answer was that budgets only notify, and you had to build your own brake pedal. That changed in July. There are now three mechanisms, and you should think of them as layers.

  1. A spend cap budget (Preview). Cloud Billing budgets can now enforce rather than just email. Set a spend cap on a project and, when usage costs cross 100% of the budget, Google pauses the service until you manually lift it. Agent Platform is explicitly on the eligible list, alongside the Gemini API, Cloud Run, and Cloud Run functions. Alerts still fire at 50% and 80%, so the pause isn't a surprise.

Three things to know before you rely on it:

  • Each cap covers one project and one eligible service. It is not account-wide protection. If you want Agent Platform and Cloud Run both capped, that's two caps. 

  • Enforcement is not instant and is based on estimated costs. Overages past the cap are billed as normal, so set the number below your real ceiling. Lifting it is manual, and service resumption can take up to an hour. It also pauses Provisioned Throughput usage, so if you've prepaid for capacity, a cap hit stops that too.

  • It's in Preview as of publication, and the eligible-service list is documented as growing. Check the current list before you design around it.

2. A billing budget with a Pub/Sub trigger that disables billing. Still the right tool when you need blast radius the spend cap can't give you: multiple services at once, an entire project, or a service that isn't eligible yet. When the budget hits a threshold, Pub/Sub fires a Cloud Function that detaches the billing account, which stops all billable activity within minutes. Blunter and more dangerous than the native cap — it can leave resources unrecoverable — so reach for it second, not first. Full walkthrough: Automatically respond to budget notifications.

code_block
<ListValue: [StructValue([('code', '# Sketch: create a budget SCOPED TO ONE PROJECT that publishes to Pub/Sub at 50%, 90%, 100%.\r\ngcloud billing budgets create \\\r\n --billing-account=012345-6789AB-CDEF01 \\\r\n --display-name="my-startup-prod hard stop" \\\r\n --budget-amount=2000USD \\\r\n --filter-projects=projects/my-startup-prod \\\r\n --threshold-rule=percent=0.5 \\\r\n --threshold-rule=percent=0.9 \\\r\n --threshold-rule=percent=1.0,basis=current-spend \\\r\n --notifications-rule-pubsub-topic=projects/my-startup-prod/topics/budget-alerts'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7ff5e41eabd0>)])]>

Sources:  Manage spend cap budgets, Set up programmatic notifications gcloud billing budgets create reference, Cloud Billing budgets concepts, Disable billing with notifications walkthrough, Programmatic notification payload schema.

Two things to get ahead of  for, as the defaults can cause unexpected issues: 

  1. Limit your budget scope: Without --filter-projects, your budget applies to your entire billing account. A spike in any project will trigger the kill switch for everything. 

  2. Deploy locally: The budget notification doesn't specify which project is affected. To ensure the kill switch only affects the intended project, deploy your Cloud Function in the same project you're protecting (e.g., my-startup-prod).

Then wire up a tiny Cloud Function to that topic that calls projects.updateBillingInfo to unlink the billing account when the 100% threshold fires. That is your circuit breaker.

Mechanical ceilings via quota overrides. Even if you never set up the above kill switch, you can cap the rate at which cost can accumulate by setting explicit per-model, per-region quotas below the platform default. If your app never legitimately needs more than 500 requests per minute for gemini-2.5-pro, cap it there in the Cloud Quotas console; a leaked key can't burn what the quota flatly refuses to serve.

#9 Where should I actually keep secrets? (Not in .env files!)

The short answer is: Secret Manager. Not  in environment variables, not in .env files, and never in your repo. Grant read access via IAM only to the service account that needs it.

code_block
<ListValue: [StructValue([('code', '# Store a third-party API key (Stripe, OpenAI, whatever).\r\necho -n "sk_live_xxx" | gcloud secrets create stripe-live-key --data-file=-\r\n\r\n# Grant only the runtime service account access to read it.\r\ngcloud secrets add-iam-policy-binding stripe-live-key \\\r\n --member=serviceAccount:agent-runtime@my-startup-prod.iam.gserviceaccount.com \\\r\n --role=roles/secretmanager.secretAccessor'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7ff5e41e8250>)])]>
code_block
<ListValue: [StructValue([('code', '# Application code fetches it at startup; nothing lives on disk.\r\nfrom google.cloud import secretmanager\r\nsm = secretmanager.SecretManagerServiceClient()\r\nresp = sm.access_secret_version(\r\n name="projects/my-startup-prod/secrets/stripe-live-key/versions/latest"\r\n)\r\nstripe_key = resp.payload.data.decode("utf-8")'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7ff5e41eae10>)])]>

Then two little disciplines that pay for themselves the first time you need them:

  • Rotation on a schedule and on suspicion. Secret Manager versions are cheap; treat them as immutable and roll forward. 

  • Detection when a secret leaks. Secret Manager notifications and Google Cloud's Sensitive Data Protection can catch keys checked into a repo or pasted into a log stream,  before an attacker does.

For any AI application that acts on a user's behalf, calls Gmail on their behalf, reads a Drive folder, hits a third-party SaaS with the user's credentials, do not store a long-lived token. Use OAuth 2.0 with short-lived access tokens and a refresh flow, so that when a user rage-quits or a compromised account gets revoked, the agent loses access at the same time. 

#10  How do I stop my brand new AI agent from doing something it absolutely shouldn't?

An agent that can call tools, browse the web, or execute code needs the same defense-in-depth thinking as any other production service, arguably more, because it makes decisions that neither you nor the model can fully predict in advance.

Four layers, none optional once you have real users:

1. Identity for the agent itself. Give the agent its own service account, scoped only to the resources and tools it genuinely needs,  the exact same least-privilege principle as any other workload. Agent Engine supports first-class agent identity so every action can be attributed to a specific agent instance in your audit logs.

2. Sandboxed code execution. If your agent runs generated code,  a common pattern for data-analysis or "run this Python for me" flows, do not run it in your application process. Use an isolated sandbox so a bad combination can't touch your production data.

code_block
<ListValue: [StructValue([('code', '# Enable server-side code execution inside a sandbox for a request.\r\nfrom google import genai\r\nfrom google.genai import types\r\n\r\nclient = genai.Client(vertexai=True, project="my-startup-prod", location="us-central1")\r\nresp = client.models.generate_content(\r\n model="gemini-2.5-pro",\r\n contents="Compute the correlation between these two columns: ...",\r\n config=types.GenerateContentConfig(\r\n tools=[types.Tool(code_execution=types.ToolCodeExecution())],\r\n ),\r\n)'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7ff5e41ea5d0>)])]>

3. Prompt and response filtering. Model Armor sits in front of your model calls and screens for prompt injection, jailbreaks, sensitive-data exfiltration, and off-brand output,  all of which are essentially guaranteed the moment you have real users being real users.

4. Behavioral monitoring. Security Command Center with threat detection flags anomalies in agent behavior,  a service account suddenly calling an API it's never touched before, an agent reaching out to an unfamiliar external host, an unexpected spike in privileged operations. In near-real-time.

None of these are optional once your agent is acting on behalf of a real user or handling real money.

Your homework, so to speak:

  1. Audit for raw API keys in your repo, your notebooks, and your production runtime. Rotate anything that shouldn't be there.

  2. Move any workload that doesn't need a synchronous response to the Batch API.

  3. Turn on the Model observability dashboard and put one alert on capacity errors, so the next 429 reaches you before it reaches a customer.

  4. Set a spend cap on the project, and keep an eye out for 50% and 80% alerts. If usage crosses 100% of the budget, Google will pause the service until you manually lift it.

Do those four things this week and you're already ahead of most startups shipping AI features. 

Have a scenario you'd like us to cover next? Reach us at Google Cloud for Startups.

Expanding Google Antigravity for enterprise customers

詳細を表示

Since announcing Google Antigravity in Gemini Enterprise Agent Platform at I/O in May, we’ve heard helpful feedback from our customers. Your developers want easy access to coding agents across surfaces. Your enterprise governance team wants security controls and license management. And your finance team wants pooled usage so that no prepaid token ever goes unused. Now, everybody finally gets what they want:

  • Antigravity is available now as part of eligible Gemini Enterprise app subscriptions, including out-of-the-box administrative and spend controls.

  • New IDE extensions let developers use Antigravity in the IDEs of their choice, including VS Code.

Unify AI developer tools and enterprise-grade controls in one subscription

Equipping your developers with advanced agentic tools shouldn't mean managing separate add-on licenses, invoices, billing consoles or security settings. With AI developer tools included in Gemini Enterprise subscriptions, administrators can easily enable Antigravity and Android Studio for users with eligible Gemini Enterprise Standard, Plus, and Standard Emerging Market licenses, and maintain full governance with spend, security, observability and usage metrics consolidated in the Gemini Enterprise admin console.

Unblock your developers while controlling spend

With billing flexibility and cost management tools in Gemini Enterprise, you can ensure your developers have the resources they need while managing costs:

  • Granular spend thresholds: Administrators can set monthly project-level budget caps directly in the Billing console, with additional per-user and team controls rolling out later this year.

  • Pooled quotas: Shared token pools provide flexibility to high-demand teams, preventing purchased quota from sitting idle across the organization.

  • Overage enablement: To maintain continuous developer workflows when pooled quotas are met, administrators can opt into overages with monthly spend caps, smoothly transitioning excess usage to standard consumption-based rates. 

  • Usage metrics: Centralized usage tracking provides visibility into token consumption, API calls, and developer activity, enabling organizations to continuously optimize their AI investments.

<div class="article-module h-c-page">
  <div class="h-c-grid">


<figure class="article-image--large
  
  
    h-c-grid__col
    h-c-grid__col--6 h-c-grid__col--offset-3
    
    
  ">

  
  
    
    <img alt="Gif 1" src="https://storage.googleapis.com/gweb-cloudblog-publish/original_images/Gif_1_vK8C2Nf.gif" />
    
    </a>
  
</figure>


  </div>
</div>

Safeguard your organization’s code and data with built-in privacy and security

Gemini Enterprise subscriptions bring Google Antigravity under Google Cloud’s standard security and compliance protections. Administrators and IT teams can set clear boundaries around workspace access, enable full audit logging, and enforce data privacy from a single console:

  • Configurable security policies: Enforce security and compliance controls, such as workspace sandboxing, and browser and MCP server access, to help ensure AI agents operate safely within authorized enterprise environments.

  • Central audit logging: Enable comprehensive audit logging with a single toggle, capturing prompts, agent responses, and metadata for compliance reporting.

  • Data privacy: Maintain data ownership under Google Cloud’s Terms of Service, ensuring all agent activity executes strictly within your secure cloud boundary.

Bring agentic coding directly into your team’s preferred development environments 

Starting today your developers can use Antigravity across the surfaces they already know and use — including Visual Studio Code, Visual Studio (preview), Jetbrains (preview) and Zed IDEs (preview)  via the new IDE extensions as well as the Antigravity 2.0 desktop app, and the Antigravity CLI.

Throughout all surfaces, administrators can enforce corporate identity standards while removing setup friction for technical teams via native support for Workforce Identity Federation (WIF) and Application Default Credentials (ADC).

<div class="article-module h-c-page">
  <div class="h-c-grid">


<figure class="article-image--large
  
  
    h-c-grid__col
    h-c-grid__col--6 h-c-grid__col--offset-3
    
    
  ">

  
  
    
    <img alt="VSCode IDE Plugin" src="https://storage.googleapis.com/gweb-cloudblog-publish/original_images/VSCode_IDE_Plugin_gif_v3_720p.gif" />
    
    </a>
  
    <figcaption class="article-image__caption "><p>AGY IDE extension demo</p></figcaption>
  
</figure>


  </div>
</div>

What our customers are saying

From rapid code generation to end-to-end task automation, Google Antigravity is giving engineering teams the momentum of cutting edge AI development backed by the stability, governance, and scale of Google Cloud. Here is how leading enterprise customers and partners are driving measurable outcomes in production:

“Deploying Antigravity in Gemini Enterprise allows Accenture to arm our engineers with Google DeepMind’s premier technology on the secure, trusted foundation of Google Cloud. Abstracting away operational complexity ensures our teams don't have to choose between developer speed and enterprise-grade governance — freeing them to deliver high-velocity engineering and transformative value for our clients.” — Chetna Sehgal, Global Practice Lead, Accenture Google Business Group

“At AirAsia and across the Group, we’re all about empowering our people. Bringing highly capable Gemini models directly into our daily workflows with Antigravity 2.0 does exactly that. We are putting the most advanced AI capabilities into the hands of our entire workforce, from software engineering to finance, marketing, legal, HR and much more. This empowers both our developers and critical back-office teams to innovate at an unprecedented pace and drive proven time-savings across the board.” — Nikunj Shanti, CTO, AirAsia Next

“Enterprises are moving beyond AI experimentation and expecting measurable business outcomes. With Gemini Enterprise and next-generation developer tools like Antigravity 2.0 and Antigravity CLI, we see significant opportunities to further embed agentic AI, particularly the advanced reasoning capabilities of Gemini models, directly into software delivery workflows. This goes beyond productivity as it enables faster decision-making, higher code quality, and reduced technical debt at scale. What stands out is how these capabilities are helping our teams evolve from writing code to orchestrating outcomes, strengthening every phase of the software development lifecycle while scaling innovation securely and responsibly.” — Rakesh Aerath, President, Asia Pacific Global Delivery Centers of Excellence, CGI

“Every developer workflow is unique, and agentic AI should adapt to the engineer, not the other way around. With Google Antigravity supported across developers' preferred IDEs, the desktop app, and the CLI, Cognizant can seamlessly embed agentic engineering across our global delivery centers. It gives our teams the freedom to choose their preferred surface while delivering high-velocity, secure software for our clients.”— Rajesh Varrier, President, Global Operations and Chairman & Managing Director, Cognizant India

“Antigravity has played a key role in advancing our AI-first strategy at Datamatics. Over the last few months, our teams have used it to rapidly build and deploy multiple applications, accelerate solution development, and embed AI into core business processes. From AI Impact Hub to analytics and sales enablement solutions, it has helped us move beyond experimentation to real execution, delivering measurable business outcomes while enabling teams to innovate faster and at scale.”— Vijay Venkatachalam, Vice President, Information Systems Group (ISG), Datamatics

"Embedding Google Antigravity's autonomous capabilities directly into Gemini Enterprise development environments allows our internal and Forward Deployed Engineering teams to automate complex tasks. Supported by the platform’s new FinOps and governance controls, our teams can confidently focus on orchestrating high-value, secure and cost-efficient outcomes for our clients at scale."  — Faruk Muratovic, US AI & Engineering Strategy and Services Leader, Deloitte

“Adopting Antigravity places Wipro at the leading edge of the AI-driven software development lifecycle. Combining Antigravity 2.0, CLI, and the new IDE extensions with a seamless developer experience and superior code accuracy fits naturally into our AI-first engineering strategy. Working with Google Cloud allows us to accelerate software delivery and bring next-generation value to our global enterprise clients.”  — Debashish Ghosh, Vice President and Global Head, Google Partnership, Wipro

Get started with Antigravity in Gemini Enterprise

Google Antigravity in Gemini Enterprise is available today for eligible Gemini Enterprise Standard, Plus, and Standard Emerging Market licenses, with broader support coming soon.

  • For administrators: Visit the enterprise setup guide to enable AI Developer tools for Google Antigravity and Android Studio. 

  • For developers: Start building with Antigravity 2.0, the Antigravity CLI, or your preferred IDEs via Antigravity IDE extensions.