GCP News - 2026-08-26
2026-08-26
最終更新: 2026-08-27 21:31:31 JST
Google Cloud Release Notes
August 26, 2026
- Link: https://docs.cloud.google.com/release-notes#August_26_2026
- Published: 2026-08-26 16:00:00
- Fetched: 2026-08-27 21:31:25
詳細を表示
Apigee hybrid
Announcement
v1.14.8
On August 26, 2026 we released an updated version of the Apigee hybrid software, v1.14.8.
- For information on upgrading, see Upgrading Apigee hybrid to version v1.14.8.
- For information on new installations, see The big picture.
Security
| Bug ID | Description |
|---|---|
| N/A | Security fixes for apigee-asm-ingress. This addresses the following vulnerabilities: |
| N/A | Security fixes for apigee-asm-istiod. This addresses the following vulnerabilities: |
| N/A | Security fixes for apigee-connect-agent. This addresses the following vulnerabilities: |
| N/A | Security fixes for apigee-fluent-bit. This addresses the following vulnerabilities:
|
| N/A | Security fixes for apigee-hybrid-cassandra. This addresses the following vulnerabilities: |
| N/A | Security fixes for apigee-hybrid-cassandra-client. This addresses the following vulnerabilities: |
| N/A | Security fixes for apigee-mart-server. This addresses the following vulnerabilities: |
| N/A | Security fixes for apigee-mint-task-scheduler. This addresses the following vulnerabilities: |
| N/A | Security fixes for apigee-operators. This addresses the following vulnerabilities: |
| N/A | Security fixes for apigee-prom-prometheus. This addresses the following vulnerability: |
| N/A | Security fixes for apigee-prometheus-adapter. This addresses the following vulnerabilities: |
| N/A | Security fixes for apigee-redis. This addresses the following vulnerabilities: |
| N/A | Security fixes for apigee-runtime. This addresses the following vulnerabilities: |
| N/A | Security fixes for apigee-synchronizer. This addresses the following vulnerabilities: |
| N/A | Security fixes for apigee-watcher. This addresses the following vulnerabilities: |
BigQuery
Security
An Improper Input Validation vulnerability was discovered in the JDBC driver in BigQuery Data Transfer Service versions prior to May 1, 2026. An authenticated attacker could use crafted JDBC connection string parameters to achieve remote code execution in the connector container and escalate privileges in the tenant project. For more information, see the GCP-2026-056 security bulletin.
Feature
You can now view real-time logs for your Python UDFs in Cloud Logging. This feature is generally available.
Cloud Load Balancing
Feature
SSL policy cross-project referencing is now available for Application Load Balancers and proxy Network Load Balancers in Preview. You can use cross-project referencing to define and maintain a central SSL policy in an administrative project and reference it from target HTTPS proxies or target SSL proxies in different projects.
Cross-project referencing is supported for global and regional SSL policies. You can use cross-project referencing with the following load balancers:
- Global external Application Load Balancer
- Regional external Application Load Balancer
- Cross-region internal Application Load Balancer
- Regional internal Application Load Balancer
- Global external proxy Network Load Balancer
For more information, see Cross-project SSL policy referencing.
Cloud Logging
Change
VM Extension Manager extension policies for the Ops Agent are Generally Available (GA). Extension policies provide zonal and project-wide Ops Agent installation, version upgrades, and configuration management. For more information, see Install and manage the Ops Agent by using VM Extension Manager policies.
Cloud Monitoring
Change
VM Extension Manager extension policies for the Ops Agent are Generally Available (GA). Extension policies provide zonal and project-wide Ops Agent installation, version upgrades, and configuration management. For more information, see Install and manage the Ops Agent by using VM Extension Manager policies.
Cloud Service Mesh
Announcement
1.29.7-asm.2 is now available for in-cluster Cloud Service Mesh.
For details on upgrading Cloud Service Mesh, see Upgrade Cloud Service Mesh. Cloud Service Mesh 1.29.7-asm.2 uses Envoy v1.35.14.
This release resolves the security vulnerabilities listed in Security Bulletin GCP-2026-057.
Fixed
Patch 1.29.7-asm.2 contains the fix for the following platform CVEs:
| CVE | Proxy | Control Plane | Distroless | CNI | Severity |
|---|---|---|---|---|---|
| CVE-2026-5704 | Yes | Yes | No | Yes | Medium (5.5) |
Announcement
1.28.10-asm.24 is now available for in-cluster Cloud Service Mesh.
For details on upgrading Cloud Service Mesh, see Upgrade Cloud Service Mesh. Cloud Service Mesh 1.28.10-asm.24 uses Envoy v1.36.10.
This release resolves the security vulnerabilities listed in Security Bulletin GCP-2026-057.
Fixed
Patch 1.28.10-asm.24 contains the fix for the following platform CVEs:
| CVE | Proxy | Control Plane | Distroless | CNI | Severity |
|---|---|---|---|---|---|
| CVE-2026-5704 | Yes | Yes | No | Yes | Medium (5.5) |
Announcement
1.27.9-asm.34 is now available for in-cluster Cloud Service Mesh.
For details on upgrading Cloud Service Mesh, see Upgrade Cloud Service Mesh. Cloud Service Mesh 1.27.9-asm.34 uses Envoy v1.35.14.
This release resolves the security vulnerabilities listed in Security Bulletin GCP-2026-057.
Fixed
Patch 1.27.9-asm.34 contains fixes for the following platform CVEs:
| CVE | Proxy | Control Plane | Distroless | CNI | Severity |
|---|---|---|---|---|---|
| CVE-2026-10536 | Yes | Yes | No | Yes | Low (9.8) |
| CVE-2026-42151 | No | No | No | Yes | High (7.5) |
| CVE-2026-42154 | No | No | No | Yes | High (7.5) |
| CVE-2026-40179 | No | No | No | Yes | Medium (6.1) |
| CVE-2026-44903 | No | No | No | Yes | Medium (6.1) |
| CVE-2026-5704 | Yes | Yes | No | Yes | Medium (5.5) |
Feature
For clusters using the TRAFFIC_DIRECTOR implementation, configuring the trace
sampling rate with randomSamplingPercentage with the Telemetry API is now
supported in the Rapid release channel. For more information, see
Accessing Cloud Trace.
Gemini Enterprise
Feature
Gemini Enterprise: Updates to A2UI Material catalog component properties
The A2UI component gallery reference has been updated to reflect the latest A2UI version v0.9 Material catalog component properties and schema:
MaterialButton: Added support for Material 3 styling properties (appearance,disableRipple, andextended), along with new variants (icon,fab, andmini-fab). Obsoletecolorand ARIA description properties have been removed.- Validation
checks: Replaced the staticrequiredboolean property across input components (MaterialCheckbox,MaterialDatepicker,MaterialInput,MaterialSelect, andMaterialTimepicker) with the reactive validationchecksrule array. MaterialIconandMaterialChips: Added thetooltipproperty onMaterialIconand theactionproperty onMaterialChips.- Layout and input types: Documented all supported
justifyalignment values forMaterialColumnandMaterialRow, and updated allowed input types forMaterialInput.
For more information, see A2UI component gallery reference.
Google Kubernetes Engine
Feature
In GKE version 1.36 and later, GCPAuthzPolicy and GCPAuthzExtension resources for GKE Gateway are now available in Preview. You can use these resources to enforce identity-based access control and zero-trust authorization on incoming traffic at the Gateway layer. These capabilities are supported on the following GatewayClasses:
- gke-l7-global-external-managed
- gke-l7-regional-external-managed
- gke-l7-rilb
For more information, see Configure the GCPAuthzExtension resource.
Google SecOps
Feature
[Spotlight Feature] Mandiant Frontline Threats rule packs
Curated Detections has been enhanced with additional Mandiant Frontline Threats detections for Linux, MacOS, and Google Cloud. The following rule packs have been added to the Content Hub:
- Mandiant Frontline Threats for Linux
- Mandiant Frontline Threats for MacOS
- Mandiant Frontline Threats for Google Cloud
Google SecOps SIEM
Feature
[Spotlight Feature] Mandiant Frontline Threats rule packs
Curated Detections has been enhanced with additional Mandiant Frontline Threats detections for Linux, MacOS, and Google Cloud. The following rule packs have been added to the Content Hub:
- Mandiant Frontline Threats for Linux
- Mandiant Frontline Threats for MacOS
- Mandiant Frontline Threats for Google Cloud
Knowledge Catalog
Feature
Knowledge Catalog support for importing metadata from dbt Core and MetricFlow is available in Preview.
You can use the gcloud alpha dataplex dbt metadata-jobs command to extract and
import technical, semantic (MetricFlow), operational, data quality, and lineage
metadata from dbt Core artifacts into Knowledge Catalog.
For more information, see Import metadata from dbt Core and About metadata connectors.
Oracle on Google Cloud Compute
Feature
Oracle on Google Cloud Compute supports running Oracle workloads on Compute Engine's M4N machine series that provides leading block storage performance with Hyperdisk Extreme. For more information, see M4N machine series and Available regions and zones.
This feature is Generally Available (GA).
Feature
Oracle on Google Cloud Compute offers in-depth documentation that describes how to deploy Oracle AI Database workloads using Google Cloud NetApp Volumes. For more information, see Overview of Oracle Database deployment using NetApp Volumes.
Google Cloud Blog
Bringing gVisor sandboxes to distributed Ray clusters
- Link: https://cloud.google.com/blog/products/containers-kubernetes/gvisor-sandboxes-for-ray-clusters-on-gke/
- Published: 2026-08-26 01:00:00
- Fetched: 2026-08-27 21:31:27
詳細を表示
The reinforcement learning (RL) ecosystem is rapidly adopting Ray as the unified compute runtime for complex post-training workflows. Across Google Cloud, we see customers using Ray for workloads ranging from multimodal data pipelines to frontier RL. But as agentic and reasoning models evolve, a critical bottleneck has emerged: orchestrating secure, isolated sandboxes at scale to safely execute dynamic rollouts, code generation, and multi-turn tool interactions. Today, in partnership with Anyscale, we are excited to introduce an experimental library for Ray that leverages agentic AI technologies being developed at Google to bring native, high-performance sandboxing directly into distributed Ray clusters.
Sandboxes as Ray Primitives
Ray has become a common runtime for orchestrating post-training workloads. Frameworks including veRL, NeMo-RL, SLIME, MILES, and SkyRL already use Ray to coordinate distributed trainers, inference engines, rollout workers, and other components.
When we designed Ray Sandboxing, an important goal was to make it fit naturally into the existing Ray programming model rather than introduce a separate abstraction for isolated execution. A sandbox has many of the same properties as other resources managed by Ray: it needs to be placed on a machine, assigned resources, created and destroyed, recovered from failures, and scaled with the surrounding workload. This led us to represent each high-level sandbox through a Ray Actor:
<div class="article-module h-c-page">
<div class="h-c-grid">
<figure class="article-image--large
h-c-grid__col
h-c-grid__col--6 h-c-grid__col--offset-3
">
<img alt="image1" src="https://storage.googleapis.com/gweb-cloudblog-publish/images/image1_SrQumpQ.max-1000x1000.png" />
</a>
</figure>
</div>
</div>
The Ray scheduler decides which node should run a sandbox and reserves the corresponding CPU and memory resources. The sandbox Actor manages its lifecycle, while gVisor provides the isolated execution environment on that node.
Starting in Ray 2.58, framework authors and researchers can manage sandboxed environments using the same Ray APIs and patterns they already use for the rest of their workload. For example:
- code_block
- <ListValue: [StructValue([('code', 'import ray\r\nfrom ray.experimental import sandbox\r\n\r\nray.init()\r\n# Create a gVisor sandbox environment and return an actor handle for a proxy actor\r\nsb = sandbox.create(\r\n cpu=1.0,\r\n memory="512Mi",\r\n image="python:3.12-slim"\r\n)\r\n# Execute code inside the sandbox\r\nresult = ray.get(sb.exec.remote("python -c \'import sys; print(sys.version)\'"))\r\nprint(result.stdout)'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7f59306cce50>)])]>
This creates a gVisor sandbox from an OCI-compatible image and returns a Ray Actor handle. Calls to exec are normal Ray Actor calls, so the sandbox can live anywhere in the cluster. The created actor is a proxy that will forward the operations to gVisor.
The sandbox API covers the basic lifecycle needed by agentic workloads:
-
Create environments from OCI container images
-
Set CPU and memory limits
-
Configure environment variables, working directories, and networking
-
Execute commands
-
Read, write, upload, and download files
-
Inspect sandbox state
-
Terminate or delete environments.
For lower-level use cases, SandboxRuntime provides direct access to local gVisor sandboxes and lets users modify the OCI specification before it is handed to gVisor. Here is an example how this API can be used to build a pool of local sandboxes inside of an actor:
- code_block
- <ListValue: [StructValue([('code', 'import ray\r\nfrom ray.experimental.sandbox.runtime import SandboxRuntime\r\n\r\n@ray.remote\r\nclass SandboxPool:\r\n def __init__(self, size: int = 3, image: str = "python:3.10-slim"):\r\n self.runtime = SandboxRuntime()\r\n self.sandboxes = [\r\n self.runtime.create(image=image, memory="512Mi")\r\n for _ in range(size)\r\n ]\r\n\r\n def run_command(self, index: int, command: str):\r\n return self.runtime.exec(self.sandboxes[index], command)\r\n\r\n def close(self):\r\n for sb_id in self.sandboxes:\r\n self.runtime.delete(sb_id)\r\n\r\n# Deploy an actor managing a pool of local sandboxes\r\npool = SandboxPool.remote(size=3)\r\nresult = ray.get(pool.run_command.remote(0, "python3 -c \'print(\\"Hello from pool!\\")\'"))\r\nprint(result.stdout)\r\nray.get(pool.close.remote())'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7f5930883950>)])]>
Why gVisor?
Running model-generated code means treating the code inside the environment as untrusted. Ray Sandboxing uses gVisor, Google's open-source application kernel, as its initial sandbox runtime. gVisor implements a substantial portion of the Linux system-call interface in userspace, putting an additional isolation boundary between workloads and the host kernel. It is OCI-compatible, works with standard container images, and does not require exposing a Docker daemon or host Docker socket to the sandbox.
This combination is particularly useful for agentic workloads: environments remain lightweight enough to create dynamically while providing stronger isolation than executing generated code directly in ordinary containers. gVisor also provides sub-second sandbox startup and low per-sandbox memory overhead, making it possible to use sandboxes as relatively fine-grained distributed resources.
In future versions of Ray, we plan to extend support to other sandboxing runtimes such as Agent Substrate or Kata Containers.
Try Ray sandboxing on GKE
Check out the Ray documentation to learn more about Ray Sandboxes. To try out these sandboxing capabilities on GKE, head over to the Ray sandboxing User Guide. Have feedback or ideas? Join the discussion on the GitHub issue to collaborate on the future of Ray for reinforcement learning.
Your chance to start building AI agents from the absolute basics
- Link: https://cloud.google.com/blog/topics/developers-practitioners/your-chance-to-start-building-ai-agents-from-the-absolute-basics/
- Published: 2026-08-26 18:07:00
- Fetched: 2026-08-27 21:31:27
詳細を表示
Have you been hearing a lot about "AI agents" lately but aren't sure how to actually start building them? You don't need a background in machine learning or years of software experience to get started. The best way to learn is by doing, which is why we built Agent Valley.
Agent Valley is a free, 5-week live learning series designed to take you from scratch to building your very own hands-on agent systems. And instead of staring at boring terminal lines, you’ll be building and playing inside a tiny, low-poly virtual world!
Meet your instructor
You’ll be learning directly from Annie Wang, one of our top Google DevRel Engineers. She designed this course from the ground up to be fully hands-on, interactive, and beginner-friendly. If you want to learn how AI systems are built by the people actually designing them at Google, this is your chance.
How we'll learn together
You’ll learn by building in a split-screen workspace on your laptop. On Day 1, you'll describe and summon a custom low-poly companion that serves as your play character and save file. As you guide your companion through the valley's five districts, a live Runtime Inspector sits right beside the game, showing you exactly what the AI is thinking, deciding, and costing in real-time. Setup is completely zero-stress. Google will provide the environment for running these exercises, so you can dive straight into building.
Agent 101 Live with 5 modular sessions (Jump in anytime!)
-
Week 1: The Summoning Grove (CONTROL) · Get started by summoning your companion and learning how to keep its memory and traits consistent across a conversation.
-
Week 2: The Buildyard (DECOMPOSE) · Learn how to break a big project down so multiple AI assistants can work together in parallel without stepping on each other's toes.
-
Week 3: Market Street (COORDINATE) · Open up a virtual shop! You'll learn how to write reliable code so transactions and returns go smoothly without crashing.
-
Week 4: The Archive (REMEMBER) · Give your companion a memory. Learn how to help your agent remember past details without getting confused or making things up.
-
Week 5: The Night Market (LIVE) · The grand finale. Learn how to make your agent react live to events in the world (like fireworks or stage lights) while keeping the system fast and affordable.
<div class="article-module h-c-page">
<div class="h-c-grid">
<figure class="article-image--large
h-c-grid__col
h-c-grid__col--6 h-c-grid__col--offset-3
">
<img alt="agent-valley-roadmap-2160x2700" src="https://storage.googleapis.com/gweb-cloudblog-publish/images/agent-valley-roadmap-2160x2700.max-1000x1000.png" />
</a>
</figure>
</div>
</div>
Join the livestream
- 5 Tue starting Sep 1 · 10:00 AM (Pacific Time)
- Anyone new to AI agents who wants to learn by coding and playing.
- RSVP Here: goo.gle/agent101
Dynamic capacity management for AI infrastructure
- Link: https://cloud.google.com/blog/topics/ai-infrastructure/best-practices-for-dynamic-capacity-management/
- Published: 2026-08-26 22:30:00
- Fetched: 2026-08-27 21:31:27
詳細を表示
The internet connected billions of people and mobile devices, putting computers in every hand. Now, we’re in the middle of the next big technology shift, deploying millions of autonomous AI agents to work alongside employees and end users. Today, we announced new FinOps controls for Gemini Enterprise to help organizations manage project-level AI spend and eliminate token shock. But the sheer scale of the agentic era is placing new constraints at every layer of the stack, including infrastructure. AI workloads are notoriously difficult to architect, resource-intensive, and bursty, which can also lead to scaling bottlenecks and large pools of underutilized — or misutilized — compute resources.
Organizations need insights to help them extract more value from their infrastructure investments. In this blog, we outline best practices for dynamic capacity management — scheduling and utilization strategies to help you run enterprise and AI applications on a single, flexible foundation with predictable cost and performance. These capabilities are designed to augment our on-demand, Spot and committed use discount (CUD) consumption models, which provide flexible pricing and discounting for your workloads. Let’s jump in.
Here's a quick summary
Three ways you can implement dynamic capacity management:
-
Schedule capacity for planned events. Schedule mission-critical resources (GPUs, TPUs and select VM families) ahead of planned events using calendar mode, or optimize costs for batch jobs with flexible start times using flex-start mode in Dynamic Workload Scheduler. Once you obtain the capacity, those resources are guaranteed for the specified duration.
-
Maintain service continuity by creating a fallback plan for every application. Define automated, prioritized hardware fallback lists using managed instance groups (MIGs) so your apps automatically pivot to the next approved compute option when your preferred option isn’t available.
-
Automate your entire capacity management lifecycle on a single, adaptive control plane. Google Kubernetes Engine (GKE) provides an agent-native environment to orchestrate the entire process — from fallback lists using Custom ComputeClasses, to granular hardware slicing with dynamic resource allocation, so agents can rapidly spin up in secure sandboxes and containers while it dynamically reallocating resources on the fly.
Why architectural flexibility matters
Ninety percent of enterprises want to deploy agents within the next three years, but only 17% of IT leaders feel confident their current IT setup can handle the load. Because these workloads have unique performance needs, organizations are racing to adopt specialized infrastructure, including accelerators (GPUs, TPUs) and CPUs with customized compute, memory, and storage ratios. However, agents also require access to enterprise applications and databases — often at a volume and scale that vastly exceeds typical human usage. Handling the intense demands of both agents and the applications they interact with requires a dynamic infrastructure. Infrastructure teams can leverage custom-designed processors like Google’s Axion to meet these needs, but hardware isn’t a complete solution. They also need ways to use that infrastructure wisely, solving execution inefficiencies to enable more flexibility across the stack.
How to overcome infrastructure constraints
Achieving this kind of flexibility requires a two-pronged approach: securing resources for the demand you can predict, and building automation to respond to the demand you can't. Combining the two, you can preschedule capacity for planned events and your infrastructure can adapt to unexpected changes without manual intervention.
1. Schedule capacity for planned events
You can secure mission-critical capacity ahead of scheduled milestones, offline training, or anticipated demand surges using Dynamic Workload Scheduler. By scheduling the resources you need up front, you optimize your spend and ensure you get access to the compute resources you need. Dynamic Workload Scheduler supports hardware accelerators (TPUs and GPUs) and select CPUs with two distinct modes:
-
Flex-start mode: Use this for latency-tolerant workloads like batch processing, model training, or offline fine-tuning. Instead of requiring resources immediately, you submit a defined duration request and the system intelligently queues your job, provisioning the resources as soon as capacity becomes available. This maximizes cost-efficiency and drastically improves your ability to obtain high-demand accelerators.
-
Calendar mode: Use this for mission-critical, time-bound events like a major product launch, a scheduled migration, or a seasonal traffic surge. By specifying the exact start and end dates of your event, you create a future reservation. This guarantees the requested capacity will be available when the event begins.
<div class="article-module h-c-page">
<div class="h-c-grid">
<figure class="article-image--large
h-c-grid__col
h-c-grid__col--6 h-c-grid__col--offset-3
">
<img alt="1" src="https://storage.googleapis.com/gweb-cloudblog-publish/images/1_SEuNvWu.max-1000x1000.png" />
</a>
</figure>
</div>
</div>
2. Maintain service continuity by creating a fallback plan for every application
Not every spike in traffic is predictable. You also need to plan for unexpected traffic from, say, a breaking news cycle or a sudden market shift that drives a surge in user activity. To help your services get the resources they need without interruption, you need a fallback plan — an automated, prioritized sequence of acceptable hardware configurations. This strategy:
-
Decouples your workloads from a single VM shape, size, or configuration. This allows them to run without manual intervention if your preferred option is unavailable
-
Allows you to execute a progressive tech refresh by adopting the newest VM generations as your primary choice while keeping older generations as an automatic fallback option.
If you run non-containerized workloads on Google Compute Engine, you can dynamically manage capacity with instance flexibility in managed instance groups (MIGs) and bulk VM creation. Instance flexibility lets you specify multiple machine types for your VM instances rather than being limited to a single machine type.
How it works: If your preferred machine type is temporarily unavailable, the MIG automatically provisions a compatible alternative from your list based on real-time capacity. When combined with location flexibility — by specifying multiple zones your MIGs can search within a region — you can drastically improve your provisioning success rate. If your MIGs use Spot VMs, Compute Engine automatically integrates with Spot capacity signals to prioritize machine types that offer longer estimated uptimes and lower risk of pre-emption.
<div class="article-module h-c-page">
<div class="h-c-grid">
<figure class="article-image--large
h-c-grid__col
h-c-grid__col--6 h-c-grid__col--offset-3
">
<img alt="2" src="https://storage.googleapis.com/gweb-cloudblog-publish/images/2_lW2ljtl.max-1000x1000.png" />
</a>
</figure>
</div>
</div>
You can also extend instance flexibility to your block storage layer by setting baseline disk defaults and configuring disk overrides so your storage adapts when a VM falls back to a different machine type.
How it works: Most of the time you can simply rely on our default options, omitting ‘disk type’ from the instance template entirely. However, for data disks that will outlive their associated VMs, it’s possible to enable a fast, durable Hyperdisk across multiple VM generations.
While Compute Engine provides instance flexibility for organizations working with virtual machines, GKE goes a step further and automates the entire capacity lifecycle from a single control plane. With GKE custom ComputeClasses, platform teams can design multi-dimensional fallback lists, automatically combine different VM machine families, sizes, and ratios, scale across multiple zones, and shift between on-demand and Spot VMs. By using Dynamic Workload Scheduler as a capacity target, and custom ComputeClasses to define the policy and priority, you can fully automate the capacity management lifecycle.
How it works: Once you’ve set up ComputeClasses, GKE automatically detects when a preferred node configuration is unavailable and falls back to your pre-approved alternative options in order of priority. When active migration is enabled, GKE gracefully migrates workloads back to higher-priority node configurations as capacity becomes available. For short-lived disks such as boot disks, GKE dynamically picks the right defaults based on the instance family. However, for long-term disks that will outlive the VM, you can use Hyperdisk.
<div class="article-module h-c-page">
<div class="h-c-grid">
<figure class="article-image--large
h-c-grid__col
h-c-grid__col--6 h-c-grid__col--offset-3
">
<img alt="3" src="https://storage.googleapis.com/gweb-cloudblog-publish/images/3_bBzgKas.max-1000x1000.png" />
</a>
</figure>
</div>
</div>
Another GKE feature, dynamic resource allocation, helps eliminate wasteful, all-or-nothing hardware assignments by letting developers define advanced rules that dictate how resources are consumed.
How it works: Instead of claiming an entire GPU or TPU, your application specifies its exact parameters — such as total memory or number of cores — and the system allocates the perfect slice of hardware, helping to maximize utilization and reduce costs.
<div class="article-module h-c-page">
<div class="h-c-grid">
<figure class="article-image--large
h-c-grid__col
h-c-grid__col--6 h-c-grid__col--offset-3
">
<img alt="4" src="https://storage.googleapis.com/gweb-cloudblog-publish/images/4_RtJb1xh.max-1000x1000.png" />
</a>
</figure>
</div>
</div>
Take the next step toward dynamic infrastructure
Scaling AI shouldn’t mean linearly scaling your infrastructure budget or accumulating more tech debt. As these examples show, the right tools can help you overcome constraints and dramatically alter the value you get from your compute investments. Here are three steps to get started:
-
Audit your workloads for immediate cost-savings: Identify any applications currently tightly coupled to a single VM family, machine type, or availability zone, and map out viable alternative hardware shapes. Look beyond your existing configurations to evaluate new compute options that might better serve or act as alternatives based on your workload-level objectives. Then use Compute Engine MIGs, bulk VM creation or GKE Custom ComputeClasses to adopt them automatically, integrating them into your fallback lists.
-
Commit to a minimum spend for deeply discounted prices: Receive automatic discounts for sustained use, or up to 63% off when you sign up for Compute flexible committed use discounts, where your discount is tied to the resources you use regardless of the specific machine type or location.
-
Engage your account team: Reach out to your Google Cloud account team to craft a tailored capacity management strategy and configure your automated fallback lists.
FinOps for the AI era: New flexible billing and cost controls for agents
- Link: https://cloud.google.com/blog/products/ai-machine-learning/flexible-billing-and-cost-controls-for-agents-on-google-cloud/
- Published: 2026-08-26 22:30:00
- Fetched: 2026-08-27 21:31:27
詳細を表示
Editor's note: A product image was updated after initial publication.
As AI takes on more complex work, business leaders face a new challenge: enabling rapid innovation using agents while protecting their margins and budgets. To get a real return on AI, financial operations (FinOps) and cost management must evolve alongside technology, giving you clear visibility, proactive cost controls, and flexible payment models that fit your needs.
That’s why today we’re introducing expanded billing flexibility and new cost management tools for agent workloads across Gemini Enterprise and developer tools like Google Antigravity in Gemini Enterprise and Android Studio.
-
Flexible payment options: You can mix our existing, predictable per-user seat subscriptions with a new pay-as-you-go option in Gemini Enterprise app that lets you run agent workloads without hitting quota limits mid-task.
-
Developer access, one place to manage your AI: Google Antigravity and Android Studio AI use is now included in your Gemini Enterprise subscription (available for select customers and rolling out broadly soon), giving your developers more without giving you more to manage. Usage across Antigravity, the platform, and the app rolls up into a single view instead of separate licenses and billing silos.
-
Pay less as your usage grows: If your AI workloads are steady or climbing, Flexible Savings Plans let you commit to a monthly spend you're comfortable with and take 10–20% off your token costs — no minimums, no maximums, and no new billing silo to manage.
-
Consolidated spend guardrails: You can now set hard monthly caps on AI spend and projects, estimate agent runtime costs, and catch sudden budget spikes before they hit your invoice.
Give your teams flexibility without losing control over spend in Gemini Enterprise
Every organization operates differently. Even within the same business, no two teams consume AI in the same way. Your business users might rely on steady, everyday productivity tools. Meanwhile, your technical teams might run AI agent workloads in bursts.
To help align costs with how work actually gets done, you can combine these payment and licensing choices and features across Gemini Enterprise:
|
Option |
How it works |
Why it helps optimize spend |
|
Gemini Enterprise app per-user seat subscription |
You pay a fixed monthly fee per user, which includes daily quota pools that are shared across your entire project. |
Predictable budgeting. It provides finance teams with a clear, steady monthly baseline for teams with consistent daily productivity needs. |
|
[New] Gemini Enterprise app pay-as-you-go consumption edition *available for select customers and rolling out broadly soon |
There is no upfront commitment or base subscription fee, meaning you pay strictly for the compute and tokens your teams consume at standard model API rates. |
Only pay for what you use. Your spend scales up and down automatically with real usage, ensuring you never pay for empty seats when project demand dips. |
|
[New for Antigravity in Gemini Enterprise] Consolidated pooled quotas |
Daily usage allowances are pooled project-wide, letting business apps, developer tools, and custom agents draw from the same shared quota. Pooled quota is always exhausted first, and admins can control if overages are allowed, at which point it’s charged at pay-as-you-go rates. |
Maximized resource usage: Unused daily allowances from business users automatically absorb heavy developer or custom API agent demands, so no quota allowance goes to waste. |
|
[Coming soon] Deferred execution pricing *available for select workloads soon |
Mark eligible agent workloads as deferred, and our intelligent scheduler in the Gemini Enterprise Agent Platform runs them during off-peak capacity windows. |
Substantial discounts for work that can wait: AI workloads can run on separate, off-peak capacity, you pay up to half the inference cost and bypass standard quota limits entirely – letting you run substantially more agentic volume under the same budget. |
Equip developers with advanced agentic tooling under a single Gemini Enterprise subscription
We’re rolling out access to Google Antigravity in Gemini Enterprise, an agent-first developer platform that brings powerful agentic coding and agent-building capabilities to technical teams, included with Gemini Enterprise subscriptions for eligible customers. In addition, Android developers can leverage the Google Antigravity quota included in their Gemini Enterprise subscriptions natively in Android Studio, the agentic IDE for professional Android development.
To be more efficient with agentic coding costs, we are pooling developer tools quota included in each Gemini Enterprise subscription and making it available across the whole Google Cloud project so your teams can benefit from the capacity you’re already purchasing. Your developers get access to advanced agentic tools, while you maintain centralized governance and control.
For a closer look into what’s new with Antigravity in Gemini Enterprise and how customers are putting it to work in production, take a look at our deep-dive.
Budget smarter with Gemini Enterprise Flexible Savings Plans (FSPs)
If your organization has steady or growing AI workloads, Gemini Enterprise Flexible Savings Plans offer a simple, spend-based commitment model across Gemini Enterprise usage. FSPs are designed to lower token costs while keeping budgets flexible:
-
Programmatic savings: Receive 10% off for 1-year or 20% off for 3-year commitments for monthly spending across Gemini Enterprise.
-
Tailored to your pace: With no minimum or maximum spend requirements, you can determine a monthly commitment that fits your current traffic and make adjustments as your usage increases over time.
-
Enterprise Agreement (EA) friendly: FSP spend seamlessly draws down against your existing Google Cloud EA, giving lines of business dedicated budget control without fragmenting your broader cloud commitments.
Gemini Enterprise Flexible Savings Plans are already available for self-serve customers and customers on enterprise agreements.
Give your teams the freedom to build while maintaining financial discipline
As a leader, your goal isn't to restrict the potential value of AI – it's to remove the financial and operational risk that you face without managed AI costs. You should be able to give engineering, marketing, and operational teams the freedom to innovate with agents, but you should also have the visibility to trust what those agents are doing and the safety nets to protect your budget.
To bridge this gap, we've built robust, native governance tooling directly into the Google Cloud Billing Console around three simple goals:
1. Plan before you scale: The Google Cloud Pricing Calculator lets you estimate anticipated costs in Gemini Enterprise across per-user licenses, developer tools, and background agent runtimes. It gives you the numbers you need to build clear business cases upfront before project work begins.
2. Enforce boundaries without micromanaging spend: Instead of spending time tracking daily usage variations across project teams, let these tools do the monitoring for you:
-
Early anomaly detection: If a project’s AI spending trends higher than normal, the system flags the deviation with root cause analysis and pinpoints the top 3 SKUs driving the increase so you can see exactly what changed.
<div class="article-module h-c-page">
<div class="h-c-grid">
<figure class="article-image--large
h-c-grid__col
h-c-grid__col--6 h-c-grid__col--offset-3
">
<img alt="1 Jul22_Anomalies_Image1" src="https://storage.googleapis.com/gweb-cloudblog-publish/images/1_Jul22_Anomalies_Image1.max-1000x1000.png" />
</a>
<figcaption class="article-image__caption "><p>Billing Console showing an Early Anomaly alert with the Root Cause Analysis (RCA) breakdown highlighting the driving SKUs</p></figcaption>
</figure>
</div>
</div>
-
- Project-level spend caps: When a project needs defined financial boundaries, you can set a firm monthly spend limit directly in the Google Cloud Billing Console. If a project hits its limit, the agent's API calls temporarily pause – protecting your budget without affecting the rest of your production infrastructure. Automated email alerts at 50%, 80% and 100% of the budget keep you informed of your progress against the spend limit.
-
- Overage controls: If a spend cap triggers, you can choose to resume work with a single click in the console. Alternatively, if your priority is continuous operation, you can turn on overages so excess usage smoothly transitions to consumption rates, which can draw directly against your FSP to keep overage unit costs heavily discounted.
<div class="article-module h-c-page">
<div class="h-c-grid">
<figure class="article-image--large
h-c-grid__col
h-c-grid__col--6 h-c-grid__col--offset-3
">
<img alt="3 PAYG Overage Enabled" src="https://storage.googleapis.com/gweb-cloudblog-publish/images/3_PAYG_Overage_Enabled.max-1000x1000.png" />
</a>
<figcaption class="article-image__caption "><p>Enabling overage pay-as-you-go for a project.</p></figcaption>
</figure>
</div>
</div>
3. Get visibility into business value: Use centralized billing reports paired with the FinOps agent to generate natural-language cost insight summaries of where your budget went, making it simple to show ROI to leadership.
<div class="article-module h-c-page">
<div class="h-c-grid">
<figure class="article-image--large
h-c-grid__col
h-c-grid__col--6 h-c-grid__col--offset-3
">
<img alt="Cost overview FinOps" src="https://storage.googleapis.com/gweb-cloudblog-publish/images/Billing_overage_-_dashboard_-_new_afternoo.max-1000x1000.png" />
</a>
<figcaption class="article-image__caption "><p>AI spending reporting in Google Cloud Console</p></figcaption>
</figure>
</div>
</div>
Go deeper with AI cost optimization
To build a full-stack FinOps strategy that optimizes the cost, latency, and performance of your models and infrastructure, explore our detailed architecture specifications and frameworks:
-
How to outsmart infrastructure constraints with dynamic capacity management: Discover how to optimize your compute investments with capabilities in Google Kubernetes Engine and Google Compute Engine that automatically schedule and reallocate resources to avoid interruptions, over-provisioning, and over-reliance on any one hardware configuration.
-
Expanding Google Antigravity for Enterprise Customers: Read our developer tooling deep-dive to see how technical teams are accelerating software delivery with agent-first workflows.
-
What sports cars can teach us about optimizing AI spend: More tokens doesn't always mean better AI. Read our conversation with Mike Clark, Director of Product Management for Gemini Enterprise Agent Platform, on how to balance horsepower with efficiency and get the highest return out of every dollar you spend on AI.
-
Protection during usage spikes: Your heavy workloads can surge during peak hours without forcing you to pay for expensive, dedicated infrastructure that sits idle the rest of the time. As your AI usage grows, Gemini models can automatically scale on demand without hitting artificial rate limits – processing up to 50 million tokens per minute. Read more about Provisioned Throughput.
Google Cloud Japan Blog
Box が Gemini Embeddings 2 でマルチモーダル エンタープライズ エージェントを実現
- Link: https://cloud.google.com/blog/ja/topics/partners/box-ai-agents-gemini-embeddings-multimodal-enterprise-ai/
- Published: 2026-08-26 11:00:00
- Fetched: 2026-08-27 21:31:29
詳細を表示
※この投稿は米国時間 2026 年 8 月 19 日に、Google Cloud blog に投稿されたものの抄訳です。
エンタープライズ コンテンツ管理は、クラウド移行の時代以来、最も大きなアーキテクチャの転換期を迎えています。
長年にわたり、企業は財務モデル、臨床試験プロトコル、M&A デュー デリジェンス ルーム、エンジニアリング スキーマ、法令遵守ハンドブックなど、数兆ギガバイトに及ぶ重要なデータを Box に保存してきました。これまでは、テキストベースの検索と検索拡張生成(RAG)によって、これらのリポジトリ内の膨大なナラティブ ナレッジが効果的に活用され、エンタープライズ AI インテリジェンスの強力かつ非常に効果的なベースラインが確立されてきました。
従来の RAG アーキテクチャはテキスト処理に優れていますが、エージェントの時代にはそれ以上のものが求められます。次なる論理的な進化は、このフレームワークを拡張して、テキストと並存する本質的にマルチモーダルで、高度な空間情報や構造を持つ要素を捉えることです。テキスト エンベディングは文章のインデックス作成に優れていますが、マルチモーダル アーキテクチャは、これまでにない画期的な機能を実現します。たとえば、財務表における厳密な行と列のセマンティクスを保持したり、臨床データのような視覚的情報を解釈したり、複数ページにわたるフローチャートのロジックを空間レイアウトを維持したままマッピングしたりすることが可能になります。
あらゆるデジタル コンテンツに対応する次世代の機能を提供するため、Google Cloud と Box は Gemini Multimodal Embeddings 2 を活用して高度なマルチモーダル機能を Box の Agentic Platform に統合し、業界をリードする Box のインテリジェント コンテンツ管理プラットフォームと Google Cloud の高度な AI エンべディングを融合させています。
エンベディングの改善によるメリット: ドキュメント コンテンツの次元の拡張
-
視覚的および空間的なジオメトリの保持: 複数列の表や財務マトリックスなどの複雑なドキュメント要素は、その空間的なレイアウトによって意味を成しています。これらの要素をフラットなテキスト文字列に変換すると、列ヘッダーと対応するデータポイントとの関連付けが失われてしまうおそれがあります。マルチモーダル エンベディングにより、システムは人間とまったく同じようにドキュメントを解釈できるようになり、空間的な位置関係の整合性が維持されます。
-
視覚的なモダリティの解明: 企業のドキュメントには、技術的な図表、プロセス フローチャート、ブランディング アセット、製品写真などの視覚的なインジケーターが多数含まれています。マルチモーダル機能により、これらの要素は検索システムにとって不可視なものではなくなり、ユーザーは画像とテキストを組み合わせた同時検索が可能になります。
-
ハイブリッド ファイル形式の接続: 実際のビジネス ワークフローが、単一のドキュメント形式で完結することはほとんどありません。エージェントは、PDF ポリシー、ログを追跡するスプレッドシート、プレゼンテーション資料を相互参照する必要がある場合があります。マルチモーダル エンベディングによって RAG を拡張することで、これら多様な形式を横断した統一的な理解が可能になります。
アーキテクチャ ソリューション: Gemini Multimodal Embeddings 2
Google Cloud の Gemini Multimodal Embeddings 2 は、テキスト、ラスター画像、ドキュメント ページ、レンダリングされたスプレッドシートの表、ビジュアル チャートを同一のセマンティック表現空間に埋め込むことができる、統合されたマルチモーダル ベクトル空間を実現します。
<div class="article-module h-c-page">
<div class="h-c-grid">
<figure class="article-image--large
h-c-grid__col
h-c-grid__col--6 h-c-grid__col--offset-3
">
<img alt="GIF_1" src="https://storage.googleapis.com/gweb-cloudblog-publish/original_images/image_bu28HMu.gif" />
</a>
</figure>
</div>
</div>
gemini-embeddings-2 によって実現される主なプロダクト機能:
-
クロスモーダル検索(テキストからビジュアル / ビジュアルからテキスト): 自然言語クエリにより、手動のタグ付けをすることなく、膨大なスライド ライブラリから特定のチャートや図解といったピンポイントなビジュアル要素を検索および抽出できるようになります。
-
レイアウトを考慮したドキュメント エンベディング: ファイルを任意のテキスト ブロックに分割するのではなく、ドキュメント ページのレンダリングを直接埋め込むことで、視覚的な階層構造、コールアウト ボックス、構造的なコンテキストをそのまま保持できます。
-
異種形式のブリッジ: モダリティ固有の構造情報を損なうことなく、.docx、.xlsx、.pdf、.pptx、.png、.csv 間のコンテンツをシームレスにつなぐネイティブ サポート。
マルチモーダル エンタープライズ エージェントの 3 つのコアパターン
Box 内でマルチモーダル エンベディングを活用することで、従来の RAG を拡張して複雑なビジュアル ワークフローをサポートするための、主要な 3 つの独自デザイン パターンを見出しました。
パターン 1: 複雑な財務レポートと分析レポート
課題
企業の財務、調査、監査の各チームは、高度に構造化されたドキュメントを分析します。こうしたドキュメントでは、埋め込みの表、成長チャート、脚注アノテーションの中に重要なデータが存在しています。テキストのみのインデックス登録では、これらの数字がコンテキストから分離されるため、自動分析が困難になります。
マルチモーダルの利点
-
構造の整合性: エンベディング モデルは表やグラフの物理的な構造を捉えるため、財務エージェントは、どの列ヘッダーが特定の指標の行に対応しているのかを正確に理解できるようになります。
-
視覚的な傾向分析: エージェントは、記述された要約と、付随する棒グラフや折れ線グラフの視覚的な傾向を相互参照し、文書内の主張とソースデータとの不一致を特定して指摘できます。
-
コンテキストに応じたソースの特定: ユーザーは複雑なポートフォリオに対してクエリを実行し、特定の指標を裏付ける正確なページ、表、グラフを即座に取得できます。
<div class="article-module h-c-page">
<div class="h-c-grid">
<figure class="article-image--large
h-c-grid__col
h-c-grid__col--6 h-c-grid__col--offset-3
">
<img alt="2" src="https://storage.googleapis.com/gweb-cloudblog-publish/images/2_9rTykxw.max-1000x1000.png" />
</a>
</figure>
</div>
</div>
パターン 2: マルチモーダルによる臨床意思決定支援と診断補助
課題
医療および臨床環境では、重要な患者データが、外部の外観写真(視覚的証拠)、顕微鏡の病理スライド(検査レポート)、構造化されたリスク マトリクス(トリアージ グリッド)など、さまざまな非構造化の視覚およびテキスト形式に断片化されています。従来のテキストベースのシステムや個別の分析ツールでは、こうしたクロスモーダルな関係性を同時に統合して理解することはできません。その結果、重要な診断の遅れを招いたり、生命に直結する処置上の合併症を見逃したりするリスクが生じます。
マルチモーダルの利点
-
クロスモーダルな臨床情報の統合: 臨床写真、組織病理画像、トリアージ グリッドを同一のスペースにインデックス化することで、身体症状と細胞レベルの検査エビデンスを同時に評価します。
-
微細な異常の特定: 顕微鏡下で観察される微細な視覚パターン(寄生虫の嚢胞壁など)を医学的知識と結び付け、希少な疾患を迅速に特定します。
-
リスクを考慮した意思決定支援: 得られた所見をトリアージ フレームワークと照合し、生命を脅かすアナフィラキシー ショックなど、患者の差し迫ったリスクについて即座に警告を発します。
<div class="article-module h-c-page">
<div class="h-c-grid">
<figure class="article-image--large
h-c-grid__col
h-c-grid__col--6 h-c-grid__col--offset-3
">
<img alt="3" src="https://storage.googleapis.com/gweb-cloudblog-publish/images/3_ZPNWwdP.max-1000x1000.png" />
</a>
</figure>
</div>
</div>
パターン 3: ドキュメント間のマルチモーダル統合とデータ調整
課題
企業内の情報は、接続されていないファイルや形式(PDF の議事録、Excel のグラフ、PNG のチラシ、メールスレッドなど)に断片化されています。従来のツールはこれらのファイルを個別に分析するため、独立したドキュメントを横断して詳細を確認したり、データの矛盾を解消したりする際に、情報を結びつけて全体像を把握することができません。
マルチモーダルの利点
-
ファイル間の統合: 完全に異なる形式(PDF、スプレッドシート、画像、メール)の情報を同時に結び付けて、複雑なビジネス クエリに回答します。
-
不整合の解決: アセット間の不整合を特定して解消します。たとえば、最新の財務スプレッドシートと照合することで、画像内に記載された古い価格情報を見つけ出すといったことが可能です。
-
画像とテキストの照合監査: 画像やスキャンされたファイルをテキストベースの記録と照らし合わせ(署名済みの PDF 契約書を法的レビューのメールで検証するなど)、条項の欠落や変更箇所を特定します。
<div class="article-module h-c-page">
<div class="h-c-grid">
<figure class="article-image--large
h-c-grid__col
h-c-grid__col--6 h-c-grid__col--offset-3
">
<img alt="4" src="https://storage.googleapis.com/gweb-cloudblog-publish/images/4_IAwu96l.max-1000x1000.png" />
</a>
</figure>
</div>
</div>
エージェント型エンタープライズ コンテンツ管理の未来
gemini-embeddings-2 と Box のエージェント プラットフォームのインテグレーションは、次世代のコンテンツ インテリジェンスを向上させる重要な新機能です。マルチモーダル エンベディングにより、Box は従来の検索の枠を超え、能動的かつインテリジェントなコラボレーションへと進化します。Box のインテリジェント コンテンツ管理プラットフォームは、エンタープライズ AI インフラストラクチャの根本的な転換を体現するものです。単なる受動的なドキュメント ストレージから、ガバナンスの効いたセマンティック インデックスによる推論レイヤに移行することで、AI エージェントは、既存のコンプライアンスとセキュリティ管理を維持したまま、コンテンツの照会や相互参照、アクションの実行を自在に行えるようになります。
マルチモーダル エンベディングと、検索、メタデータ抽出、調査、分析、作成にわたるネイティブ AI エージェント スイートを搭載した Box により、組織は、古い価格データの特定、契約条項の期限切れ、ドキュメント間の矛盾といったインサイトを、ビジネスリスクが顕在化する前にプロアクティブに抽出できるようになります。金融サービス、ライフ サイエンス、法務業務といった複雑性の高い業界では、マルチモーダル理解が競争上の必須要件となっています。Box なら、テキスト、表、グラフ、画像を横断して推論することが可能です。
より広範なエンタープライズ AI エコシステムとの相互運用を前提に設計された Box は、ガバナンスの効いた一元的なコンテンツ基盤として機能します。これにより、AI を活用したあらゆるワークフローが、承認済みで監査可能な企業データに基づいていることが保証されます。
考えてみれば、企業のデータ環境はもともとマルチモーダルなものでした。そして今、それを最大限に活用するためのテクノロジーが現実のものとなりました。gemini-embeddings-2 とのインテグレーションにより、Box は、ユーザーが非構造化エンタープライズ コンテンツからこれまでにない価値を引き出せるよう支援します。マルチモーダル ファーストのアーキテクチャ、厳格な精度ベンチマーク、そして監査対応のグラウンディングを積極的に取り入れるプロダクト リーダーが、企業の生産性とイノベーションにおける次なる波を牽引することになるでしょう。
このプロジェクトに携わってくださった Ken Ikeda、Afshaan Mazagonwalla、Samip Thakkar の各氏に感謝いたします。
- Google、エージェント プロダクト コンサルティング リード、Sandhya Patil
- Box、スタッフ AI プロダクト マネージャー、Darryl Sladden 氏
Gemini Enterprise、業界特化型ソリューション向けに進化
- Link: https://cloud.google.com/blog/ja/products/ai-machine-learning/introducing-gemini-enterprise-for-financial-services/
- Published: 2026-08-26 12:00:00
- Fetched: 2026-08-27 21:31:29
詳細を表示
※この投稿は米国時間 2026 年 8 月 26 日に、Google Cloud blog に投稿された Gemini Enterprise for Financial Services および Gemini Enterprise for Legal の抄訳です。
Google Cloud は本日、金融サービス向け AI Gemini Enterprise for Financial Services および法務向け AI Gemini Enterprise for Legal の提供開始を発表しました。あらゆる専門家が日々の業務プロセスに合わせて Google AI を活用できるようにネイティブ設計し、安全でガバナンスの行き届いた Gemini Enterprise プラットフォーム上に構築した、業界特化型のパッケージ ソリューションのシリーズの第一弾です。
各業界向けソリューションは、すぐに導入可能な領域特化型の AI 機能を備えています。これには調整済みのエージェント、多様なワークフローの管理を簡素化する専用の「スキル」、データコネクタ、特定業界向けに最適化したモデルなどが含まれます。 Google Cloud のインフラストラクチャと Gemini Enterprise プラットフォーム上に構築したこれらのソリューションは、エンタープライズ プラットフォームに求められるセキュリティ、ガバナンス、コンプライアンス、コスト マネジメント機能を提供します。顧客データ、ビジネス ルール、知的財産、カスタム エージェント、モデルの出力結果は組織固有のプライベートな状態を維持し、リスクマネジメント、監査ログ、ガバナンスをプラットフォーム アーキテクチャに直接組み込んでいます。さらに、インフラストラクチャやモデルからアプリケーション レイヤーに至る Google Cloud のフルスタック アプローチにより、組織はパフォーマンスとコストを最適化し、AI 投資の価値を最大限に引き出せます。
Gemini Enterprise for Financial Services によるイノベーション、セキュリティ、コンプライアンスの融合
資産の保護にスピードが求められる現在の市場において、汎用的な AI ツールには規制対象となる金融機関が求めるセキュリティ、リアルタイムな精度、データリネージが不足しています。また、手作業によるレポート作成によって、アナリストは迅速なインサイトの抽出ではなく、定型的なデータ収集に貴重な時間を費やしています。エージェンティック AI が業界標準となりつつある中、従来のマニュアル作業に依存したままの機関は、提案、価格設定、市場変化への対応において不利な状況に直面します。
Gemini Enterprise for Financial Services は、モデルと金融サービス業界に必要なデータ、システム、専門知識、統制機能を統合するプラットフォームを提供し、これらの課題を解決します。
Google Cloud の CEO である Thomas Kurian は次のように述べています。「金融分野の専門家は、特定のモデルやエコシステムにロックインされることなく、日常的に使用する IT システムと接続し、高度なセキュリティとコンプライアンスを備えた AI プラットフォームを求めています。 Gemini Enterprise for Financial Services により、検証可能なデータの説得力を備えた金融ワークフロー向けのエージェント機能を提供し、金融機関が AI イノベーションを確実な競合優位性へと変えられるよう支援します。」
Gemini Enterprise for Financial Services の主な構成要素:
-
Financial Research エージェント: Google が管理するこのエージェントは、完全な説明可能性を備えたエンドツーエンドのリサーチを自動化します。50 以上の基礎的なスキルを備え、確信度スコア、明示的な手法、監査を容易にするデータ スナップショット、確実な検証を可能にする正確な情報源の引用を通じて、推論プロセスの透明性を提供します。ユーザーは Gemini Enterprise app から直接エージェントにアクセスできるほか、 Agent2Agent (A2A) API を介して既存のエージェント ワークフローに接続できます。エージェントは Model Context Protocol (MCP) 統合を介して多種多様なエンタープライズ データソースに直接接続し、個別のニーズに合わせたカスタマイズ可能なレポートやドキュメントを生成します。
-
業界特化型金融スキル: 業界特化型スキルにより、ブランド ガイドラインの適用からレポートのフォーマット作成、特定データへのアクセスに至るまで、反復的な作業のショートカットを作成できます。プライベート エクイティの専門家、資産運用担当者、コンプライアンス チームのいずれが使用する場合でも、本ソリューションは信用リスク評価、ポートフォリオ モニタリング、市場ニュースの要約、調査型財務リサーチなど、多様なワークフローに適応します。スキルは Financial Research エージェント内で利用できるほか、お客様が自社で構築するエージェントからもアクセス可能です。
-
コネクタ: 本ソリューションは、CoinDesk Data & Indices、Daloopa、Dun & Bradstreet、FactSet、Finnhub、Fiscal.ai、Guidepoint、LSEG、Moody’s、MSCI、PitchBook、S&P Global、SEC Edgar などのライセンス取得済みデータソースに対する 13 のセキュアな統合を提供し、顧客環境内で直接設定できます。
-
サードパーティ エージェントおよびパートナー エコシステム: Gemini Enterprise for Financial Services は、法人顧客のオンボーディング加速と Know Your Customer (KYC) コンプライアンス強化を実現する D&B Business Verification エージェント、ローン パックの完全性確認、書類抽出、キャピタル マーケット向けのデータ照合を行う FlowX エージェント、プライベート マーケットや法人融資チームによる複雑な財務分析の加速を支援する Obin Financial エージェント、多段階の分析、レポート作成、リサーチ ワークフローに対応する S&P Global Data Retrieval エージェント、複雑なエネルギーやサステナビリティに関するデータを財務ワークフロー向けの迅速なインサイトに変換する S&P Global Energy Horizons エージェントなど、 Google Cloud のパートナー エコシステムが提供する複数のサードパーティ エージェントへ簡単にアクセスできます。
Gemini Enterprise for Financial Services は、すでに CME Group や、 Financial Research エージェントの主要な設計パートナーである Deutsche Bank をはじめとするグローバル金融機関が導入しています。 BNY、Citi Wealth、Lloyds Banking Group、Macquarie Bank、Signal Iduna などの多くの大手金融機関が、生産性と効率性を向上させる先進的なエージェンティック ワークフロー ツールとして利用を進めており、 Gemini Enterprise の導入の勢いは急速に高まっています。
法務実務向けに特化した AI である Gemini Enterprise for Legal
企業が複雑なエンドツーエンド プロセスの自動化を進める一方で、汎用ツールには厳格な機密保持、倫理的障壁(エシカル ウォール)、検証済みの法的根拠、完全なデータ隔離など、法務作業が求める統制機能が欠けています。分断されたポイント ソリューションや単体のモデルを組み合わせてこの問題に対処しようとすると、システム統合にかかるコストが増大し、脆弱なワークフローにつながります。
Gemini Enterprise for Legal は、これらの要件に焦点を当てて設計しており、契約レビュー、デュー デリジェンス、規制モニタリング、プライバシー リクエストなどの法務ワークフロー全体を実行する特化型スキル、パートナー エージェント、コネクタを提供します。主要なドキュメントおよび案件管理システムへのセキュアなコネクタを通じて、組織のエシカル ウォールを自動的に引き継ぎます。リサーチ出力はモデルの学習データだけでなく一次法源に基づき、組織のデータ、プレイブック、クライアント ファイル、交渉ポジションは常にプライベートな領域内に保護します。
Gemini Enterprise for Legal の主な構成要素:
-
法務エージェント向け特化型スキル: Gemini Enterprise for Legal には、準備書面の作成、引用の検証、コントラクト ライフサイクル マネジメント、規制動向の監視、データ 主体 アクセス リクエスト (DSAR) への対応など、複雑で高度な法務業務で AI エージェントを導く特化型スキルが含まれています。
-
コネクタ: 本ソリューションは、 Docusign、Everlaw、Harvey、iManage、Legora、NetDocuments、RelativityOne、Solve Intelligence、Thomson Reuters など、主要な法務業界プラットフォームへのセキュアな Model Context Protocol (MCP) 統合を提供します。
-
サードパーティ エージェントおよびパートナー エコシステム: Gemini Enterprise for Legal は、本プラットフォーム上で開発を行う法務技術プロバイダーのネットワークに加え、Accenture、Deloitte、Devoteam、Eudia、Factor Law、KPMG、Tribe AI、Valtech、Zazmic、Zencore、66degrees などの大手コンサルティングパートナーやシステム インテグレータへのスムーズなアクセスを提供します。
Google Cloud の CEO である Thomas Kurian は次のように述べています。「エージェンティック AI により、法務の専門家は過去の判例のリサーチ、複雑な論理展開の構築、定型作業の自動化を実現し、クライアントに大きな価値を提供できるようになります。一方で、これらのエージェンティック ワークフローのすべての側面が正確かつ事実に基づき、法的根拠に裏付けられていることが非常に重要です。 Gemini Enterprise for Legal は、現代の法務における複雑な実務に対応しながら、チームが自信を持って戦略的なアドバイスに集中できるよう、重要な根拠付けとガバナンス機能を提供します。」
提供時期
Gemini Enterprise for Financial Services および Gemini Enterprise for Legal は、本日より世界中の Google Cloud をご利用のお客様向けにプレビュー版として提供を開始します。ソリューションの詳細については、 Gemini Enterprise for Financial Services ブログおよび Gemini Enterprise for Legal ブログをご覧ください。
Google Cloud について
Google Cloud は、AI インフラストラクチャ、Gemini をはじめとする先進的なモデル、データ管理、マルチクラウド セキュリティ、開発ツール、さらにエージェントやアプリなど、強力で最適化された AI スタックを提供し、エージェント時代に向けた組織の変革を支援します。200 以上の国と地域で、信頼されるテクノロジー パートナーとして選ばれています。
Google Cloud Blog (AI & ML)
FinOps for the AI era: New flexible billing and cost controls for agents
- Link: https://cloud.google.com/blog/products/ai-machine-learning/flexible-billing-and-cost-controls-for-agents-on-google-cloud/
- Published: 2026-08-26 22:30:00
- Fetched: 2026-08-27 21:31:31
詳細を表示
Editor's note: A product image was updated after initial publication.
As AI takes on more complex work, business leaders face a new challenge: enabling rapid innovation using agents while protecting their margins and budgets. To get a real return on AI, financial operations (FinOps) and cost management must evolve alongside technology, giving you clear visibility, proactive cost controls, and flexible payment models that fit your needs.
That’s why today we’re introducing expanded billing flexibility and new cost management tools for agent workloads across Gemini Enterprise and developer tools like Google Antigravity in Gemini Enterprise and Android Studio.
-
Flexible payment options: You can mix our existing, predictable per-user seat subscriptions with a new pay-as-you-go option in Gemini Enterprise app that lets you run agent workloads without hitting quota limits mid-task.
-
Developer access, one place to manage your AI: Google Antigravity and Android Studio AI use is now included in your Gemini Enterprise subscription (available for select customers and rolling out broadly soon), giving your developers more without giving you more to manage. Usage across Antigravity, the platform, and the app rolls up into a single view instead of separate licenses and billing silos.
-
Pay less as your usage grows: If your AI workloads are steady or climbing, Flexible Savings Plans let you commit to a monthly spend you're comfortable with and take 10–20% off your token costs — no minimums, no maximums, and no new billing silo to manage.
-
Consolidated spend guardrails: You can now set hard monthly caps on AI spend and projects, estimate agent runtime costs, and catch sudden budget spikes before they hit your invoice.
Give your teams flexibility without losing control over spend in Gemini Enterprise
Every organization operates differently. Even within the same business, no two teams consume AI in the same way. Your business users might rely on steady, everyday productivity tools. Meanwhile, your technical teams might run AI agent workloads in bursts.
To help align costs with how work actually gets done, you can combine these payment and licensing choices and features across Gemini Enterprise:
|
Option |
How it works |
Why it helps optimize spend |
|
Gemini Enterprise app per-user seat subscription |
You pay a fixed monthly fee per user, which includes daily quota pools that are shared across your entire project. |
Predictable budgeting. It provides finance teams with a clear, steady monthly baseline for teams with consistent daily productivity needs. |
|
[New] Gemini Enterprise app pay-as-you-go consumption edition *available for select customers and rolling out broadly soon |
There is no upfront commitment or base subscription fee, meaning you pay strictly for the compute and tokens your teams consume at standard model API rates. |
Only pay for what you use. Your spend scales up and down automatically with real usage, ensuring you never pay for empty seats when project demand dips. |
|
[New for Antigravity in Gemini Enterprise] Consolidated pooled quotas |
Daily usage allowances are pooled project-wide, letting business apps, developer tools, and custom agents draw from the same shared quota. Pooled quota is always exhausted first, and admins can control if overages are allowed, at which point it’s charged at pay-as-you-go rates. |
Maximized resource usage: Unused daily allowances from business users automatically absorb heavy developer or custom API agent demands, so no quota allowance goes to waste. |
|
[Coming soon] Deferred execution pricing *available for select workloads soon |
Mark eligible agent workloads as deferred, and our intelligent scheduler in the Gemini Enterprise Agent Platform runs them during off-peak capacity windows. |
Substantial discounts for work that can wait: AI workloads can run on separate, off-peak capacity, you pay up to half the inference cost and bypass standard quota limits entirely – letting you run substantially more agentic volume under the same budget. |
Equip developers with advanced agentic tooling under a single Gemini Enterprise subscription
We’re rolling out access to Google Antigravity in Gemini Enterprise, an agent-first developer platform that brings powerful agentic coding and agent-building capabilities to technical teams, included with Gemini Enterprise subscriptions for eligible customers. In addition, Android developers can leverage the Google Antigravity quota included in their Gemini Enterprise subscriptions natively in Android Studio, the agentic IDE for professional Android development.
To be more efficient with agentic coding costs, we are pooling developer tools quota included in each Gemini Enterprise subscription and making it available across the whole Google Cloud project so your teams can benefit from the capacity you’re already purchasing. Your developers get access to advanced agentic tools, while you maintain centralized governance and control.
For a closer look into what’s new with Antigravity in Gemini Enterprise and how customers are putting it to work in production, take a look at our deep-dive.
Budget smarter with Gemini Enterprise Flexible Savings Plans (FSPs)
If your organization has steady or growing AI workloads, Gemini Enterprise Flexible Savings Plans offer a simple, spend-based commitment model across Gemini Enterprise usage. FSPs are designed to lower token costs while keeping budgets flexible:
-
Programmatic savings: Receive 10% off for 1-year or 20% off for 3-year commitments for monthly spending across Gemini Enterprise.
-
Tailored to your pace: With no minimum or maximum spend requirements, you can determine a monthly commitment that fits your current traffic and make adjustments as your usage increases over time.
-
Enterprise Agreement (EA) friendly: FSP spend seamlessly draws down against your existing Google Cloud EA, giving lines of business dedicated budget control without fragmenting your broader cloud commitments.
Gemini Enterprise Flexible Savings Plans are already available for self-serve customers and customers on enterprise agreements.
Give your teams the freedom to build while maintaining financial discipline
As a leader, your goal isn't to restrict the potential value of AI – it's to remove the financial and operational risk that you face without managed AI costs. You should be able to give engineering, marketing, and operational teams the freedom to innovate with agents, but you should also have the visibility to trust what those agents are doing and the safety nets to protect your budget.
To bridge this gap, we've built robust, native governance tooling directly into the Google Cloud Billing Console around three simple goals:
1. Plan before you scale: The Google Cloud Pricing Calculator lets you estimate anticipated costs in Gemini Enterprise across per-user licenses, developer tools, and background agent runtimes. It gives you the numbers you need to build clear business cases upfront before project work begins.
2. Enforce boundaries without micromanaging spend: Instead of spending time tracking daily usage variations across project teams, let these tools do the monitoring for you:
-
Early anomaly detection: If a project’s AI spending trends higher than normal, the system flags the deviation with root cause analysis and pinpoints the top 3 SKUs driving the increase so you can see exactly what changed.
<div class="article-module h-c-page">
<div class="h-c-grid">
<figure class="article-image--large
h-c-grid__col
h-c-grid__col--6 h-c-grid__col--offset-3
">
<img alt="1 Jul22_Anomalies_Image1" src="https://storage.googleapis.com/gweb-cloudblog-publish/images/1_Jul22_Anomalies_Image1.max-1000x1000.png" />
</a>
<figcaption class="article-image__caption "><p>Billing Console showing an Early Anomaly alert with the Root Cause Analysis (RCA) breakdown highlighting the driving SKUs</p></figcaption>
</figure>
</div>
</div>
-
- Project-level spend caps: When a project needs defined financial boundaries, you can set a firm monthly spend limit directly in the Google Cloud Billing Console. If a project hits its limit, the agent's API calls temporarily pause – protecting your budget without affecting the rest of your production infrastructure. Automated email alerts at 50%, 80% and 100% of the budget keep you informed of your progress against the spend limit.
-
- Overage controls: If a spend cap triggers, you can choose to resume work with a single click in the console. Alternatively, if your priority is continuous operation, you can turn on overages so excess usage smoothly transitions to consumption rates, which can draw directly against your FSP to keep overage unit costs heavily discounted.
<div class="article-module h-c-page">
<div class="h-c-grid">
<figure class="article-image--large
h-c-grid__col
h-c-grid__col--6 h-c-grid__col--offset-3
">
<img alt="3 PAYG Overage Enabled" src="https://storage.googleapis.com/gweb-cloudblog-publish/images/3_PAYG_Overage_Enabled.max-1000x1000.png" />
</a>
<figcaption class="article-image__caption "><p>Enabling overage pay-as-you-go for a project.</p></figcaption>
</figure>
</div>
</div>
3. Get visibility into business value: Use centralized billing reports paired with the FinOps agent to generate natural-language cost insight summaries of where your budget went, making it simple to show ROI to leadership.
<div class="article-module h-c-page">
<div class="h-c-grid">
<figure class="article-image--large
h-c-grid__col
h-c-grid__col--6 h-c-grid__col--offset-3
">
<img alt="Cost overview FinOps" src="https://storage.googleapis.com/gweb-cloudblog-publish/images/Billing_overage_-_dashboard_-_new_afternoo.max-1000x1000.png" />
</a>
<figcaption class="article-image__caption "><p>AI spending reporting in Google Cloud Console</p></figcaption>
</figure>
</div>
</div>
Go deeper with AI cost optimization
To build a full-stack FinOps strategy that optimizes the cost, latency, and performance of your models and infrastructure, explore our detailed architecture specifications and frameworks:
-
How to outsmart infrastructure constraints with dynamic capacity management: Discover how to optimize your compute investments with capabilities in Google Kubernetes Engine and Google Compute Engine that automatically schedule and reallocate resources to avoid interruptions, over-provisioning, and over-reliance on any one hardware configuration.
-
Expanding Google Antigravity for Enterprise Customers: Read our developer tooling deep-dive to see how technical teams are accelerating software delivery with agent-first workflows.
-
What sports cars can teach us about optimizing AI spend: More tokens doesn't always mean better AI. Read our conversation with Mike Clark, Director of Product Management for Gemini Enterprise Agent Platform, on how to balance horsepower with efficiency and get the highest return out of every dollar you spend on AI.
-
Protection during usage spikes: Your heavy workloads can surge during peak hours without forcing you to pay for expensive, dedicated infrastructure that sits idle the rest of the time. As your AI usage grows, Gemini models can automatically scale on demand without hitting artificial rate limits – processing up to 50 million tokens per minute. Read more about Provisioned Throughput.