GCP News - 2026-08-19

2026-08-19
最終更新: 2026-08-27 21:31:31 JST

Google Cloud Release Notes

August 19, 2026

詳細を表示

App Engine flexible environment Go

Feature

Support for the Go 1.27 runtime is in Preview.

Feature

Starting from Go runtime version 1.26 and later, the lifecycle support dates align more closely with the Go community release cycle. For more information, see Runtime support schedule.

App Engine standard environment Go

Feature

Support for the Go 1.27 runtime is in Preview.

Feature

Starting from Go runtime version 1.26 and later, the lifecycle support dates align more closely with the Go community release cycle. For more information, see Runtime support schedule.

Feature

You can migrate your App Engine push queues to Cloud Tasks by updating the bundled services SDK. This method lets you upgrade your app without needing to modify your application code. For more information on how to migrate, see the push queues migration guide (Preview).

App Engine standard environment Java

Feature

You can migrate your App Engine push queues to Cloud Tasks by updating the bundled services SDK. This method lets you upgrade your app without needing to modify your application code. For more information on how to migrate, see the push queues migration guide (Preview).

App Engine standard environment Python

Feature

You can migrate your App Engine push queues to Cloud Tasks by updating the bundled services SDK. This method lets you upgrade your app without needing to modify your application code. For more information on how to migrate, see the push queues migration guide (Preview).

Batch

Deprecated

The Batch Debian 11 operating system (OS) image family has reached end of development due to the end of support (EOS) for Compute Engine Debian 11 images on August 31, 2026. The last Batch Debian 11 images—any image versions with the batch-debian-11-official prefix—are only supported until August 31, 2026. Before then, migrate any job that uses a Batch Debian 11 image to a Batch Debian 12 image (or other image) as follows:

  • For job definitions that use the batch-debian image prefix (which is the default image for jobs with any script runnables), the image that Batch automatically selects during job creation is gradually migrating to Debian 12 no later than August 31, 2026. For any jobs created before August 31, 2026, you can check whether the job uses Debian 11 or Debian 12 by describing the job.

  • For job definitions that specify either the batch-debian-11-official image family or an image version with that prefix, specify a different image during job creation. For example, to migrate to Debian 12, specify either the batch-debian-12-official image family or an image version with that prefix.

Learn more about OS images, viewing OS images, and specifying OS images.

Buildpacks

Feature

Starting from Go runtime version 1.26 and later, the lifecycle support dates align more closely with the Go community release cycle. For more information, see Runtime support schedule.

Cloud CDN

Feature

Global Front End is a unified offering that simplifies billing by consolidating pricing across networking products, including Cloud CDN, global external Application Load Balancer, Google Cloud Armor, and Service Extensions, into one solution to help deliver, scale, and secure your internet-facing applications. Cloud CDN is included in the Global Front End Enterprise billing tier. This feature is available in Preview.

For more information, see Global Front End.

Cloud Load Balancing

Feature

Global Front End combines global external Application Load Balancers, Google Cloud Armor, Cloud CDN, and Service Extensions into one solution to help deliver, scale, and secure your internet-facing applications.

For more information, see Global Front End.

This feature is available in Preview.

Cloud Run

Feature

Support for the Go 1.27 runtime is in Preview.

Feature

Starting from Go runtime version 1.26 and later, the lifecycle support dates align more closely with the Go community release cycle. For more information, see Runtime support schedule.

Cloud Run functions

Feature

Support for the Go 1.27 runtime is in Preview.

Feature

Starting from Go runtime version 1.26 and later, the lifecycle support dates align more closely with the Go community release cycle. For more information, see Runtime support schedule.

Container Optimized OS

Change

cos-125-19216-532-121

Kernel Docker Containerd GPU Drivers
COS-6.12.94 v27.5.1 v2.1.9 See List

Fixed

Updated cos-gpu-installer to v2.7.6.

Security

Fixed CVE-2026-68116 in the Linux kernel.

Security

Fixed CVE-2026-68139 in the Linux kernel.

Security

Fixed CVE-2026-68171 in the Linux kernel.

Security

Fixed CVE-2026-68325 in the Linux kernel.

Security

Fixed CVE-2026-68336 in the Linux kernel.

Security

Fixed CVE-2026-68343 in the Linux kernel.

Change

cos-129-19506-299-148

Kernel Docker Containerd GPU Drivers
COS-6.12.94 v27.5.1 v2.2.6 See List

Fixed

Updated cos-gpu-installer to v2.7.6.

Security

Fixed CVE-2026-68284 in the Linux kernel.

Security

Fixed CVE-2026-68398 in the Linux kernel.

Change

Runtime sysctl changes:

  • Changed: net.ipv4.udp_mem: 188034 250714 376068 -> 188034 250715 376068

Gemini Enterprise

Change

Gemini Enterprise: Subscription seat quantity limits

If you purchase or modify a seat-based subscription directly through the Google Cloud console, the number of seats is capped according to your Cloud Billing account type:

  • Self-serve (online) and resold accounts: Up to 25 seats per account.
  • Invoiced (offline) accounts: Up to 1,000 seats per account.

For more information, see Subscription seat quantity limits.

Google Cloud Armor

Feature

Global Front End is a unified offering that simplifies billing by consolidating pricing across networking products. Cloud Armor is included in the Global Front End Enterprise billing tier. Enabling Global Front End Enterprise in a project enables specific Cloud Armor Enterprise features for your global external Application Load Balancers. For more information, see Global Front End.This feature is available in Preview.

Google Cloud Contact Center as a Service

Announcement

Google Cloud CCaaS 6.3

We've released version 6.3 of Google Cloud CCaaS.

The timing of the update to your instance depends on the deployment schedule that you have chosen. For more information, see Deployment schedules.

Feature

Agent desktop supports parameters in custom panel URLs

In the agent desktop, you can now configure fixed and dynamic parameters to include in the URLs of custom panels. This lets you pass relevant session, agent, and customer context into custom panels.

For more information, see Parameters.

Feature

Callback offer restrictions

You now have greater control over when callbacks are offered. You can configure the following:

  • Prevent callback offers from being made outside of callback hours.

  • Prevent callback offers that will likely occur outside of callback hours. If conditions improve (that is, EWT decreases), the system adjusts and can offer callbacks.

Administrators: We've added the following checkboxes to the CCAI Platform portal:

  • Restrict callback offer outside of callback window

  • Restrict callback offer that will exceed hours of operation. If queue condition improve offer callbacks

These checkboxes are available in the following locations:

  • The Settings > Call > Callback Settings pane (to configure globally).

  • The Settings > Queue > IVR (Interactive Voice Response) > Edit / View > QUEUE_NAME > Callback Settings > Configure > Callback Management pane (to configure a queue).

For more information, see Manage callbacks.

Fixed

This release addresses the following issues:

  • Fixed an issue where chat transcripts incorrectly displayed undefined joined instead of the agent's name during a chat transfer when real-time redaction was enabled.

  • Fixed an issue where agents were incorrectly placed into Unresponsive status after answering a call if the customer declined the system's callback attempt.

  • Fixed an issue where overcapacity deflection didn't trigger when agent-initiated outbound or direct-inbound calls were transferred to a queue, causing users to wait indefinitely.

  • Fixed an issue where the Answer button didn't appear for incoming calls in the agent desktop, preventing agents from accepting calls.

  • Fixed an issue where live translation didn't activate (or translated in the wrong direction) after a chat was transferred between queues.

  • Fixed an issue where predictive campaign calls became stuck in a queued state, causing agents to appear available despite being unable to receive new calls.

  • Fixed an issue where web chats became stuck in a queued state and were never assigned to an agent.

  • Fixed an issue where previously closed chat sessions briefly reappeared and gained focus when launching agent desktop.

  • Fixed an issue where agents couldn't submit disposition codes and notes during wrap-up.

  • Fixed an issue where agents couldn't change their status after a call ended abnormally, causing them to remain stuck in the wrap-up state.

  • Fixed an issue where the Call History list incorrectly displayed the same customer phone number for all Acqueon campaign calls.

  • Fixed an issue where an agent's status incorrectly remained Available during outbound calls and wrap-up periods, allowing the routing engine to offer new inbound calls to occupied agents.

  • Fixed an issue where agents were forced to re-authenticate when opening the chat adapter or email adapter despite having an active session.

  • Fixed an issue where the text screen in the email adapter suddenly re-rendered while typing, causing characters to disappear or be displaced.

  • Fixed an issue where the headings in Generative AI session summaries in chat wrap-up notes lost their bold formatting when saved.

  • Fixed an issue where the global after-hours deflection message played despite a queue-level custom redirect to a phone number being configured.

  • Fixed an issue where calls abandoned by a customer while waiting in a transfer queue were incorrectly reported as queue failures.

  • Fixed an issue where transient CRM errors caused significant delays in retrieving ticket IDs during active calls.

  • Fixed an issue where the CCAI Platform portal incorrectly displayed a chat status as Unknown (instead of Check In Timeout) when a consumer didn't check in.

  • Fixed an issue where the call adapter incorrectly displayed Portuguese BR instead of Portuguese (Portugal) during call handoffs.

  • Fixed an issue where agents making outbound calls remained in an Available status, which incorrectly allowed new inbound calls to be routed to them while they were already engaged.

  • Fixed an issue where Agent Assist live transcription and generative summary stopped working mid-call following a transfer or a hold-and-resume cycle.

  • Fixed an issue where the transfer menu delayed loading queues.

  • Fixed an issue where the agent adapter incorrectly displayed an agent's status as Unavailable when a custom status, such as Break or Special Task, was selected.

  • Fixed an issue where chats escalated from a virtual agent to a human agent queue bypassed menu-level after-hours and over-capacity deflection messages.

  • Fixed an issue where the call adapter defaulted to English in the Outbound call screen regardless of the agent's system language.

  • Fixed an issue where Call and Chat each appeared twice in the Dashboard menu when using high browser zoom levels or small window resolutions.

  • Fixed an issue where legacy dashboards were restricted to English-only labels.

  • Fixed an issue where users in SAML-only or SSO-enabled environments received invitation emails directing them to a non-existent Forgot Password flow.

  • Fixed an issue where the user activity logs incorrectly recorded an end-user ID instead of the agent's ID when a chat disconnected.

  • Fixed an issue where customer calls were abandoned during payment transactions when DTMF inputs were provided.

  • Fixed an issue where agents could see and select outbound caller IDs that weren't assigned to their teams or queues.

  • Fixed an issue where post-session virtual agent transfers stalled if a customer left the chat while still in a queue for a human agent.

  • Fixed an issue where agents and end-users could hear each other's voices despite the agent putting the call on hold.

  • Fixed an issue where the agent adapter incorrectly reverted to displaying English when initiating an outbound call in a non-default language.

  • Fixed an issue where the Hide Agent Assist button wasn't appearing in the user interface during the call disposition phase.

  • Fixed an issue where cascade agent availability conditions weren't enforced, causing regional agents to be incorrectly routed into international queues and leaving local queues understaffed.

  • Fixed an issue where agents appeared available but weren't receiving calls.

  • Fixed an issue where administrators received multiple email notifications instead of a single notification after deactivating a call channel.

  • Fixed an issue where Telnyx error handling was incorrectly configured, causing system alerts to fail during call disconnect or hold actions.

  • Fixed an issue where inbound IVR calls were stuck in a queued state if a caller disconnected during initial call processing.

  • Fixed an issue where callbacks became permanently stuck in a queued state if an agent missed a projected call.

  • Fixed an issue where calls using Telnyx or Nexmo numbers received an "An Application error has occurred" message during overcapacity deflection or automatic redirection.

  • Fixed an issue where call recordings weren't delivered or processed correctly.

  • Fixed an issue under Settings > Operation Management > Localization > Manage Location Setting where configured locations with Portuguese (Portugal) or Spanish (Spain) selected incorrectly displayed as Unknown in the Language column, and locations with Spanish (Mexico) selected incorrectly displayed as Spanish (Spain).

  • Fixed an issue in iOS and Android SDKs where the end user's initial message and custom data weren't correctly passed to the routing API during the chat menu fetch.

  • Fixed an issue where inbound calls cancelled by the caller within the first second incorrectly remained active in the system for several hours.

  • Fixed an issue where custom SIP headers were dropped. This occurred when a directly dialed agent was over capacity, the call was redirected to the agent's queue, and that queue was configured to redirect to a SIP URI.

  • Fixed an agent desktop issue where the chat screen went blank when an agent accepted or dismissed a chat.

  • Fixed an issue where the Deflections - Calls dashboard incorrectly reported the queues that calls were redirected to when using percent allocation.

  • Fixed an issue where auto-answered interactions became stuck in a Queued status despite being accepted by an agent.

  • Fixed an issue where the chat adapter in a CRM integration didn't post outbound messages when using rich text formatting.

  • Fixed an issue where the agent_activity_logs endpoint experienced timeouts and degraded performance when processing large data requests.

  • Fixed an agent desktop issue where outbound calls appeared to start successfully but didn't connect to the end-user.

  • Fixed an issue where New photo received notifications appeared whenever an agent switched between chat tabs.

  • Fixed an issue where short outbound calls incorrectly displayed a This call was abandoned by the customer message.

  • Fixed an issue where agents experienced significant delays when switching between multiple concurrent web chats.

  • Fixed an issue where attaching PDF or text files in the chat adapter failed or timed out.

  • Fixed an issue where calls transferred from a virtual agent to an agent extension were incorrectly deflected.

  • Fixed an issue where reordering queues in the CCAI Platform portal experienced extreme latency and didn't update visually without a manual page refresh.

  • Fixed an issue where chat history for added agents became unavailable after a page refresh.

  • Fixed an issue where calls to the user_activity_logs endpoint experienced significant delays.

  • Fixed an issue where repeated voicemail re-reads caused duplicate participant entries and inflated session data.

  • Fixed an issue where callback sessions remained in a Connected status indefinitely after completion.

  • Fixed an issue where agents using IdP-initiated SAML SSO were automatically routed to the default home page instead of the agent desktop.

  • Fixed an issue where DTMF options in the softphone didn't register with external IVR systems during outbound calls.

  • Fixed an issue where agents were unable to receive or fetch incoming calls.

  • Fixed an issue where supervisors using Telnyx who ended a monitoring session were unable to monitor subsequent calls.

  • Fixed an issue where empty chat bubbles appeared in the chat adapter when an end-user used suggestion chips to respond to a virtual agent.

  • Fixed an issue where duplicated contact handle-duration events caused inaccurate reporting for call and chat interactions.

  • Fixed an issue where agents didn't receive the correct error messages when their microphone was disabled or inaccessible during a call.

  • Fixed an issue where incoming calls were incorrectly multicasted and didn't auto-answer when an agent was already in an active chat session.

  • Fixed an issue where calls deflected to external SIP or phone destinations were missing the ends_at timestamp in call data.

  • Fixed an issue where the chat timeout event wasn't correctly emitted to the headless web SDK, causing sessions to remain active for up to 120 minutes, regardless of the configured settings.

  • Fixed an issue where agents using Salesforce-Lightning or Zendesk embeds couldn't maintain a stable presence connection.

  • Fixed an issue where headless web SDK client methods didn't work after a mid-session authentication update.

  • Fixed an issue where agents were unexpectedly logged out of all active sessions.

  • Fixed an issue where calls producing multiple recording segments resulted in duplicate entries in Customer Experience Insights.

  • Fixed an issue where the parent_id field was missing from callback call responses in the manager API.

  • Fixed an issue where virtual agent transcripts were lost during escalations to human agents.

  • Fixed an issue in the agent desktop where the sentiment score appeared in the Call details panel despite sentiment analysis being turned off in the conversation profile.

  • Fixed an issue where the Escalated To Language column in the Escalations table of the Virtual Agent - Calls dashboard incorrectly displayed Unknown for French (Canada) calls.

  • Fixed an issue where calls didn't advance to the next cascade group after the timer threshold was reached.

Google Distributed Cloud (software only) for VMware

Announcement

Google Distributed Cloud (software only) for VMware 1.34.800-gke.90 is now available for download. To upgrade, see Upgrade clusters. Google Distributed Cloud 1.34.800-gke.90 runs on Kubernetes v1.34.7-gke.200.

If you use a third-party storage vendor, check the listing of our previously-qualified storage partners.

After a release, it takes approximately 7 to 14 days for the version to become available for use with GKE On-Prem API clients: the Google Cloud console, the gcloud CLI, and Terraform.

Fixed

The following issues were fixed in 1.34.800-gke.90:

  • Fixed vulnerabilities listed in Vulnerability fixes.
  • Fixed an issue where user clusters remained stuck in a Reconciling state after an admin cluster upgrade. The admin cluster controller skipped reconciling legacy cluster lifecycle components during upgrades unless an initial migration annotation was set. If legacy user clusters still existed on the admin cluster, missing legacy API discovery (cluster.k8s.io/v1alpha1) caused controller reconciliation to stall. With this fix, the controller preserves legacy components as long as any legacy user clusters exist, and prunes them only after all user clusters have migrated to advanced clusters.
  • Fixed an issue where gkectl prepare failed with a permission denied error when attempting to read a private registry CA certificate. The certificate file permissions are now set to 644 so non-root processes can read it.
  • Fixed an issue where retrying a failed upgrade to an Advanced Cluster (such as re-running with an existing bootstrap cluster) could wipe or strip the encryption keys in the generated-key-kms-plugin-config secret, preventing the control plane from decrypting existing Kubernetes secrets in etcd.

Google Distributed Cloud (software only) for bare metal

Announcement

Google Distributed Cloud (software only) for bare metal 1.34.800-gke.90 is now available for download. To upgrade, see Upgrade clusters. Google Distributed Cloud for bare metal 1.34.800-gke.90 runs on Kubernetes v1.34.7-gke.200.

After a release, it takes approximately 7 to 14 days for the version to become available for installations or upgrades with the GKE On-Prem API clients: the Google Cloud console, the gcloud CLI, and Terraform.

If you use a third-party storage vendor, check the listing of our previously-qualified storage partners.

Fixed

The following issues were fixed in 1.34.800-gke.90:

Google SecOps Marketplace

Feature

Wiz: Version 9.0

  • Added the following new job:

    • Wiz and Google SecOps Bi-directional Sync Job

Change

Exchange: Version 125.0

  • Updated the parsing logic for nested S/MIME email attachments (.eml) sent from macOS and Windows in the following connector:

    • Exchange Mail Connector v2 with Oauth Authentication

Change

Microsoft Teams: Version 39.0

  • Updated error handling in the following job:

    • Refresh Token Renewal Job

Change

Palo Alto Cortex XDR: Version 32.0

  • Fixed an issue where the job repeatedly logged errors when a case was merged or deleted in Google SecOps in the following job:

    • Sync Incidents

Change

Proofpoint Cloud Threat Response: Version 5.0

  • Fixed an issue where null, missing, or unmapped priority values in API payloads caused log ingestion errors in the following connector:

    • Proofpoint Cloud Threat Response - Incidents Connector

Change

Pub/Sub: Version 4.0

  • Added support for pubsub_message_id in the Unique ID Field parameter in the following connector:

    • Pub/Sub - Messages Connector

Change

Splunk: Version 67.0

  • Updated the lookback timestamp progression logic when all fetched alerts in a cycle have already been processed in the following connector:

    • Splunk ES - Notable Events Connector

Change

Zscaler: Version 16.0

  • Fixed an issue where legacy API key and password authentication failed due to URL path normalization issues.

Managed Service for Apache Spark

Announcement

New Managed Service for Apache Spark (formerly Dataproc on Compute Engine) subminor cluster image versions:

  • 2.1.118-debian11, 2.1.118-rocky8, 2.1.118-ubuntu20, 2.1.118-ubuntu20-arm
  • 2.2.86-debian12, 2.2.86-rocky9, 2.2.86-ubuntu22, 2.2.86-ubuntu22-arm
  • 2.3.35-debian12, 2.3.35-ml-ubuntu22, 2.3.35-rocky9, 2.3.35-ubuntu22, 2.3.35-ubuntu22-arm
  • 3.0.1-debian13, 3.0.1-ml-ubuntu24, 3.0.1-rocky9, 3.0.1-ubuntu24

Key updates in these image versions include:

  • Iceberg updates: In the 2.3 image version, 2.3 clusters with Lightning Engine now use Iceberg version 1.10 by default.
  • OpenLineage updates: In the 2.2 and 2.3 image versions:
    • Upgraded OpenLineage to version 1.49 to support lineage for tables created using the Lakehouse Runtime catalog.

Fixed

Managed Service for Apache Spark (formerly Dataproc on Compute Engine): Fixed a segmentation fault when OpenLineage parses complex SQL query strings.

Oracle Database@Google Cloud

Feature

Oracle Database@Google Cloud supports provisioning VM file system storage (VM images) and VM backups on Exascale storage for Exadata VM Clusters. This feature lets you to offload VM artifacts to Exascale, freeing up capacity and reducing dependence on local DB Server storage. For more information, see Configure Exascale Storage Vault for Exadata Infrastructure and Create Exadata VM Clusters with Exascale Storage Vaults.

This feature is Generally Available (GA).

Virtual Private Cloud

Feature

General Availability: You can use Private Service Connect endpoints and backends to access multi-regional service endpoints such as storage.us.rep.googleapis.com.

Google Cloud Blog (AI & ML)

How Box is unlocking multimodal enterprise agents with Gemini Embeddings 2

詳細を表示

Enterprise content management is experiencing its biggest architectural shift since the cloud migration era. 

For years, enterprises have stored trillions of gigabytes of critical data in Box: financial models, clinical trial protocols, M&A due diligence rooms, engineering schematics, and legal compliance playbooks. Up to this point, text-based search and retrieval-augmented generation (RAG) have successfully unlocked the vast narrative knowledge within these repositories, establishing a powerful and highly effective baseline for enterprise AI intelligence.

Traditional RAG architectures have mastered text processing, but the agentic era demands more. The next logical evolution is to extend this framework to capture the inherently multimodal, deeply spatial, and highly structured elements that exist alongside text. While text embeddings excel at indexing prose, multimodal architectures unlock a major new capability: For example, they preserve the strict row-column semantics of financial tables, interpret visual evidence like clinical data, and map the logic of multi-page flowcharts without losing their spatial layout.

To deliver next-generation capabilities that can handle the vast universe of digital content, Google Cloud and Box are integrating advanced multimodal capabilities into Box's Agentic Platform, powered by Gemini Multimodal Embeddings 2 merging Box’s industry-leading Intelligent Content Management platform with Google Cloud’s advanced AI embeddings.

Benefits of improved embedding: Extending the dimensions of document content

  1. Preserving visual and spatial geometry: Complex document elements like multi-column tables or financial matrices rely on their spatial layout to convey meaning. Converting these elements into a flat string of text can disassociate column headers from their corresponding data points. Multimodal embeddings allow systems to interpret the document exactly as a human does, maintaining the integrity of spatial relationships.

  2. Illuminating the visual modality: Enterprise documents are filled with visual indicators: technical charts, process flowcharts, branding assets, and product photography. Multimodal capabilities ensure that these elements are no longer invisible to search systems, allowing users to query images and text simultaneously.

  3. Connecting hybrid file formats: Real-world business workflows rarely live in a single document format. An agent may need to cross-reference a PDF policy, a spreadsheet tracking log, and a presentation deck. Extending RAG with multimodal embeddings creates a unified understanding across these varied formats.

The Architectural Solution: Gemini Multimodal Embeddings 2

Google Cloud’s Gemini Multimodal Embeddings 2 introduces a unified, multimodal vector space capable of embedding text, raster images, document pages, rendered spreadsheet tables, and visual charts into the same semantic representation space.

<div class="article-module h-c-page">
  <div class="h-c-grid">


<figure class="article-image--large
  
  
    h-c-grid__col
    h-c-grid__col--6 h-c-grid__col--offset-3
    
    
  ">

  
  
    
    <img alt="GIF_1" src="https://storage.googleapis.com/gweb-cloudblog-publish/original_images/image_bu28HMu.gif" />
    
    </a>
  
</figure>


  </div>
</div>

Key product capabilities unlocked by gemini-embeddings-2:

  • Crossmodal retrieval (text-to-visual / visual-to-text): Enables natural language queries to retrieve highly specific visual components, such as locating a target chart or diagram within a massive library of slides, without requiring manual tagging.

  • Layout-aware document embedding: Rather than breaking files into arbitrary text blocks, the system can embed document page renderings directly, preserving visual hierarchies, callout boxes, and structural context.

  • Heterogeneous format bridging: Native support for seamlessly bridging content across .docx, .xlsx, .pdf, .pptx, .png, and .csv without losing modality-specific structural information.

Three core patterns of multimodal enterprise agents

By leveraging multimodal embeddings within Box, we have identified three uniqueprimary design patterns that illustrate how organizations can extend traditional RAG to support complex, visual workflows.

Pattern 1: Complex financial & analytical reporting

The challenge

Corporate finance, research, and audit teams analyze highly structured documents where vital data resides in embedded tables, growth charts, and footnote annotations. Text-only indexing can separate these numbers from their context, making automated analysis challenging.

The multimodal advantage

  • Structural alignment: The embedding model captures the physical structure of tables and charts, allowing financial agents to understand that a column header applies to a specific row of metrics.

  • Visual trend analysis: Agents can cross-reference written summaries with visual trends in accompanying bar or line charts, identifying and pointing out discrepancies between written claims and source data.

  • Contextual sourcing: Users can query complex portfolios and instantly retrieve the exact page, table, or chart supporting a specific metric.

<div class="article-module h-c-page">
  <div class="h-c-grid">


<figure class="article-image--large
  
  
    h-c-grid__col
    h-c-grid__col--6 h-c-grid__col--offset-3
    
    
  ">

  
  
    
    <img alt="2" src="https://storage.googleapis.com/gweb-cloudblog-publish/images/2_9rTykxw.max-1000x1000.png" />
    
    </a>
  
</figure>


  </div>
</div>

Pattern 2: Multimodal clinical decision support & assisted diagnosis

The challenge

In healthcare and clinical environments, critical patient data is fragmented across vastly different, unstructured visual and textual formats — ranging from external physical photos (visual evidence) and microscopic pathology slides (lab reports) to structured risk matrices (triage grids). Traditional text-based systems or isolated analysis tools cannot synthesize these cross-modal relationships simultaneously, which can delay critical diagnoses or risk missing immediate, life-threatening procedural complications.

The multimodal advantage

  • Cross-modal clinical synthesis: Evaluates physical symptoms alongside cellular-level laboratory evidence simultaneously by indexing clinical photos, histopathology imagery, and triage grids into a single space.

  • Granular anomaly identification: Connects niche visual patterns under a microscope (like parasitic cyst walls) with medical knowledge to rapidly isolate rare conditions.

  • Risk-aware decision support: Cross-references findings against triage frameworks to deliver instant warnings about immediate patient risks, such as life-threatening anaphylactic shock.

<div class="article-module h-c-page">
  <div class="h-c-grid">


<figure class="article-image--large
  
  
    h-c-grid__col
    h-c-grid__col--6 h-c-grid__col--offset-3
    
    
  ">

  
  
    
    <img alt="3" src="https://storage.googleapis.com/gweb-cloudblog-publish/images/3_ZPNWwdP.max-1000x1000.png" />
    
    </a>
  
</figure>


  </div>
</div>

Pattern 3: Cross-document multimodal synthesis & data reconciliation

The challenge

Enterprise information is fragmented across disconnected files and formats (e.g., PDF minutes, Excel charts, PNG flyers, and email threads). Traditional tools analyze these files in isolation, failing to connect the dots when verifying details or resolving data contradictions across independent documents.

The multimodal advantage

  • Cross-file synthesis: Connects information across entirely different formats (PDFs, spreadsheets, images, emails) simultaneously to answer complex business queries.

  • Conflict resolution: Flags and resolves contradictions between assets, such as catching outdated pricing on an image by cross-checking it against the latest financial spreadsheets.

  • Visual-to-text auditing: Audits visual or scanned files against text-based records (e.g., verifying a signed PDF contract against a legal review email) to catch missing clauses or changes.

<div class="article-module h-c-page">
  <div class="h-c-grid">


<figure class="article-image--large
  
  
    h-c-grid__col
    h-c-grid__col--6 h-c-grid__col--offset-3
    
    
  ">

  
  
    
    <img alt="4" src="https://storage.googleapis.com/gweb-cloudblog-publish/images/4_IAwu96l.max-1000x1000.png" />
    
    </a>
  
</figure>


  </div>
</div>

The future of agentic enterprise content management

The integration of gemini-embeddings-2 into Box’s Agentic Platform is an important new capability to improve the next era of content intelligence. Multimodal embeddings help Box to move beyond basic search to active, intelligent collaboration.Box's Intelligent Content Management platform represents a fundamental shift in enterprise AI infrastructure — moving beyond passive document storage to deliver a governed, semantically indexed reasoning layer where AI agents can interrogate, cross-reference, and act on content with full compliance and security controls already in place. 

Powered by multimodal embeddings and a suite of native AI agents spanning search, metadata extraction, research, analysis, and composition, Box enables organizations to proactively surface insights such as flagging stale pricing data, expiring contract clauses, or cross-document contradictions before they become business risks. For high-complexity industries like financial services, life sciences, and legal operations, Box's ability to reason across text, tables, charts, and images makes multimodal understanding a competitive requirement. 

Designed to interoperate with the broader enterprise AI ecosystem, Box serves as the single governed content foundation that ensures every AI-driven workflow is grounded in authorized, auditable enterprise data.

When you think about it, the enterprise data landscape was always multimodal. Now we have the technology to make the most of it. By integrating gemini-embeddings-2, Box helps its users unlock unprecedented value from unstructured enterprise content. Product leaders who embrace multimodal-first architectures, rigorous precision benchmarking, and audit-ready grounding will lead the next wave of enterprise productivity and innovation.

The team would like to thank Ken Ikeda, Afshaan Mazagonwalla, and Samip Thakkar for their work on this project.

Building cost-effective, high-throughput gen AI workflows in Google Dataflow

詳細を表示

Real-time streaming pipelines are the operational backbone of modern enterprises, continuously processing everything from customer support interactions to transaction logs. Traditionally, streaming DAGs are static; once deployed, their processing logic and execution paths are fixed. However, by integrating generative AI agents, we can move beyond static logic to adaptive execution. This allows streaming workflows to dynamically construct plans, query databases, and trigger custom remediation paths at runtime depending on the content of the data.

For example, when a customer sends an angry message about a damaged order, a pipeline shouldn't just log the error or flag a dashboard. It should look up the order in the database that holds customer order and inventory records, decide on a remediation action (like shipping a replacement or issuing a refund), email the customer, and log the final resolution.

However, streaming systems face a fundamental engineering hurdle when executing gen AI workflows: scale, latency, and cost. Sending every raw event directly to a heavyweight model or multi-step agent equipped with external database and email tools is prohibitively expensive, introduces high latency, and quickly exhausts API rate limits.

This pattern addresses the scale and complexity challenge by combining Google Dataflow, Google Cloud's fully managed, serverless execution service for Apache Beam, and the Agent Development Kit (ADK) to build a hybrid streaming pipeline. By using a lightweight, CPU-bound machine learning model upstream to filter and qualify events, we keep the pipeline highly cost-effective, routing only the complex cases to the downstream agent. There, the agent dynamically decides what actions to take, introducing dynamic branching to the stream without hardcoding thousands of conditional steps into the pipeline's static DAG.

A universal blueprint for high-volume streams

While we use a customer support triage scenario below, this pre-filter + agentic action pattern is a universal paradigm. It applies to any stream where a high volume (>9X%) of events are routine, and only a small number require complex, contextual reasoning.

  • IT Operations & DevOps: Filtering millions of routine system logs on CPU, and triggering an agent to run diagnostics and open bug tickets only when a critical anomaly is flagged.

  • Financial Fraud Triaging: Passing millions of transactions through lightweight, local rules, and calling an agent to execute multi-database lookup tools only for highly suspicious patterns.

  • Industrial IoT: Monitoring normal telemetry on the edge, and routing erratic spikes to an agent to coordinate equipment shutdowns and email field engineers.

The architecture: Why pre-filter streaming events?

In a high-throughput stream, the vast majority of messages do not require complex reasoning or remediation. They might be positive feedback, neutral inquiries, or simple queries.

Routing every single event to a heavyweight LLM workflow creates three primary bottlenecks:

  1. API cost: Frontier models charge per token. Under high throughput, cost scales linearly with stream volume.

  2. Latency: Multi-step workflows (which involve database lookups and external API calls) take seconds, creating a bottleneck in streaming DAGs.

  3. Quotas: External APIs have strict rate limits that streaming workers can easily exhaust.

To prevent this, we build a pre-filtered pipeline in Apache Beam/Dataflow:

<div class="article-module h-c-page">
  <div class="h-c-grid">


<figure class="article-image--large
  
  
    h-c-grid__col
    h-c-grid__col--6 h-c-grid__col--offset-3
    
    
  ">

  
  
    
    <img alt="image1" src="https://storage.googleapis.com/gweb-cloudblog-publish/images/image1_XYX8VCT.max-1000x1000.jpg" />
    
    </a>
  
</figure>


  </div>
</div>

Pipeline flow

  1. Ingestion: Read raw customer messages from Google Pub/Sub.

  2. Lightweight sentiment classifier (CPU): Run all messages through a lightweight, CPU-based Hugging Face model (distilbert-base-uncased-finetuned-sst-2-english) using Apache Beam’s RunInference transform. This executes locally on the Dataflow worker CPUs, avoiding external API costs.

  3. Pre-qualification Gate: A simple DoFn filters the stream. Messages with POSITIVE or NEUTRAL sentiment are acknowledged and dropped.

  4. Automated Remediation (ADK): If and only if a message is classified as NEGATIVE, we trigger the gen AI agent backed by gemini-3.5-flash using the ADKAgentModelHandler. The agent uses tools to look up the user in BigQuery, fetch orders, choose a remediation plan, and send a notification email via the Gmail API.

Adaptive execution: Making the Beam DAG dynamic

In traditional streaming architectures, the pipeline's Directed Acyclic Graph (DAG) is rigid. Once deployed to Dataflow, the sequence of transforms is set. If you need to handle new types of alerts or change how specific events are routed, you have to modify, test, and redeploy the entire pipeline.

By placing a gen AI agent downstream of our sentiment pre-filter, we introduce a dynamic, adaptive node inside the static DAG.

For the 95% of records that are positive or neutral, the pipeline runs along a fast, static path. But when the filter gates a negative record, the agent evaluates the payload and dynamically selects the correct sequence of API tools (e.g., database query, inventory check, or email notification) at runtime. This allows the pipeline to execute complex decision trees dynamically, eliminating the need to build and maintain thousands of hardcoded conditional branches in the static Apache Beam code.

Implementing the pipeline

Here is an example implementation in Apache Beam using the Google Agent Development Kit (ADK) and the RunInference framework.

1. Defining the lightweight sentiment model

We define the upstream CPU model using HuggingFacePipelineModelHandler. This model classifies sentiment into POSITIVE, NEUTRAL, or NEGATIVE on the worker instance.

code_block
<ListValue: [StructValue([('code', 'model_handler = HuggingFacePipelineModelHandler(\r\n task="sentiment-analysis",\r\n model="distilbert-base-uncased-finetuned-sst-2-english"\r\n)'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7ff5dd6cc8d0>)])]>

2. Building the heavyweight ADK agent

The ADK agent acts as our remediation assistant. We equip it with three tools:

  • lookup_user: Queries BigQuery for the customer's email.

  • lookup_orders: Queries BigQuery for the customer's orders and current product inventory.

  • send_email: Sends a remediation email to the customer using the Gmail API.

code_block
<ListValue: [StructValue([('code', 'def make_adk_tools(project: str, dataset: str = "sentiment_demo"):\r\n def lookup_user(user_id: int) -> dict:\r\n """Look up user information (email address) from BigQuery by user ID."""\r\n from google.cloud import bigquery\r\n\r\n client = bigquery.Client(project=project)\r\n query = (\r\n f"SELECT user_id, user_email "\r\n f"FROM `{project}.{dataset}.users` "\r\n f"WHERE user_id = @user_id"\r\n )\r\n job_config = bigquery.QueryJobConfig(\r\n query_parameters=[bigquery.ScalarQueryParameter("user_id", "INT64", user_id)]\r\n )\r\n try:\r\n results = list(client.query(query, job_config=job_config).result())\r\n if results:\r\n row = results[0]\r\n return {"user_id": row.user_id, "user_email": row.user_email}\r\n return {"error": f"No user found with user_id={user_id}"}\r\n except Exception as exc:\r\n return {"error": str(exc)}\r\n\r\n def lookup_orders(user_id: int) -> dict:\r\n """Look up a user\'s orders and current product inventory from BigQuery."""\r\n from google.cloud import bigquery\r\n\r\n client = bigquery.Client(project=project)\r\n query = (\r\n f"SELECT p.order_id, p.product_id, pr.remaining_inventory, pr.price "\r\n f"FROM `{project}.{dataset}.purchases` p "\r\n f"JOIN `{project}.{dataset}.products` pr ON p.product_id = pr.product_id "\r\n f"WHERE p.user_id = @user_id"\r\n )\r\n job_config = bigquery.QueryJobConfig(\r\n query_parameters=[bigquery.ScalarQueryParameter("user_id", "INT64", user_id)]\r\n )\r\n try:\r\n results = list(client.query(query, job_config=job_config).result())\r\n orders = [\r\n {\r\n "order_id": row.order_id,\r\n "product_id": row.product_id,\r\n "remaining_inventory": row.remaining_inventory,\r\n "price": float(row.price),\r\n }\r\n for row in results\r\n ]\r\n return {"orders": orders}\r\n except Exception as exc:\r\n return {"error": str(exc)}\r\n\r\n def send_email(to_address: str, subject: str, body: str) -> str:\r\n """Send a plain-text email to the customer via the Gmail API."""\r\n import google.auth\r\n import googleapiclient.discovery\r\n import email.mime.text\r\n import base64\r\n\r\n try:\r\n creds, _ = google.auth.default(\r\n scopes=["https://www.googleapis.com/auth/gmail.send"]\r\n )\r\n service = googleapiclient.discovery.build("gmail", "v1", credentials=creds)\r\n\r\n mime_msg = email.mime.text.MIMEText(body)\r\n mime_msg["to"] = to_address\r\n mime_msg["subject"] = subject\r\n raw = base64.urlsafe_b64encode(mime_msg.as_bytes()).decode("utf-8")\r\n service.users().messages().send(userId="me", body={"raw": raw}).execute()\r\n return "Email sent successfully"\r\n except Exception as exc:\r\n return f"Failed to send email: {exc}"\r\n\r\n return [lookup_user, lookup_orders, send_email]'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7ff5dd536bd0>)])]>

We configure the LlmAgent and package it in the ADKAgentModelHandler:

code_block
<ListValue: [StructValue([('code', 'adk_agent = LlmAgent(\r\n name="remediation_agent",\r\n model="gemini-3.5-flash",\r\n instruction=(\r\n "You are a customer service remediation assistant with access to "\r\n "BigQuery lookup tools and an email sending tool. "\r\n "When given a prompt describing a customer situation, follow the "\r\n "numbered steps exactly and use your tools to complete the task."\r\n ),\r\n tools=adk_tools,\r\n)\r\n\r\n# RunInference handler for the ADK agent\r\nadk_handler = ADKAgentModelHandler(agent=adk_agent)'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7ff5dd536250>)])]>

3. Assembling the Dataflow DAG

The entire pipeline is declared cleanly. The upstream sentiment inference feeds directly into the filtering step (FilterNegativeADK), which then conditionally executes the downstream ADKInference:

code_block
<ListValue: [StructValue([('code', 'with beam.Pipeline(options=pipeline_options) as p:\r\n # 1. Read from Pub/Sub and classify sentiment on CPU\r\n sentiment_results = (\r\n p\r\n | "ReadFromPubSub" >> beam.io.ReadFromPubSub(topic=known_args.input_topic)\r\n | "DecodeMessages" >> beam.Map(lambda x: x.decode(\'utf-8\'))\r\n | "SentimentInference" >> RunInference(model_handler)\r\n )\r\n\r\n # 2. Filter out non-negative sentiment and invoke the ADK Agent\r\n _ = (\r\n sentiment_results\r\n | "FilterNegativeADK" >> beam.ParDo(FilterNegativeAndPromptADK())\r\n | "ADKInference" >> RunInference(adk_handler)\r\n | "LogADKResults" >> beam.ParDo(LogADKResponse())\r\n )'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7ff5e418ddd0>)])]>

Cost and performance advantages

By introducing this filtering step, we gain major engineering and operational advantages:

1. Significant cost reductions

Instead of paying for Gemini input/output tokens on 100% of incoming events, we pay only for the fraction that represent negative customer sentiment (typically < 5% of messages). The other 95% are classified locally on CPU instances at zero incremental API cost.

2. High streaming throughput

Dataflow distributes the CPU classification workload across many instances. Since CPU inference takes milliseconds, the pipeline scales horizontally to handle high-throughput event streams. The heavyweight LLM agent, which can take seconds per request due to tool execution, is called sparingly, preventing backlog.

3. Native Apache Beam integration

Adding the agent into the DAG requires no complex orchestration logic or manual thread pools. Using ADKAgentModelHandler with Beam's native RunInference transform handles parallel worker threads, batching, and integration automatically, keeping the codebase maintainable and clean.

Key takeaways

Streaming data is fast and high-volume, while heavyweight generative AI reasoning is slow and costly.

By building a pre-filtered pipeline with Google Dataflow and the ADK, you get the best of both worlds: the cost and speed of local CPU-based models, and the deep, automated capabilities of Gemini-backed agents.

To see the complete codebase and deploy this yourself, check out the next-2026-demo GitHub repository.


Apache Beam is a trademark of the Apache Software Foundation