GCP News - 2026-08-11

2026-08-11
最終更新: 2026-08-27 21:31:31 JST

Google Cloud Release Notes

August 11, 2026

詳細を表示

Access Approval

Feature

Datastream is generally available (GA).

Access Transparency

Feature

Datastream is generally available (GA).

Apigee hybrid

Announcement

v1.16.9

On August 11, 2026 we released an updated version of the Apigee hybrid software, v1.16.9.

Fixed

Fixed in this release

Bug ID Description
514973778 Fixed an issue where the SanitizeUserPrompt and SanitizeModelResponse policies failed to tolerate unknown fields while parsing responses from the Model Armor Service.
543171828 Fixed an issue where the apigee-logger DaemonSet failed to schedule on cluster nodes without custom node labels due to a default logger.nodeSelector in the Helm chart.

Security

Various security and CVE fixes are included in this release.

BigQuery

Feature

Query templates for data clean rooms are generally available (GA). Query templates allow data clean room owners and publishers to share predefined queries without exposing the underlying tables and views.

Additionally, table parameters in table-valued functions (TVFs) are generally available (GA). You can use the ANY TABLE type as a table parameter to create generic functions that accept tables of any structure.

Cloud Run

Feature

Cloud Run NVIDIA L4 GPU driver version 580.x.x is available for services, jobs, and worker pools.

Cloud SQL for MySQL

Feature

The Cloud SQL remote MCP server now supports specialized endpoint URLs (toolsets) for /readonly, /instance_manage, and /query_execution. Specialized toolset URLs let you restrict the set of exposed MCP tools based on your security and workflow requirements.

For more information, see Available toolsets.

Feature

You can now use the create_instance tool in the Cloud SQL remote MCP server to provision free trial instances for testing and development by setting the free_trial parameter to true.

For more information, see Create a free trial instance.

Feature

When executing SQL queries using the execute_sql or execute_sql_readonly tool, setting the sql_commenter_enabled parameter to true automatically appends sqlcommenter tags (mcp.tool, mcp.server, user.identity, mcp.client) to SQL statements for enhanced database observability.

For more information, see sqlcommenter tags.

Cloud SQL for PostgreSQL

Feature

The Cloud SQL remote MCP server now supports specialized endpoint URLs (toolsets) for /readonly, /instance_manage, and /query_execution. Specialized toolset URLs let you restrict the set of exposed MCP tools based on your security and workflow requirements.

For more information, see Available toolsets.

Feature

You can now use the create_instance tool in the Cloud SQL remote MCP server to provision free trial instances for testing and development by setting the free_trial parameter to true.

For more information, see Create a free trial instance.

Feature

When executing SQL queries using the execute_sql or execute_sql_readonly tool, setting the sql_commenter_enabled parameter to true automatically appends sqlcommenter tags (mcp.tool, mcp.server, user.identity, mcp.client) to SQL statements for enhanced database observability.

For more information, see sqlcommenter tags.

Cloud SQL for SQL Server

Feature

You can now use the update_user tool in the Cloud SQL remote MCP server to update passwords for Cloud SQL for SQL Server database users.

For more information, see Available tools.

Feature

The Cloud SQL remote MCP server now supports specialized endpoint URLs (toolsets) for /readonly and /instance_manage. Specialized toolset URLs let you restrict the set of exposed MCP tools based on your security and workflow requirements.

For more information, see Available toolsets.

Compute Engine

Feature

Generally available: Compute flexible committed use discounts (CUDs) are available for G2 and G4 GPU accelerator-optimized machine series. The supported resources include vCPUs, memory, Local SSD disks, and GPUs.

Compute flexible CUDs are spend-based CUDs that apply to eligible Google Cloud spend across Compute Engine, GKE, and Cloud Run. For G2 and G4 machine series, compute flexible commitments provide the flexibility to switch between eligible machine series and regions depending on your workload needs. For GPUs that belong to these machine series, compute flexible commitments don't require attached reservations.

For more information, see Compute flexible CUDs.

Security

A vulnerability (CVE-2026-6726) in the Trusted Computing Group's TPM 2.0 reference implementation code was discovered and is being addressed. For more information, see the GCP-2026-054 security bulletin.

Confidential VM

Security

A vulnerability affecting Intel TDX firmware was discovered and is being addressed. For more information, see the GCP-2026-053 security bulletin.

Container Optimized OS

Change

cos-beta-133-19999-0-28

Kernel Docker Containerd GPU Drivers
COS-6.18.39 v29.4.3 v2.3.2 See List

Change

cos-129-19506-299-116

Kernel Docker Containerd GPU Drivers
COS-6.12.94 v27.5.1 v2.2.6 See List

Change

cos-dev-138-20035-0-0

Kernel Docker Containerd GPU Drivers
COS-6.18.41 v29.4.3 v2.3.2 See List

Feature

Added support for installing the Vast 4.5.8 NFS client drivers with cos-dkms.

Fixed

Added kernel patch to reduce bcache garbage collection sleep interval to prevent I/O stalls.

Feature

Added support for installing the Vast 4.5.8 NFS client drivers with cos-dkms.

Fixed

Added kernel patch to reduce bcache garbage collection sleep interval to prevent I/O stalls.

Fixed

Fixed CVE-2026-33186 in google-guest-agent.

Fixed

Added kernel patch to reduce bcache garbage collection sleep interval to prevent I/O stalls.

Fixed

Mask nfttables-restore.service to address time to ssh regression.

Security

Fixed CVE-2026-64227 in the Linux kernel.

Fixed

Mask nfttables-restore.service to address time to ssh regression.

Security

Fixed KCTF-8173f7e in the Linux Kernel.

Security

Fixed CVE-2026-64279 in the Linux kernel.

Fixed

Updated app-admin/node-problem-detector to v0.8.25.

Security

Fixed CVE-2026-64286 in the Linux kernel.

Security

Fixed KCTF-8173f7e in the Linux Kernel.

Security

Fixed CVE-2026-64287 in the Linux kernel.

Security

Upgraded net-libs/nghttp2 to 1.69.0 and fixed CVE-2026-58055.

Security

Fixed CVE-2026-64352 in the Linux kernel.

Security

Fixed CVE-2026-64375 in the Linux kernel.

Security

Fixed CVE-2026-64401 in the Linux kernel.

Security

Fixed CVE-2026-64413 in the Linux kernel.

Security

Fixed CVE-2026-64416 in the Linux kernel.

Security

Fixed CVE-2026-64476 in the Linux kernel.

Security

Fixed CVE-2026-64508 in the Linux kernel.

Security

Fixed CVE-2026-64530 in the Linux kernel.

Security

Fixed CVE-2026-64532 in the Linux kernel.

Security

Fixed CVE-2026-64533 in the Linux kernel.

Security

Fixed CVE-2026-64534 in the Linux kernel.

Security

Fixed CVE-2026-64535 in the Linux kernel.

Security

Fixed CVE-2026-64538 in the Linux kernel.

Security

Fixed CVE-2026-64542 in the Linux kernel.

Security

Fixed CVE-2026-64545 in the Linux kernel.

Security

Fixed CVE-2026-64546 in the Linux kernel.

Security

Fixed CVE-2026-64548 in the Linux kernel.

Security

Fixed CVE-2026-64552 in the Linux kernel.

Security

Fixed CVE-2026-64554 in the Linux kernel.

Security

Fixed CVE-2026-64555 in the Linux kernel.

Security

Fixed KCTF-8173f7e in the Linux Kernel.

Change

Runtime sysctl changes:

  • Changed: net.ipv4.udp_mem: 188034 250714 376068 -> 188034 250715 376068

Change

cos-121-18867-528-58

Kernel Docker Containerd GPU Drivers
COS-6.6.143 v27.5.1 v2.0.10 See List

Fixed

Update dev-lang/go to 1.25.12.

Security

Fixed CVE-2026-64279 in the Linux kernel.

Security

Fixed CVE-2026-64319 in the Linux kernel.

Security

Fixed CVE-2026-64352 in the Linux kernel.

Security

Fixed CVE-2026-64375 in the Linux kernel.

Security

Fixed CVE-2026-64401 in the Linux kernel.

Security

Fixed CVE-2026-64413 in the Linux kernel.

Security

Fixed CVE-2026-64474 in the Linux kernel.

Security

Fixed CVE-2026-64476 in the Linux kernel.

Security

Fixed CVE-2026-64535 in the Linux kernel.

Security

Fixed KCTF-8173f7e in the Linux Kernel.

Change

cos-117-18613-675-48

Kernel Docker Containerd GPU Drivers
COS-6.6.143 v24.0.9 v1.7.34 See List

Fixed

Update dev-lang/go to 1.25.12.

Security

Fixed CVE-2026-64279 in the Linux kernel.

Security

Fixed CVE-2026-64319 in the Linux kernel.

Security

Fixed CVE-2026-64352 in the Linux kernel.

Security

Fixed CVE-2026-64375 in the Linux kernel.

Security

Fixed CVE-2026-64401 in the Linux kernel.

Security

Fixed CVE-2026-64413 in the Linux kernel.

Security

Fixed CVE-2026-64535 in the Linux kernel.

Security

Fixed CVE-2026-64548 in the Linux kernel.

Security

Fixed CVE-2026-64556 in the Linux kernel.

Security

Fixed KCTF-8173f7e in the Linux Kernel.

Cortex Framework

Announcement

Release 7.0.2

Fixed

  • Resolved security vulnerabilities in transitive dependencies by updating the following corresponding direct dependencies: google-auth, google-cloud-bigquery, google-cloud-dataform, google-cloud-resource-manager, google-cloud-service-usage and google-cloud-storage.

Firestore

Feature

Firestore now supports the asia-southeast3 Bangkok region.

For a full list of supported locations, see Locations.

Firestore in Datastore mode

Feature

Firestore in Datastore mode (Datastore) now supports the asia-southeast3 Bangkok region.

For a full list of supported locations, see Locations.

Gemini Enterprise

Feature

Gemini Enterprise: New data stores and support for new actions (Public Preview)

The following data stores are available in Public Preview in Gemini Enterprise:

You can search and read data from these data stores using natural language.

Additionally, the following data stores support new actions in Public Preview:

  • Airtable: Update records for a table.
  • Hex: Create threads and continue threads.
  • Miro: Create documents and update documents.
  • Smartsheet: Add rows.

Feature

Gemini Enterprise: Manage overages, spend limits, and costs for invoiced Cloud Billing accounts

If your project has an invoiced Cloud Billing account and at least one active, non-free-trial subscription, administrators can enable overages, configure monthly spend limits, and monitor feature usage and costs in Gemini Enterprise:

  • Enable overages: Allow users to continue using features at pay-as-you-go rates after reaching pooled quotas. Overages are supported for Standard, Plus, and Standard Emerging Market editions for customers with an invoiced Cloud Billing account and at least one active, non-free-trial subscription.
  • Set spend limits: Configure monthly project spending caps and budget alert thresholds in Cloud Billing to prevent unexpected charges.
  • View feature usage and costs: Track pooled quota consumption, pay-as-you-go usage, and 30-day billing trends on the Usage & Spending page in the Gemini Enterprise console and Cloud Billing console.

For more information, see:

Google Distributed Cloud (software only) for VMware

Announcement

Google Distributed Cloud (software only) for VMware 1.35.400-gke.81 is now available for download. To upgrade, see Upgrade clusters. Google Distributed Cloud 1.35.400-gke.81 runs on Kubernetes v1.35.3-gke.400.

If you use a third-party storage vendor, check the listing of our previously-qualified storage partners.

After a release, it takes approximately 7 to 14 days for the version to become available for use with GKE On-Prem API clients: the Google Cloud console, the gcloud CLI, and Terraform.

Fixed

The following issues were fixed in 1.35.400-gke.81:

  • Link to Vulnerability fixes for the list of security vulnerabilities addressed in this release.
  • Fixed an issue where user clusters remained stuck in a Reconciling state after an admin cluster upgrade. The admin cluster controller skipped reconciling legacy cluster lifecycle components during upgrades unless an initial migration annotation was set. If legacy user clusters still existed on the admin cluster, missing legacy API discovery (cluster.k8s.io/v1alpha1) caused controller reconciliation to stall. With this fix, the controller preserves legacy components as long as any legacy user clusters exist, and prunes them only after all user clusters have migrated to advanced clusters.
  • Fixed an issue where gkectl prepare failed with a permission denied error when authenticating against a private container registry.
  • Fixed an issue where retrying a user or admin cluster upgrade to advanced clusters caused etcd secret decryption failures.

Google Distributed Cloud (software only) for bare metal

Announcement

Google Distributed Cloud (software only) for bare metal 1.35.400-gke.81 is now available for download. To upgrade, see Upgrade clusters. Google Distributed Cloud for bare metal 1.35.400-gke.81 runs on Kubernetes v1.35.3-gke.400.

After a release, it takes approximately 7 to 14 days for the version to become available for installations or upgrades with the GKE On-Prem API clients: the Google Cloud console, the gcloud CLI, and Terraform.

If you use a third-party storage vendor, check the listing of our previously-qualified storage partners.

Fixed

The following issues were fixed in 1.35.400-gke.81:

  • Link to Vulnerability fixes for the list of security vulnerabilities addressed in this release.

Spanner

Feature

For DML statements, Spanner now enforces the 80,000 mutation limit (including indexes) per statement rather than cumulatively across the transaction. This allows a transaction to execute multiple DML statements that collectively exceed the limit, as long as each individual statement remains under it. For the Mutation API, the 80,000 limit applies to all mutations in the commit. All transactions are still subject to the 100 MiB commit size limit.

Google Cloud Blog (AI & ML)

How WPP operationalizes platform and data engineering for AI marketing

詳細を表示

Between chaotic levels of market fragmentation and economic volatility, marketing and communications agencies can no longer rely on the human intuition they’ve traditionally used to win clients and optimize their ad spend. WPP is replacing that guesswork with an AI-powered view of shifting market dynamics, giving brands predictive certainty that lets them invest with confidence while moving at the speed of the market. That’s the value of WPP Open, its agentic marketing system.

But before it could begin applying sophisticated AI models to power those insights, WPP had to overcome a critical engineering challenge: the marketing data that made up the models was fragmented across hundreds of global agencies. While this dynamic made it nearly impossible to deploy AI tools efficiently and securely, access to models was only part of the equation. And  until it built a reliable way to ingest, clean, and serve data to those models, WPP couldn’t unlock the true potential of generative AI.

To solve this, WPP partnered with Google Cloud to construct a unified data backbone and  custom platform engineering path. Now, by standardizing its serverless compute patterns and data processing workflows, WPP is able to  securely deploy targeted marketing campaigns in days instead of months.

Architecting a centralized, service-based data foundation 

An important part of this effort was accelerating data availability and centralizing management. To do this, WPP adopted a service-based project structure for its current production environment. Rather than isolating every workload into separate silos, its engineering team centralized Google Cloud Storage (GCS) and BigQuery into dedicated, shared data projects, while also segregating the compute and processing workloads into distinct processing projects.

This structure simplified the core team’s user experience and ensured that all data consumers interacted with a unified source of truth. Because data from WPP’s various product lines lives in shared infrastructure, it was essential that security be strictly enforced at a granular level. By directly applying identity and access management (IAM) controls at the individual GCS bucket and BigQuery dataset levels, the company’s teams only see the data they’re  authorized to access.

At the same time, raw data from various partners lands in dedicated GCS buckets in order to keep the raw inputs organized and isolated. From there, Managed Service for Apache Spark executes custom Apache Spark jobs to cleanse, normalize, and canonicalize information into standardized cohort definitions (SCDs). By utilizing a serverless architecture combined with Kubeflow for pipeline orchestration, WPP’s data engineering team avoided the overhead that often results from managing cluster infrastructure. This allowed them to focus entirely on the data transformation logic fueling the downstream GCS and BigQuery layers  that ultimately feed the company’s audience & performance AI models.

<div class="article-module h-c-page">
  <div class="h-c-grid">


<figure class="article-image--large
  
  
    h-c-grid__col
    h-c-grid__col--6 h-c-grid__col--offset-3
    
    
  ">

  
  
    
    <img alt="1 WPP Data Pipeline Architecture" src="https://storage.googleapis.com/gweb-cloudblog-publish/original_images/wpp_data_flow_architecture.jpg" />
    
    </a>
  
</figure>


  </div>
</div>

What made our collaboration with Google Cloud successful was the balance they struck between uncompromising professionalism when it comes to best practices and timely delivery of incredibly pragmatic, real-world solutions.
- Jonas Dahlbaek
Senior Data Engineering Lead, WPP

Standardizing data into unified cohorts

 When raw data enters WPP’s processing zone, its platform converts it into SCDs that become core concepts used throughout the framework for keying purposes. These are based on five keys: age, gender, geo, product, and interest. But these underlying data definitions are fluid and continuously canonicalized to reflect evolving marketing concepts. As a result, this uniform structure allows WPP to join and aggregate data on a global scale without exposing sensitive underlying particulars or relying on shared identifiers.

The platform's core processing engine was built in type-safe Scala to ensure comprehensive visibility and compliance This custom framework tightly controls how data is transformed, and it inherently supports full source traceability while guaranteeing that every data point within the curated datasets can be traced back to its origin. This is a crucial level of traceability when building enterprise AI applications, as data scientists and auditors must understand exactly what information feeds into the models, even as WPP concurrently prepares to transition to Google Cloud Knowledge Catalog for automated, enterprise-wide data governance in the future.

Working with Google Cloud has been instrumental in accelerating and standardizing our engineering efforts. In a world where massive volumes of fragmented data present a daily challenge, having the right infrastructure is paramount to thriving in the AI age and helps our developers and AI marketers alike.
-Suleman Khan
Product Manager for OI & Google Partnerships, WPP

Standardizing the enterprise software lifecycle

For WPP, even with all these steps in place, processing data is only half the battle. To serve applications and manage the underlying infrastructure, the company’s platform engineering team developed a suite of reusable and centralized GitLab continuous integration and continuous deployment (CI/CD) templates. With this, WPP reduced the cognitive load on individual development teams and ensured that all deployments met strict corporate security standards.

These templates manage various enterprise workloads autonomously. The suite includes universal Cloud Run templates for full-stack web applications and  batch data processing and scheduled pipelines. It also includes a deploy-only template for multi-stage workflows and a Cloud Run functions deployment template for event-driven microservices.

<div class="article-module h-c-page">
  <div class="h-c-grid">


<figure class="article-image--large
  
  
    h-c-grid__col
    h-c-grid__col--6 h-c-grid__col--offset-3
    
    
  ">

  
  
    
    <img alt="2 WPP Cloud Platform Engineering" src="https://storage.googleapis.com/gweb-cloudblog-publish/images/2_WPP_Cloud_Platform_Engineering.max-1000x1000.png" />
    
    </a>
  
</figure>


  </div>
</div>

Implementing zero-rebuild promotion

Rebuilding container images in a production environment can introduce unnecessary risk and the potential for configuration drift. In order to maintain environmental consistency, WPP embraced a "build once, deploy many" methodology that applied cross-project IAM logic and Google Cloud Artifact Registry configurations.

As part of this process, developers build and test container images in the development environment. Once those exact, immutable container images are validated, they’re promote  directly to production. This zero-rebuild promotion ensures total parity across deployment stages and eliminates unexpected production behaviors. The CI/CD templates also facilitate progressive traffic migration, which allowed teams to route a small percentage of traffic to new revisions before initiating a full rollout.

Immutable deployments. Traceable data. Unshakable trust. When you know exactly what goes into your AI, you can ship at the speed of light.
- Ranjith K Poldas
Associate Director , Devops (I&P), WPP Media

Automating security and intelligent networking

With this modern architecture, enterprise security acts as a foundational enabler for WPP, so it integrated Wiz security scanning directly into the pre-push phase of the CI/CD pipeline to catch vulnerabilities before code merges. The company also utilized Google Cloud Identity-Aware Proxy to enforce zero-trust access across its  internal applications.

To further simplify operations, WPP adopted templates with intelligent virtual private cloud (VPC) logic. This configuration automatically identifies and resolves networking conflicts between legacy VPC connectors and modern Direct VPC access. This automated networking prevents deployment failures and accelerates the release cycle.

Monitoring operational health and driving ROI

Because a resilient platform foundation requires deep observability, WPP’s engineering team now monitors strict operational metrics instead of relying solely on deployment frequency. The team tracks request latency across p50, p95, and p99 percentiles, alongside 4xx and 5xx error rates. It  also monitors container startup times to mitigate cold starts, while tracking overall CPU and memory utilization. This granularity ensures that both data pipelines and serverless infrastructure always remain highly available.

"Navigating a transformation of this scale across multiple complex workstreams—spanning data engineering, platform infrastructure, and AI integration—required more than just alignment; it demanded deep, mutual trust. Working as true partners, Google Cloud and WPP moved in lockstep to deliver production-ready platform capabilities on time."
Yang Yue , Program Manager , Google Cloud

For WPP, operationalizing its data and AI stacks at this velocity provided the necessary infrastructure for its advanced workloads, and the business impact was clear and quantifiable. By building this dual foundation, the company reduced creative and strategy time from four weeks to just three hours. It also saw a 70% gain in production efficiency, a 33x increase in content volume, and  a 2.8x increase in campaign return on investment. In short, by partnering with Google Cloud and implementing a broad suite of products and tools, WPP was able to quickly realize a significant ROI and boost productivity, efficiency, reliability, and security across the company.

How Malachyte solves retail’s cold-start problem with managed real-time AI

詳細を表示

What’s the best way to recommend products to little-known users? 

We’ve spent our careers trying to solve this problem for major companies like Spotify and Priceline, and it’s why Sidd founded Malachyte, an AI-powered ecommerce recommendation platform. These days, consumers have come to expect content that feels personalized and relevant, and online services competing for their attention have no choice but to do this exceptionally well.  

Malachyte was inspired by some unique insights into how advanced AI models, and large language models in particular, could be applied in new ways to old challenges like personalization and recommendations. 

As Malachyte set out to win potential customers’ business, we needed secure, scalable, reliable and, above all, leading-edge AI infrastructure to continue building the personalization algorithm we had always envisioned. By utilizing Google Cloud tools like Bigtable and Managed Service for Apache Kafka, Malachyte has been able to help some of its retailers double and sometimes even triple their sales. 

This is the story of how we built it, and the ways any founder can use services like these to start deploying AI foundation models in new ways.

How Malachyte lifted sales for their users 

For Malachyte, the aha moment was discovering that it could use neural networks with attention mechanisms — the same concept powering large language models — to personalize retail search and product pages. This approach is what enables LLMs to derive meaning from the relative order of items in a sequence, in their case the order of words and syllables in a sentence. When it comes to a retail website or app, what Malachyte wanted to capture was the sequence of customer interactions with the site.

<div class="article-module h-c-page">
  <div class="h-c-grid">


<figure class="article-image--large
  
  
    h-c-grid__col
    h-c-grid__col--6 h-c-grid__col--offset-3
    
    
  ">

  
  
    
    <img alt="1 - Malachyte blog" src="https://storage.googleapis.com/gweb-cloudblog-publish/images/1_-_Malachyte_blog_.max-1000x1000.png" />
    
    </a>
  
    <figcaption class="article-image__caption "><p>What if we predicted the next thing a user wants on an ecommerce website just like LLMs predict the next word in a sentence?</p></figcaption>
  
</figure>


  </div>
</div>

A pre-GPT language model might have tried to look at a specific sequence of words or even fragments of words (what we now know of as tokens), but those earlier models wouldn’t examine what happens if the words were in the comparable order but weren’t contiguous or were re-arranged. The breakthrough came — in part through Google’s work on transformers — when LLMs gained the ability to understand complex and long-range dependencies within a sequence of items. 

This more sophisticated method has delivered dramatic results — both for the proliferation of gen AI in general, and for Malachyte’s application of the technology.

To make this work in practice, Malachyte creates a vector of everything known about a visitor when they arrive on a site.  Most users are visiting for the first time, so little is known about them. This is what’s known as the  “cold start” problem. The trick is to use every interaction with a user to refine this vector. Each new addition to the vector, like a click or a query, does two things: it drives a prediction about the next thing the user wants, and it provides more information about the user.  

Malachyte’s platform then updates the user vector and the prediction at the same time. This not only enhances the understanding of the individual user and their preferences, it also improves the overall model with the anonymized user data. With every inference, the context of both the average and the specific shopper grows. 

The company further innovates by not just using attention-based neural networks but combining that with updating user profiles 100 milliseconds at time.

<div class="article-module h-c-page">
  <div class="h-c-grid">


<figure class="article-image--large
  
  
    h-c-grid__col
    h-c-grid__col--6 h-c-grid__col--offset-3
    
    
  ">

  
  
    
    <img alt="2 - Malachyte Blog" src="https://storage.googleapis.com/gweb-cloudblog-publish/images/2_-_Malachyte_Blog.max-1000x1000.png" />
    
    </a>
  
    <figcaption class="article-image__caption "><p>Malachyte’s recommendation and search agents populate the next page’s search results or recommendation carousels based on what users clicked on previous pages.</p></figcaption>
  
</figure>


  </div>
</div>

To be sure, this idea isn’t in itself new. Retailers have long used collaborative filtering recommendations systems to identify similar users and items that required massive sets of interaction history. These models typically required a lot of data, including third-party cookie-based profiles and demographics. 

By focusing on the sequence of interactions in a session, retailers can achieve far more personalization — with less required data or spend — than by focusing only on a user’s profile. As a bonus, retailers can now offer their users more privacy by not relying on long-term cookie data.

This works because of the model structure and multimodal vectors that encode everything they know about a user, including browser data, click history and searches. The output, too, is multimodal: The same model can be applied to on-site search product pages, category pages, and add-to-cart carousels.  

To make this work, each product in the catalog is embedded into the same space as the user vector, which gets updated and subsequently moves the vector closer to relevant products and further from those that aren’t. The neural network computing the embedding is being continuously trained across retailers who work with Malachyte, improving the quality for everyone. The system effectively becomes a data cooperative with each retailer's user helping make the model smarter for everyone.

<div class="article-module h-c-page">
  <div class="h-c-grid">


<figure class="article-image--large
  
  
    h-c-grid__col
    h-c-grid__col--6 h-c-grid__col--offset-3
    
    
  ">

  
  
    
    <img alt="3 - Malachyte blog vector space - high res" src="https://storage.googleapis.com/gweb-cloudblog-publish/images/3_-_Malachyte_blog_vector_space_-_high_res.max-1000x1000.png" />
    
    </a>
  
    <figcaption class="article-image__caption "><p>A user session represented as a vector in a space of products.</p></figcaption>
  
</figure>


  </div>
</div>

To make this delivery for every user at every inference in 100 milliseconds, Malachite found real benefits in building onGoogle Cloud’s real-time AI stack.   

With this system, every behavioral event streams into a Managed Service for Apache Kafka cluster. Rather than queuing for a future training job, each event immediately becomes an update to the user’s profile in Bigtable. The Kafka cluster allows the customer’s front-end to persist, so the user session signals quickly with little worry about how they fit into the user vector.  

Bigtable allows Malachyte’s services to look up and update the right user vectors, and it and Kafka operate at the order of 10 milliseconds per step, which allows the entire recommendation loop to complete with no disruption to the user experience. 

<div class="article-module h-c-page">
  <div class="h-c-grid">


<figure class="article-image--large
  
  
    h-c-grid__col
    h-c-grid__col--6 h-c-grid__col--offset-3
    
    
  ">

  
  
    
    <img alt="4 - Malachyte Blog" src="https://storage.googleapis.com/gweb-cloudblog-publish/images/4_-_Malachyte_Blog.max-1000x1000.png" />
    
    </a>
  
    <figcaption class="article-image__caption "><p>The three layer real-time AI architecture: a retailer’s website, Malachyte’s AI models and serving front ends, and context management infrastructure.</p></figcaption>
  
</figure>


  </div>
</div>

In addition to a fast core, a second layer of product catalog updates, inventory signals, and retailer-specific dimensional data keeps product data up to date. This flows through Cloud Pub/Sub, which offers globally accessible REST APIs that enable connections retailers can use without deep integration work. Malachyte agents run on Google Kubernetes Engine (GKE), with model inference on Google Compute Engine (GCE).

<div class="article-module h-c-page">
  <div class="h-c-grid">


<figure class="article-image--large
  
  
    h-c-grid__col
    h-c-grid__col--6 h-c-grid__col--offset-3
    
    
  ">

  
  
    
    <img alt="5 Malachyte blog" src="https://storage.googleapis.com/gweb-cloudblog-publish/images/5_Malachyte_blog.max-1000x1000.png" />
    
    </a>
  
    <figcaption class="article-image__caption "><p>Continuous ingestion of external data, such as product catalog updates, operates through Pub/Sub’s global messaging system.</p></figcaption>
  
</figure>


  </div>
</div>

With its migration to Google Cloud’s AI architecture, Malachyte demonstrated that production AI inference and training are about more than GPUs and storage. They require real-time continuous learning infrastructure that includes a fast key-value store, a streaming layer, and a managed messaging system, all integrated with the foundation model architecture. 

This approach also shows that even a small team like Malachyte’s can have a big impact in an industry. It just needs access to powerful infrastructure and core AI managed services.

Try it for yourself 

Looking to shake up your industry or stay ahead of the competition like Malachyte? Try Managed Service for Apache Kafka, Cloud Pub/Sub, and Bigtable. New customers can receive $300 in Google Cloud credits.

Google named a Leader in The Forrester Wave™: AI Platforms, Q3 2026

詳細を表示

At Google Cloud, we help organizations of all sizes build and operationalize complex agentic workflows with total confidence. By combining world-class AI research with an open, fully integrated AI platform, we give customers the flexibility to innovate and the foundation to deliver measurable business value. 

At the center of it all is Gemini Enterprise, a unified platform designed to power the agentic enterprise, meet builders where they are, and deliver enterprise trust by default

We believe this integrated approach is why Google has been named a Leader in The Forrester Wave™: AI Platforms, Q3 2026 report, and received the highest score in the Strategy category.

<div class="article-module h-c-page">
  <div class="h-c-grid">


<figure class="article-image--large
  
  
    h-c-grid__col
    h-c-grid__col--6 h-c-grid__col--offset-3
    
    
  ">

  
  
    
    <img alt="Image_AI-Platforms,-Q3-2026" src="https://storage.googleapis.com/gweb-cloudblog-publish/images/Image_AI-Platforms-Q3-2026.max-1000x1000.png" />
    
    </a>
  
</figure>


  </div>
</div>

Powering the Agentic Era

As agents become embedded across every part of the business, organizations need a unified platform. Gemini Enterprise serves as the front door to AI for your entire organization, removing the silos between business users, developers, and IT leaders by connecting them all with shared, universal context.

Gemini Enterprise Agent Platform is the foundation that enables technical teams and IT leaders to safely, securely, and cost-effectively build and deploy production-grade agents. Any agent built in Agent Platform can be deployed across your entire workforce through the Gemini Enterprise app, putting custom agentic capabilities directly into the hands of every employee. 

By unifying enterprise data, frontier model capabilities, developer tooling, and IT operations under one roof, Gemini Enterprise empowers teams to transform products, services, and complex agentic tasks while maintaining centralized control every step of the way.

Meeting builders where they are

No two development teams build agents in the same way. Some work in coding environments, others rely on low-code tools, and many organizations use a mix of models. Gemini Enterprise is built for that reality. We offer a platform tailored to every team's skill set, enabling high-code developers to build complex agentic systems while giving business and operational teams low-cod and no-code tools to rapidly design and test agent behaviors. And with access to over 200 native and third-party models, alongside pre-built agent templates, engineering teams can move from prototype to deployment with speed. 

True enterprise AI should extend far beyond text. Gemini Enterprise is multi-modal by design, meaning it natively understands, reasons across, and generates text, code, audio, image, and video inputs in a single workflow. By securely connecting this multi-modal intelligence to your enterprise data wherever it lives, your engineering teams can turn this information into contextual business experiences.

Enterprise trust, built in

Trust is foundational to AI adoption, which is why Agent Platform is built with enterprise-grade governance, security, and observability by default. To help you scale with confidence, Agent Platform features built-in guardrails, continuous evaluation, and real-time tracking for costs, latency, and token usage. This operational transparency gives organizations the control they need to safely expand agentic workflows across the business. Builders can also leverage capabilities like Knowledge Catalog to establish a universal context engine across the enterprise, aggregating metadata with zero-copy federation to improve agent accuracy.

Looking ahead

We believe this recognition reflects our commitment to delivering an open, scalable, and powerful AI platform. As agentic systems reshape how enterprise software is built, Gemini Enterprise will continue to provide the foundation developers need to build with confidence.

Want to dive deeper into the report? Access the full Forrester Wave™: AI Platforms, Q3 2026 report here.


Forrester does not endorse any company, product, brand, or service included in its research publications and does not advise any person to select the products or services of any company or brand based on the ratings included in such publications. Information is based on the best available resources. Opinions reflect judgment at the time and are subject to change. This report is part of a broader collection of Forrester resources, including interactive models, frameworks, tools, data, and access to analyst guidance. For more information, read about Forrester’s objectivity here.