New Relic integrations with the Google Cloud Platform (GCP) include one that reports Google Cloud Vertex AI data to New Relic. This document explains how to activate the GCP Vertex AI integration and describes the data it reports.
Features
Vertex AI is Google Cloud's unified platform for building, deploying, and managing machine-learning models and generative-AI applications. New Relic Vertex AI integration collects performance, throughput, latency, and resource utilization metrics across locations, endpoints, indexes, feature stores, and feature online stores.
Activate integration
To enable the integration, follow standard procedures to connect your GCP service to New Relic:
Polling frequency
New Relic integrations query your GCP services according to a polling interval that varies by integration. The polling frequency for Google Cloud Vertex AI is 5 minutes. The resolution is 1 data point every minute.
Workload Identity Federation
Find and use data
After you enable the integration, your Vertex AI resources appear as entities in the New Relic entity explorer. To see dashboards and manage services, go to one.newrelic.com > All capabilities > Infrastructure > GCP.
All Vertex AI metrics available in GCP Cloud Monitoring are collected as dimensional metrics in the Metric event type. Additional metrics beyond this table are collected automatically. See Google's Vertex AI metrics documentation for the complete list.
Entities
Metric data
Key metrics — Location
Metric name | Unit | Description |
|---|---|---|
| Count | Number of pipeline jobs currently being executed in the location. |
| Count | Number of pipeline tasks currently being executed in the location. |
| Count | Number of online prediction requests per project per base model. |
| Count | Current quota usage for online prediction requests per base model. |
| Count | Current quota limit for online prediction requests per base model. |
| Count | Number of attempts that exceeded the online prediction requests quota. |
| Count | Online prediction input tokens per minute per project per base model. |
| Count | Online prediction output tokens per minute per project per base model. |
| Count | Total tokens served by publisher models in the location. |
| Milliseconds | Latency to first response token from publisher models. |
For the complete list of location metrics, see Google's Vertex AI metrics documentation.
Key metrics — Endpoint
Metric name | Unit | Description |
|---|---|---|
| Percent | Average fraction of time over the past sample period during which the accelerators were actively processing. |
| Bytes | Accelerator memory allocated by the deployed model replica. |
| Percent | CPU utilization of the deployed model replica. |
| Bytes | Memory in use by the deployed model replica. |
| Bytes | Network bytes received by the deployed model replica. |
| Bytes | Network bytes sent by the deployed model replica. |
| Count | Number of online predictions served by the endpoint. |
| Milliseconds | Online prediction latency of the deployed model. |
| Count | Number of online prediction errors returned by the endpoint. |
| Count | Number of online prediction responses returned by the endpoint, faceted by response code. |
| Count | Number of active replicas serving the deployed model. |
| Count | Target number of active replicas for the deployed model. |
For the complete list of endpoint metrics, see Google's Vertex AI metrics documentation.
Key metrics — Index
Metric name | Unit | Description |
|---|---|---|
| Count | Current number of shards backing the index. |
| Count | Number of stream-update requests sent to the index. |
| Count | Number of datapoints successfully upserted or removed via stream update. |
| Milliseconds | Latency between a stream-update response and the update taking effect. |
For the complete list of index metrics, see Google's Vertex AI metrics documentation.
Key metrics — Featurestore
Metric name | Unit | Description |
|---|---|---|
| Percent | Average CPU load across nodes in the Featurestore online storage. |
| Percent | CPU load on the hottest node in the Featurestore online storage. |
| Count | Number of nodes provisioned for the Featurestore online storage. |
| Count | Number of entities updated in the Featurestore online storage. |
| Count | Number of online serving requests handled by the Featurestore, faceted by EntityType. |
| Milliseconds | Online serving request latency at the Featurestore, faceted by EntityType. |
| Bytes | Online serving request size at the Featurestore, faceted by EntityType. |
| Bytes | Online serving response size at the Featurestore, faceted by EntityType. |
| Bytes | Total data stored in the Featurestore. |
| Bytes | Billable bytes processed for Featurestore offline data. |
| Count | Number of streaming-write requests processed into offline storage. |
| Seconds | Time from when the streaming-write API is called until the record reaches offline storage. |
For the complete list of featurestore metrics, see Google's Vertex AI metrics documentation.
Key metrics — Feature Online Store
Metric name | Unit | Description |
|---|---|---|
| Count | Number of online serving requests handled by the Feature Online Store, faceted by FeatureView. |
| Milliseconds | Online serving request latency at the Feature Online Store, faceted by FeatureView. |
| Bytes | Online serving response size at the Feature Online Store, faceted by FeatureView. |
| Count | Number of syncs currently running at the Feature Online Store. |
| Seconds | Age of the data being served by the Feature Online Store. |
| Count | Breakdown of records in the Feature Online Store by synced timestamp. |
| Percent | Average CPU load across Bigtable nodes backing the Feature Online Store. |
| Percent | CPU load on the hottest Bigtable node backing the Feature Online Store. |
| Count | Number of Bigtable nodes backing the Feature Online Store. |
| Bytes | Total data stored in the Feature Online Store. |
For the complete list of feature online store metrics, see Google's Vertex AI metrics documentation.
Key metrics — Pipeline Job
Metric name | Unit | Description |
|---|---|---|
| Seconds | Runtime seconds of the pipeline job being executed, from creation to end. |
| Count | Total number of completed pipeline tasks in the pipeline job. |
For the complete list of pipeline job metrics, see Google's Vertex AI metrics documentation.
Service account or user account
Find and use data
After activating the integration and waiting a few minutes (based on the polling frequency), data will appear in the New Relic UI. To find and use your data, including links to your and alert settings, go to one.newrelic.com > All capabilities > Infrastructure > GCP > (select an integration).
Data is attached to the following event types:
Entity | Event Type | Provider |
|---|---|---|
Endpoint |
|
|
Feature store |
|
|
Feature Online Store |
|
|
Location |
|
|
Index |
|
|
PipelineJob |
|
|
For more on how to use your data, see Understand and use integration data.
VertexAI Endpoint data
Metric | Unit | Description |
|---|---|---|
| Percent | Average fraction of time over the past sample period during which the accelerator(s) were actively processing. |
| Bytes | Amount of accelerator memory allocated by the deployed model replica. |
| Count | Number of online prediction errors. |
| Bytes | Amount of memory allocated by the deployed model replica and currently in use. |
| Bytes | Number of bytes received over the network by the deployed model replica. |
| Bytes | Number of bytes sent over the network by the deployed model replica. |
| Count | Number of online predictions. |
| Milliseconds | Online prediction latency of the deployed model. |
| Milliseconds | Online prediction latency of the private deployed model. |
| Count | Number of active replicas used by the deployed model. |
| Count | Number of different online prediction response codes. |
| Count | Target number of active replicas needed for the deployed model. |
VertexAI Featurestore data
Metric | Unit | Description |
|---|---|---|
| Percent | The average CPU load for a node in the Featurestore online storage. |
| Percent | The CPU load for the hottest node in the Featurestore online storage. |
| Count | The number of nodes for the Featurestore online storage. |
| Count | Number of entities updated on the Featurestore online storage. |
| Milliseconds | Online serving latencies by EntityType. |
| Bytes | Request size by EntityType. |
| Count | Featurestore online serving count by EntityType. |
| Bytes | Response size by EntityType. |
| Bytes | Number of bytes billed for offline data processed. |
| Bytes | Bytes stored in Featurestore. |
| Count | Number of streaming write requests processed for offline storage. |
| Seconds | Time (in second) since the write API is called until it is written to offline storage. |
VertexAI FeatureOnlineStore data
Metric | Unit | Description |
|---|---|---|
| Count | Number of serving count by FeatureView. |
| Bytes | Serving response size by FeatureView. |
| Milliseconds | Online serving latencies by FeatureView. |
| Milliseconds | Number of running syncs at given point of time. |
| Seconds | Measure of the serving data age in seconds. |
| Count | Breakdown of data in Feature Online Store by synced timestamp. |
| Percent | The average CPU load of nodes in the Feature Online Store. |
| Percent | The CPU load of the hottest node in the Feature Online Store. |
| Count | The number of nodes for the Feature Online Store(Bigtable). |
| Count | Bytes stored in the Feature Online Store. |
VertexAI Location data
Metric | Unit | Description |
|---|---|---|
| Count | Number of requests per base model. |
| Count | Number of attempts to exceed the limit on quota metric. |
| Count | Current limit on quota metric. |
| Count | Current usage on quota metric. |
| Count | Number of pipeline jobs being executed. |
| Count | Number of pipeline tasks being executed. |
VertexAI Index data
Metric | Unit | Description |
|---|---|---|
| Count | Number of successfully upserted or removed datapoints. |
| Milliseconds | The latencies between the user receives a UpsertDatapointsResponse or RemoveDatapointsResponse and that update takes effect. |
| Count | Number of stream update requests. |
VertexAI Pipeline Job data
Metric | Unit | Description |
|---|---|---|
| Seconds | Runtime seconds of the pipeline job being executed (from creation to end). |
| Count | Total number of completed Pipeline Tasks. |