Kubernetes KPIs

📊 Kubernetes Metrics Overview: Pod, Container, and Node Statuses

🧭 Overview

GermainUX collects Kubernetes KPIs to monitor the health, availability, resource utilization, and stability of clusters, nodes, pods, and containers.

These KPIs can be used for real-time monitoring, dashboards, SLA tracking, alerting, analysis, and automation.

They also provide supporting technology evidence when investigating user-experience, application, or business-process issues in the Analysis Dashboard.

📋 Kubernetes KPIs

KPI

Name

Example

Notes

Kubernetes:PodStatus

Germain UX DS name : metadata.name of the POD

d710a3b4:lyjvms-sts-3

1 metric per POD

Kubernetes:ContainerStatus

Germain UX DS name : metadata.name of the POD : Container Name

d710a3b4:lyjvms-sts-2:apache-pocjvms

1 metric per Container

Kubernetes:ContainerRestartCount

Germain UX DS name

d710a3b4

1 metric per Cluster

Kubernetes:Node CPU Usage

metadata.name of the Node

10.230.92.57

1 metric per Node

Kubernetes:Node Memory Usage

metadata.name of the Node

10.230.92.57

1 metric per Node

Kubernetes:Container CPU Usage

metadata.name of the container

-

1 metric per container

Kubernetes:Container Memory Usage

metadata.name of the container

-

1 metric per container

🌱 Pod Status

Kubernetes monitors the status of each Kubernetes Pod.

One metric is generated for each Pod, allowing GermainUX to identify Pods that are not operating in their expected state.

The Data Source name identifies both the GermainUX Kubernetes Data Source and the Pod:

<GermainUX Data Source>:<Pod Name>

Example:

d710a3b4:lyjvms-sts-3

Use this KPI to:

Use

Monitor Pod health

Detect unavailable or unhealthy Pods

Identify affected applications or services

Track Pod status over time

Trigger alerts or automated actions when configured conditions are detected

📦 Container Status

Kubernetes monitors the status of individual containers running within Kubernetes Pods.

One metric is generated for each Container.

The Data Source name identifies the Kubernetes Data Source, Pod, and Container:

<GermainUX Data Source>:<Pod Name>:<Container Name>

Example:

d710a3b4:lyjvms-sts-2:apache-pocjvms

Use this KPI to identify containers that are unavailable or not operating in their expected state.

🔁 Container Restart Count

Kubernetes monitors container restarts within the Kubernetes cluster.

The Data Source name identifies the monitored Kubernetes cluster.

Example:

d710a3b4

Use restart count to identify container instability and investigate whether repeated restarts are associated with application availability, errors, or performance issues.

🖥️ Node CPU Usage

Kubernetes CPU Usage monitors CPU utilization for each Kubernetes Node.

One metric is generated for each Node.

The Node is identified using its metadata.name.

Example:

10.230.92.57

Use this KPI to identify sustained or abnormal CPU utilization that may affect Pods, Containers, services, or applications running on the Node.

💾 Node Memory Usage

Kubernetes Memory Usage monitors memory utilization for each Kubernetes Node.

One metric is generated for each Node and identified using its metadata.name.

Use this KPI to identify memory pressure or abnormal memory utilization that may affect workloads running on the Node.

🚀 Container CPU Usage

Kubernetes CPU Usage monitors CPU utilization at the individual Container level.

One metric is generated for each Container and identified using the Container's metadata.name.

Use this KPI to:

Use

Identify CPU-intensive Containers

Detect abnormal CPU utilization

Compare resource consumption across Containers

Investigate application or service performance issues

🧠 Container Memory Usage

Kubernetes Memory Usage monitors memory utilization at the individual Container level.

One metric is generated for each Container and identified using the Container's metadata.name.

Use this KPI to:

Use

Identify memory-intensive Containers

Detect abnormal memory utilization

Identify potential resource constraints

Compare memory consumption across Containers

Investigate application or service degradation

🔍 Analyze Kubernetes Health

Kubernetes KPIs can be monitored in GermainUX dashboards and analyzed alongside application, service, user-experience, and business KPIs.

For example:

Application Performance Degradation
→ Technology Health
→ Kubernetes Node CPU Usage
→ Affected Node
→ Pods & Containers
→ Related KPIs
→ Analysis
→ Likely Root Cause

Or:

Service Failure
→ Kubernetes
→ Affected Container
→ Pod
→ Container Restart Count
→ Related Application Errors
→ Analysis

💓 Technology Health

Kubernetes KPIs can be surfaced in the Technology Health view of the Operational Dashboard to provide real-time visibility into the infrastructure supporting monitored applications.

This can help identify whether an application or user issue is associated with:

Area

Cluster health

Pod availability

Container availability

Container restarts

Node CPU utilization

Node memory utilization

Container CPU utilization

Container memory utilization

🔔 Alerts & Automation

Kubernetes KPIs can be associated with SLAs, alerts, Rules, and Automation to detect and respond to abnormal conditions.

For example:

Container Status Changes
→ Detect condition
→ Analyze affected application/service
→ Alert appropriate team
→ Create or update ticket
→ Continue monitoring
→ Validate recovery

The available actions depend on the GermainUX configuration and permissions.

Document

Kubernetes Monitoring

Technology Health Dashboard

KPIs

SLAs

Analysis Dashboard

Root-Cause Analysis

Monitoring

Automation

📦 Additional Resources

ℹ️ Get Help

The Germain Team can help you set this up. Contact GermainUX Support.

Component: Engine

Feature Availability: 2022.3 or later