Kubernetes Observability

📡 Monitor or Automate Kubernetes with GermainUX

GermainUX provides real-time monitoring, analytics, alerting, and automation for Kubernetes environments.

Monitor the health, availability, and performance of Kubernetes infrastructure—from clusters and nodes to pods and containers—and quickly identify conditions affecting the applications and services running within them.

What GermainUX can help teams detect:

Issue

1

Kubernetes component availability issues

2

Unhealthy or unavailable pods

3

Container failures

4

Excessive container restarts

5

High CPU or memory utilization

6

Node resource constraints

7

Container readiness and liveness issues

8

Kubernetes infrastructure instability

Collected Metrics

The Kubernetes monitor collects the KPIs below. Which KPIs are collected depends on the monitor's Collect cluster status and Collect cluster metrics settings.

In the Name column, <dataSource> is the monitor's data source name. "Configured namespace" means the monitor's Namespace setting (default: default).

Namespace scope: KPIs for pods, containers, workloads (deployments, stateful sets, daemon sets) and unschedulable pods only cover the configured namespace. Node-level and cluster-level KPIs cover all namespaces.

Cluster status

Collected only when Collect cluster status is enabled.

KPI

Name format

Unit

Notes

Kubernetes Node Status

<dataSource>:<node>

%

One per node. 100 = ready, 50 = ready but under pressure or cordoned, 0 = not ready. Details list the problems (e.g. DiskPressure, Cordoned). Skipped with a warning if nodes can't be listed.

Kubernetes Pod Status

<dataSource>:<pod>

%

One per pod in the configured namespace. 100 = Running, 0 = any other phase.

Kubernetes Container Status

<dataSource>:<pod>:<container>

%

One per container in the configured namespace. 100 = ready, 50 = started but not ready, 0 = not running. Details give the waiting or termination reason (e.g. CrashLoopBackOff, OOMKilled). Pods with no containers created yet are skipped.

Kubernetes Container Restarts

<dataSource>:<pod>:<container>

restarts

One per container in the configured namespace. Restarts since the previous collection. Not reported until a container has been seen once, so there is no value on the first collection or for a new pod.

Kubernetes Container Restart Count

<dataSource> (cluster)

restarts

One per cluster. Sum of all Container Restarts values. Not reported until at least one container has a baseline.

Kubernetes Unschedulable Pod Count

<dataSource> (cluster)

pods

One per cluster. Pods in the configured namespace that the scheduler can't place on any node.

Kubernetes Deployment Availability

<dataSource>:<deployment>

%

One per deployment in the configured namespace. Available ÷ desired replicas; a deployment scaled to zero = 100. Details e.g. "2 of 3 available". Skipped with a warning if the monitor lacks permission.

Kubernetes StatefulSet Availability

<dataSource>:<statefulset>

%

One per stateful set in the configured namespace. Same calculation and permission handling as deployments.

Kubernetes DaemonSet Availability

<dataSource>:<daemonset>

%

One per daemon set in the configured namespace. Available ÷ desired scheduled pods. Same permission handling as deployments.

Cluster metrics: nodes

Collected only when Collect cluster metrics is enabled. Usage values come from the Kubernetes Metrics API.

KPI

Name format

Unit

Notes

Kubernetes Node CPU Usage

<dataSource>:<node>

%

One per node reported by the Metrics API. Percentage of the node's allocatable CPU.

Kubernetes Node CPU Cores

<dataSource>:<node>

cores

One per node reported by the Metrics API. Raw usage.

Kubernetes Node Memory Usage

<dataSource>:<node>

%

One per node reported by the Metrics API. Percentage of the node's allocatable memory.

Kubernetes Node Memory Bytes

<dataSource>:<node>

bytes

One per node reported by the Metrics API. Raw usage.

Kubernetes Node Pod Usage

<dataSource>:<node>

%

One per node. Non-terminated pods (all namespaces) as a percentage of the node's max pods.

Kubernetes Node Pods

<dataSource>:<node>

pods

One per node. Non-terminated pods, all namespaces; 0 for empty nodes.

Kubernetes Node CPU Request Commitment

<dataSource>:<node>

%

One per node. Sum of container CPU requests (all namespaces) as a percentage of allocatable CPU.

Kubernetes Node CPU Requests

<dataSource>:<node>

cores

One per node. Sum of container CPU requests, all namespaces.

Kubernetes Node Memory Request Commitment

<dataSource>:<node>

%

One per node. Sum of container memory requests (all namespaces) as a percentage of allocatable memory.

Kubernetes Node Memory Requests

<dataSource>:<node>

bytes

One per node. Sum of container memory requests, all namespaces.

Cluster metrics: cluster

Collected only when Collect cluster metrics is enabled.

KPI

Name format

Unit

Notes

Kubernetes Cluster Pod Usage

<dataSource> (cluster)

%

One per cluster. Total non-terminated pods as a percentage of total max pods across all nodes.

Kubernetes Cluster Pods

<dataSource> (cluster)

pods

One per cluster. Total non-terminated pods across all nodes.

Cluster metrics: containers

Collected only when Collect cluster metrics is enabled.

KPI

Name format

Unit

Notes

Kubernetes Container CPU Usage

<dataSource>:<pod>:<container>

%

One per container in the configured namespace. Percentage of the container's CPU limit, or of its node's allocatable CPU if no limit is set. Can briefly exceed 100 when a limit is set.

Kubernetes Container CPU Cores

<dataSource>:<pod>:<container>

cores

One per container in the configured namespace. Raw usage.

Kubernetes Container CPU Request

<dataSource>:<pod>:<container>

cores

One per container in the configured namespace that has a CPU request.

Kubernetes Container CPU Limit

<dataSource>:<pod>:<container>

cores

One per container in the configured namespace that has a CPU limit. Reflects the value in effect after an in-place resize.

Kubernetes Container Memory Usage

<dataSource>:<pod>:<container>

%

One per container in the configured namespace. Percentage of the container's memory limit, or of its node's allocatable memory if no limit is set.

Kubernetes Container Memory Bytes

<dataSource>:<pod>:<container>

bytes

One per container in the configured namespace. Raw usage.

Kubernetes Container Memory Request

<dataSource>:<pod>:<container>

bytes

One per container in the configured namespace that has a memory request.

Kubernetes Container Memory Limit

<dataSource>:<pod>:<container>

bytes

One per container in the configured namespace that has a memory limit. Reflects the value in effect after an in-place resize.


⏰ Availability & Health Monitoring

Continuously monitor the operational health of your Kubernetes environment.

GermainUX can monitor Kubernetes components such as:

Component

API Server

etcd

Scheduler

Controller Manager

It can also monitor the status and health of:

Resource

Nodes

Pods

Containers

Kubernetes servers

This provides teams with real-time visibility into infrastructure conditions that may affect application availability or performance.


📦 Container Monitoring

Monitor the status and resource utilization of individual containers running within Kubernetes pods.

GermainUX can provide visibility into:

🔎 Container Status

Monitor container status, including readiness and liveness, to quickly identify containers that are unavailable or unhealthy.

🔁 Container Restarts

Track container restart counts to identify instability, recurring failures, or abnormal behavior.

⚙️ Container CPU

Monitor CPU utilization to detect excessive resource consumption, capacity constraints, and potential performance bottlenecks.

💧 Container Memory

Monitor memory utilization to identify excessive consumption and potential memory-related performance issues.


📦 Pod Monitoring

Monitor the health and status of pods throughout the Kubernetes environment.

GermainUX can monitor information such as:

Metric

Pod status

Readiness

Phase

Conditions

Associated container health

This helps teams identify pods that are unhealthy, unavailable, or contributing to application performance issues.


💻 Node Monitoring

Monitor Kubernetes nodes to understand infrastructure health and resource utilization.

GermainUX can track:

🧮 Node CPU

Monitor CPU utilization across Kubernetes nodes to identify resource constraints and unusual consumption.

🧠 Node Memory

Monitor node memory utilization to identify capacity issues and memory pressure that may affect workloads.

Combining node, pod, and container monitoring provides visibility from the infrastructure layer down to individual workloads.


⚙️ Kubernetes Configuration & Control Monitoring

GermainUX can monitor additional Kubernetes information that provides context around the state and behavior of the environment.

This includes:

🚩 Control Flags

Monitor Kubernetes control information associated with operations such as:

Operation

Rolling updates

Scaling

Autoscaling

💡 Display Hints

Collect contextual information exposed by Kubernetes for containers, pods, and nodes to provide additional insight during analysis and troubleshooting.


🔍 Detect Kubernetes Issues in Real Time

GermainUX can continuously analyze Kubernetes telemetry to identify conditions that require attention.

For example:

Example

Flow

Container Failure

Detect → Alert Team → Create Ticket

Repeated Container Restarts

Detect Threshold → Alert → Investigate

High Node Memory

Detect → Identify Affected Workloads → Alert

Unhealthy Pod

Detect → Analyze → Trigger Action

This helps teams move from manually reviewing infrastructure metrics to proactively identifying the conditions that require action.


📊 Analytics, Dashboards & Reports

Kubernetes monitoring data can be analyzed through GermainUX real-time dashboards and automated reports.

Examples include:

Report / Metric

Kubernetes component availability

Container health

Container restart frequency

Container CPU utilization

Container memory utilization

Pod health and availability

Node CPU utilization

Node memory utilization

Infrastructure health and performance

Dashboards, KPIs, and reports are customizable to match the architecture and operational requirements of each Kubernetes environment.


🚨 Alerts & SLA Monitoring

Configure alerts and SLA thresholds for Kubernetes conditions such as:

Condition

Kubernetes component unavailable

Container unavailable or unhealthy

Container restart count above threshold

Container CPU above threshold

Container memory above threshold

Pod unavailable or unhealthy

Node CPU above threshold

Node memory above threshold

Kubernetes server health issues

Alert conditions, thresholds, time windows, recipients, and escalation logic can be customized according to operational requirements.


🤖 Automation

GermainUX can trigger automated actions when specific Kubernetes conditions are detected.

For example:

Detect → Analyze → Alert → Create Ticket → Trigger Action

Depending on the use case, automation can include:

Automation

Alerts and notifications

Ticket or incident creation

Script execution

Program execution

HTTP execution

SSH execution

Scheduled actions

Custom operational workflows

Remediation actions where appropriate

This allows Kubernetes monitoring to become part of a broader operational workflow rather than simply generating additional alerts.


🔗 Correlate Kubernetes with Applications & Services

Kubernetes infrastructure does not operate in isolation.

GermainUX can combine Kubernetes monitoring with telemetry from applications, APIs, databases, Java services, operating systems, and other technologies monitored by GermainUX.

This provides broader visibility across the technology stack:

Application → Service → Container → Pod → Node → Infrastructure

Teams can use this context to determine whether an application performance or availability issue originates in the application itself or in the Kubernetes infrastructure supporting it.


🔧 Configuration

Kubernetes monitoring is performed through the GermainUX Engine.

Deploy and configure a GermainUX Engine with access to the Kubernetes environment, then select the components, KPIs, thresholds, alerts, and automation appropriate for your environment.

See Deploy Monitoring for Kubernetes for detailed deployment and configuration instructions.

GermainUX itself can also be deployed within a Kubernetes container.


📦 Additional Resources


ℹ️ Get Help

The Germain Team can help you set this up. Contact GermainUX Support.

Component: Engine

Feature Availability: 2022.3 or later