📡 Monitor or Automate Kubernetes with GermainUX
GermainUX provides real-time monitoring, analytics, alerting, and automation for Kubernetes environments.
Monitor the health, availability, and performance of Kubernetes infrastructure—from clusters and nodes to pods and containers—and quickly identify conditions affecting the applications and services running within them.
What GermainUX can help teams detect:
|
Issue |
|
|---|---|
|
1 |
Kubernetes component availability issues |
|
2 |
Unhealthy or unavailable pods |
|
3 |
Container failures |
|
4 |
Excessive container restarts |
|
5 |
High CPU or memory utilization |
|
6 |
Node resource constraints |
|
7 |
Container readiness and liveness issues |
|
8 |
Kubernetes infrastructure instability |
Collected Metrics
The Kubernetes monitor collects the KPIs below. Which KPIs are collected depends on the monitor's Collect cluster status and Collect cluster metrics settings.
In the Name column, <dataSource> is the monitor's data source name. "Configured namespace" means the monitor's Namespace setting (default: default).
Namespace scope: KPIs for pods, containers, workloads (deployments, stateful sets, daemon sets) and unschedulable pods only cover the configured namespace. Node-level and cluster-level KPIs cover all namespaces.
Cluster status
Collected only when Collect cluster status is enabled.
|
KPI |
Name format |
Unit |
Notes |
|---|---|---|---|
|
Kubernetes Node Status |
|
% |
One per node. 100 = ready, 50 = ready but under pressure or cordoned, 0 = not ready. Details list the problems (e.g. DiskPressure, Cordoned). Skipped with a warning if nodes can't be listed. |
|
Kubernetes Pod Status |
|
% |
One per pod in the configured namespace. 100 = Running, 0 = any other phase. |
|
Kubernetes Container Status |
|
% |
One per container in the configured namespace. 100 = ready, 50 = started but not ready, 0 = not running. Details give the waiting or termination reason (e.g. CrashLoopBackOff, OOMKilled). Pods with no containers created yet are skipped. |
|
Kubernetes Container Restarts |
|
restarts |
One per container in the configured namespace. Restarts since the previous collection. Not reported until a container has been seen once, so there is no value on the first collection or for a new pod. |
|
Kubernetes Container Restart Count |
|
restarts |
One per cluster. Sum of all Container Restarts values. Not reported until at least one container has a baseline. |
|
Kubernetes Unschedulable Pod Count |
|
pods |
One per cluster. Pods in the configured namespace that the scheduler can't place on any node. |
|
Kubernetes Deployment Availability |
|
% |
One per deployment in the configured namespace. Available ÷ desired replicas; a deployment scaled to zero = 100. Details e.g. "2 of 3 available". Skipped with a warning if the monitor lacks permission. |
|
Kubernetes StatefulSet Availability |
|
% |
One per stateful set in the configured namespace. Same calculation and permission handling as deployments. |
|
Kubernetes DaemonSet Availability |
|
% |
One per daemon set in the configured namespace. Available ÷ desired scheduled pods. Same permission handling as deployments. |
Cluster metrics: nodes
Collected only when Collect cluster metrics is enabled. Usage values come from the Kubernetes Metrics API.
|
KPI |
Name format |
Unit |
Notes |
|---|---|---|---|
|
Kubernetes Node CPU Usage |
|
% |
One per node reported by the Metrics API. Percentage of the node's allocatable CPU. |
|
Kubernetes Node CPU Cores |
|
cores |
One per node reported by the Metrics API. Raw usage. |
|
Kubernetes Node Memory Usage |
|
% |
One per node reported by the Metrics API. Percentage of the node's allocatable memory. |
|
Kubernetes Node Memory Bytes |
|
bytes |
One per node reported by the Metrics API. Raw usage. |
|
Kubernetes Node Pod Usage |
|
% |
One per node. Non-terminated pods (all namespaces) as a percentage of the node's max pods. |
|
Kubernetes Node Pods |
|
pods |
One per node. Non-terminated pods, all namespaces; 0 for empty nodes. |
|
Kubernetes Node CPU Request Commitment |
|
% |
One per node. Sum of container CPU requests (all namespaces) as a percentage of allocatable CPU. |
|
Kubernetes Node CPU Requests |
|
cores |
One per node. Sum of container CPU requests, all namespaces. |
|
Kubernetes Node Memory Request Commitment |
|
% |
One per node. Sum of container memory requests (all namespaces) as a percentage of allocatable memory. |
|
Kubernetes Node Memory Requests |
|
bytes |
One per node. Sum of container memory requests, all namespaces. |
Cluster metrics: cluster
Collected only when Collect cluster metrics is enabled.
|
KPI |
Name format |
Unit |
Notes |
|---|---|---|---|
|
Kubernetes Cluster Pod Usage |
|
% |
One per cluster. Total non-terminated pods as a percentage of total max pods across all nodes. |
|
Kubernetes Cluster Pods |
|
pods |
One per cluster. Total non-terminated pods across all nodes. |
Cluster metrics: containers
Collected only when Collect cluster metrics is enabled.
|
KPI |
Name format |
Unit |
Notes |
|---|---|---|---|
|
Kubernetes Container CPU Usage |
|
% |
One per container in the configured namespace. Percentage of the container's CPU limit, or of its node's allocatable CPU if no limit is set. Can briefly exceed 100 when a limit is set. |
|
Kubernetes Container CPU Cores |
|
cores |
One per container in the configured namespace. Raw usage. |
|
Kubernetes Container CPU Request |
|
cores |
One per container in the configured namespace that has a CPU request. |
|
Kubernetes Container CPU Limit |
|
cores |
One per container in the configured namespace that has a CPU limit. Reflects the value in effect after an in-place resize. |
|
Kubernetes Container Memory Usage |
|
% |
One per container in the configured namespace. Percentage of the container's memory limit, or of its node's allocatable memory if no limit is set. |
|
Kubernetes Container Memory Bytes |
|
bytes |
One per container in the configured namespace. Raw usage. |
|
Kubernetes Container Memory Request |
|
bytes |
One per container in the configured namespace that has a memory request. |
|
Kubernetes Container Memory Limit |
|
bytes |
One per container in the configured namespace that has a memory limit. Reflects the value in effect after an in-place resize. |
⏰ Availability & Health Monitoring
Continuously monitor the operational health of your Kubernetes environment.
GermainUX can monitor Kubernetes components such as:
|
Component |
|---|
|
API Server |
|
etcd |
|
Scheduler |
|
Controller Manager |
It can also monitor the status and health of:
|
Resource |
|---|
|
Nodes |
|
Pods |
|
Containers |
|
Kubernetes servers |
This provides teams with real-time visibility into infrastructure conditions that may affect application availability or performance.
📦 Container Monitoring
Monitor the status and resource utilization of individual containers running within Kubernetes pods.
GermainUX can provide visibility into:
🔎 Container Status
Monitor container status, including readiness and liveness, to quickly identify containers that are unavailable or unhealthy.
🔁 Container Restarts
Track container restart counts to identify instability, recurring failures, or abnormal behavior.
⚙️ Container CPU
Monitor CPU utilization to detect excessive resource consumption, capacity constraints, and potential performance bottlenecks.
💧 Container Memory
Monitor memory utilization to identify excessive consumption and potential memory-related performance issues.
📦 Pod Monitoring
Monitor the health and status of pods throughout the Kubernetes environment.
GermainUX can monitor information such as:
|
Metric |
|---|
|
Pod status |
|
Readiness |
|
Phase |
|
Conditions |
|
Associated container health |
This helps teams identify pods that are unhealthy, unavailable, or contributing to application performance issues.
💻 Node Monitoring
Monitor Kubernetes nodes to understand infrastructure health and resource utilization.
GermainUX can track:
🧮 Node CPU
Monitor CPU utilization across Kubernetes nodes to identify resource constraints and unusual consumption.
🧠 Node Memory
Monitor node memory utilization to identify capacity issues and memory pressure that may affect workloads.
Combining node, pod, and container monitoring provides visibility from the infrastructure layer down to individual workloads.
⚙️ Kubernetes Configuration & Control Monitoring
GermainUX can monitor additional Kubernetes information that provides context around the state and behavior of the environment.
This includes:
🚩 Control Flags
Monitor Kubernetes control information associated with operations such as:
|
Operation |
|---|
|
Rolling updates |
|
Scaling |
|
Autoscaling |
💡 Display Hints
Collect contextual information exposed by Kubernetes for containers, pods, and nodes to provide additional insight during analysis and troubleshooting.
🔍 Detect Kubernetes Issues in Real Time
GermainUX can continuously analyze Kubernetes telemetry to identify conditions that require attention.
For example:
|
Example |
Flow |
|---|---|
|
Container Failure |
Detect → Alert Team → Create Ticket |
|
Repeated Container Restarts |
Detect Threshold → Alert → Investigate |
|
High Node Memory |
Detect → Identify Affected Workloads → Alert |
|
Unhealthy Pod |
Detect → Analyze → Trigger Action |
This helps teams move from manually reviewing infrastructure metrics to proactively identifying the conditions that require action.
📊 Analytics, Dashboards & Reports
Kubernetes monitoring data can be analyzed through GermainUX real-time dashboards and automated reports.
Examples include:
|
Report / Metric |
|---|
|
Kubernetes component availability |
|
Container health |
|
Container restart frequency |
|
Container CPU utilization |
|
Container memory utilization |
|
Pod health and availability |
|
Node CPU utilization |
|
Node memory utilization |
|
Infrastructure health and performance |
Dashboards, KPIs, and reports are customizable to match the architecture and operational requirements of each Kubernetes environment.
🚨 Alerts & SLA Monitoring
Configure alerts and SLA thresholds for Kubernetes conditions such as:
|
Condition |
|---|
|
Kubernetes component unavailable |
|
Container unavailable or unhealthy |
|
Container restart count above threshold |
|
Container CPU above threshold |
|
Container memory above threshold |
|
Pod unavailable or unhealthy |
|
Node CPU above threshold |
|
Node memory above threshold |
|
Kubernetes server health issues |
Alert conditions, thresholds, time windows, recipients, and escalation logic can be customized according to operational requirements.
🤖 Automation
GermainUX can trigger automated actions when specific Kubernetes conditions are detected.
For example:
Detect → Analyze → Alert → Create Ticket → Trigger Action
Depending on the use case, automation can include:
|
Automation |
|---|
|
Alerts and notifications |
|
Ticket or incident creation |
|
Script execution |
|
Program execution |
|
HTTP execution |
|
SSH execution |
|
Scheduled actions |
|
Custom operational workflows |
|
Remediation actions where appropriate |
This allows Kubernetes monitoring to become part of a broader operational workflow rather than simply generating additional alerts.
🔗 Correlate Kubernetes with Applications & Services
Kubernetes infrastructure does not operate in isolation.
GermainUX can combine Kubernetes monitoring with telemetry from applications, APIs, databases, Java services, operating systems, and other technologies monitored by GermainUX.
This provides broader visibility across the technology stack:
Application → Service → Container → Pod → Node → Infrastructure
Teams can use this context to determine whether an application performance or availability issue originates in the application itself or in the Kubernetes infrastructure supporting it.
🔧 Configuration
Kubernetes monitoring is performed through the GermainUX Engine.
Deploy and configure a GermainUX Engine with access to the Kubernetes environment, then select the components, KPIs, thresholds, alerts, and automation appropriate for your environment.
See Deploy Monitoring for Kubernetes for detailed deployment and configuration instructions.
GermainUX itself can also be deployed within a Kubernetes container.
📦 Additional Resources
-
KPIs for Kubernetes
-
AI Insights & Recommendations
-
More Analytics Features
-
More Automation Features
ℹ️ Get Help
The Germain Team can help you set this up. Contact GermainUX Support.
Component: Engine
Feature Availability: 2022.3 or later