Understanding Kubernetes: Debugging broken workloads

Understanding Kubernetes: Debugging broken workloads

In the last article, we looked at monitoring and how metrics, logs and alerts can help us understand what is happening in our cluster. And sooner or later, our monitoring will tell us: Something is broken. Now the fun begins. 😅 So let’s have a look at some of the tools Kubernetes gives us for debugging broken workloads.

Start with the Pod

When something doesn’t look right, we can start with investigating the pod:

kubectl get pods

This can already give us first information and useful hints:

NAME                         READY   STATUS             RESTARTS   AGE
my-app-7f8c6d9f7b-x2k9m      1/1     Running            0          5m
my-app-7f8c6d9f7b-p8j4q      0/1     CrashLoopBackOff   5          3m
my-app-7f8c6d9f7b-q7v2n      0/1     Pending            0          2m

Pending could mean that Kubernetes hasn’t found a suitable node yet. CrashLoopBackOff tells us that the container is repeatedly crashing. And READY 0/1 tells us that the container in those Pods isn’t currently considered ready.

But we still do not know what exactly is happening - so we’ll need to take a closer look.

One of your best friends: kubectl describe

In order to get more detailed information about the Pod, we can describe the Pod itself by running:

kubectl describe pod my-app-7f8c6d9f7b-p8j4q

This gives us a lot more information about the Pod, including its containers, mounts, probes, resource configuration and current state. The bottom of the output is particularly interesting:

Events:
  Type     Reason     Age                From     Message
  ----     ------     ---                ----     -------
  Normal   Scheduled  2m                 scheduler Successfully assigned default/my-app-... to node-2
  Normal   Pulled     2m                 kubelet   Successfully pulled image "my-app:v1"
  Warning  BackOff    30s (x5 over 2m)   kubelet   Back-off restarting failed container

Now we know that Kubernetes successfully scheduled the Pod and pulled the image. The problem happened afterwards: the container keeps crashing and Kubernetes is trying to restart it.

This is why describe is so useful: it gives us information about what Kubernetes knows about the object and what happened to it.

Checking events

Events are often one of the quickest ways to find out what went wrong. You can look at the events for a specific Pod with kubectl describe, but you can also query events directly:

kubectl get events --sort-by=.lastTimestamp

For example, you might find something like:

Warning  FailedScheduling
0/3 nodes are available: 3 Insufficient memory.

Or:

Warning  FailedMount
MountVolume.SetUp failed for volume "config"

Or:

Warning  Unhealthy
Readiness probe failed: HTTP probe failed with statuscode: 503

These messages can be surprisingly helpful. They can also tell you when the problem isn’t actually inside your application at all. Maybe the Pod can’t be scheduled because the cluster doesn’t have enough resources. Maybe a volume can’t be mounted. Maybe an image can’t be pulled.

Checking logs

If the container actually started, the next obvious question is: What does the application say?

kubectl logs my-app-7f8c6d9f7b-p8j4q

## Or, to follow logs:
kubectl logs -f my-app-7f8c6d9f7b-p8j4q

One particularly useful option when dealing with restarting containers is:

kubectl logs my-app-7f8c6d9f7b-p8j4q --previous

This shows the logs from the previous instance of the container - and is extremely useful for CrashLoopBackOff situations. The current container might have only been running for a second, while the previous instance may have printed the actual error before it crashed.

For example:

Starting application...
Connecting to database...
ERROR: password authentication failed for user "app"

Get inside the container

Sometimes the logs aren’t enough. Maybe the application looks fine, but something about the container environment isn’t what you expected. That’s where kubectl exec comes in.

kubectl exec -it my-app-7f8c6d9f7b-p8j4q -- /bin/sh

Now we’re inside the running container and can inspect things directly.

For example:

env

or:

ls -la /etc/config

or perhaps test whether another service is reachable:

curl http://my-database:8080

This can help answer questions like:

  • Is the configuration actually there?
  • Are the expected environment variables set?
  • Can the container resolve another service?
  • Can it connect to the endpoint I’m expecting?

There is one catch, though: exec only works if the container is running and contains a shell (or another executable you can use). And that’s not always the case.

So what if the container has no debugging tools?

This is where ephemeral containers can help. Instead of changing the application image just to add debugging tools, Kubernetes allows us to temporarily add another container to an existing Pod.

For example:

kubectl debug my-app-7f8c6d9f7b-p8j4q -it \
  --image=busybox \
  --target=my-app

The exact debugging setup depends on the container runtime and what we want to inspect, but the basic idea is simple: Keep the application container as it is, and add a temporary container with the tools we need. This can be especially useful for minimal or distroless images.

For example, if the application container doesn’t contain sh or curl, a debugging container can give us a shell and some basic networking tools without having to rebuild and redeploy the application image.

Ephemeral containers are not meant to become part of the normal workload. They’re there to help us investigate what is happening.

Don’t forget about probes

Sometimes a Pod is running, but Kubernetes still doesn’t consider it healthy. That’s where startup, readiness and liveness probes become important. You can see the configured probes with:

kubectl describe pod my-app-7f8c6d9f7b-p8j4q

And you might find an event such as:

Warning  Unhealthy
Readiness probe failed: HTTP probe failed with statuscode: 503

This tells us something important: The container itself might be running perfectly well, but its readiness check is failing. That means Kubernetes can keep the Pod running while the Pod is not considered ready to receive traffic.

A liveness probe failing has a different consequence: Kubernetes may restart the container.

So when a workload appears to be running but isn’t behaving as expected, check the probes. They are part of the application’s interaction with Kubernetes and can be the reason why a perfectly running container isn’t actually serving traffic.

A simple debugging checklist

At this point, we have seen quite a few commands - time for some minimal checklist to get started:

1. What is the Pod doing?

kubectl get pods

Look for:

  • Pending
  • CrashLoopBackOff
  • ImagePullBackOff
  • Error
  • unexpected restarts
  • READY not matching the expected number

2. What does Kubernetes know about it?

kubectl describe pod <pod>

Pay particular attention to the container state and the Events at the bottom.

3. What happened?

kubectl get events --sort-by=.lastTimestamp

Look for scheduling problems, failed mounts, image pull errors, probe failures and other warnings.

4. What does the application say?

kubectl logs <pod>
kubectl logs <pod> --previous

5. Can I inspect the container?

kubectl exec -it <pod> -- /bin/sh

Check configuration, files, environment variables and connectivity.

If the container is too minimal or doesn’t contain the tools you need, consider an ephemeral debugging container:

kubectl debug <pod> -it --image=busybox

6. Is Kubernetes actually considering the Pod healthy?

Check the startup, readiness and liveness probes.

And if the Pod itself isn’t the problem, start looking one level higher: the Deployment, ReplicaSet, Service, ConfigMap, Secret, PersistentVolumeClaim or even the node the Pod is running on.

One example: CrashLoopBackOff

Let’s put some of this together. Imagine we run:

kubectl get pods

and get:

NAME                       READY   STATUS             RESTARTS   AGE
web-6f7d9c8b6f-4xk2m      0/1     CrashLoopBackOff   7          4m

First:

kubectl describe pod web-6f7d9c8b6f-4xk2m

The Events show repeated container restarts, but nothing obviously wrong.

So we check the logs:

kubectl logs web-6f7d9c8b6f-4xk2m

Maybe we get:

Starting web application...
Loading configuration...
ERROR: missing required environment variable DATABASE_URL

Now we have a much better idea of what’s going on.

We can inspect the Deployment:

kubectl describe deployment web

and check how the environment variable is configured.

Maybe it should come from a Secret:

env:
  - name: DATABASE_URL
    valueFrom:
      secretKeyRef:
        name: web-config
        key: database-url

At this point, we might inspect the Secret, its name and the Deployment configuration and hopefully find the mismatch.

The important part isn’t that we knew the answer beforehand. It’s that we worked our way from the symptom to the cause.

Debugging is about narrowing things down

Kubernetes gives us a lot of information when something goes wrong. The challenge is figuring out which information is useful for the problem we’re currently looking at.

A Pod in CrashLoopBackOff needs a different investigation than a Pod stuck in Pending. A Pod that is Running but isn’t receiving traffic might lead us to its readiness probe or Service. A Pod that can’t start because of a missing volume will probably lead us somewhere completely different.

So don’t just randomly throw kubectl commands at the problem. Start with the symptom, follow the clues, and gradually make the problem smaller. Sometimes that means looking inside the Pod. Sometimes it means moving up the stack to the Deployment or Service. And sometimes the problem isn’t in the Pod at all.

Summing up

There is no single kubectl command that magically tells us why our application is broken. But there is a pretty useful toolbox as we’ve seen in this article. And the most important part is probably not a command at all: Start with what you know, make the problem smaller, and work your way towards the cause.

And quite often, one of the first places to look for clues is Kubernetes Events. They give us a glimpse into what Kubernetes itself has noticed and what its components have been doing. But there is more to Events than just the lines at the bottom of kubectl describe.

So in the next article, we’ll take a closer look at Kubernetes Events: what they are, where they come from, how to query them, and what they can (and can’t) tell us when we’re trying to understand what’s happening in our cluster.