Why Is This Process Running? The Evolution of Process Provenance Tooling
Every sysadmin has met a process that should not exist. An unknown binary eating CPU. A port with a listener nobody opened. The question is always the same: why is this running? For years the answer meant manual archaeology across three or four tools. Then containers broke the last good answer, and a new tool class formed around that single question.
2018 – The /proc years
In 2018 you answered the question by hand. ps -ef --forest showed the tree. lsof -i named the open port. /proc/PID/cmdline and /proc/PID/environ held the rest. An experienced operator stitched it together: check the parent, read the unit file, tail the log. Slow and fragile, but workable, because production boxes were small and stable.
The limit was intent. ps renders the shape of a moment, not the reason it took that shape. auditd logged who spawned what, but only when you configured rules in advance. With no rules you got a frozen tree and a story you had to invent. Archaeology, not forensics.
ps -ef --forest | grep -B2 -A2 31415
lsof -i :8080
cat /proc/31415/cmdline
readlink /proc/31415/exe
2020 – eBPF watches everything
2020 was the eBPF tipping point. bpftrace turned kernel events into one-liners. Falco and sysdig turned execve and clone streams into detection signals. For the first time you could watch a process spawn in kernel space, live, with no patches and no agent.
The cost was weight: privileges, tuning, expertise. And these tools answer what happened far better than why this is here. A trace shows the event. The operator still walks the tree to find the cause. Detection improved. Explanation stayed behind.
2022 – Containers erase the parent
Orchestration then removed the last easy answer. By 2022 the PID you hunted usually belonged to a container. Its supervisor is a runtime. The runtime runs in a pod created by a Deployment, applied from a Helm chart, executed by a pipeline nobody remembers running. systemd and the runtime respawn crashed units, so a process can have a parent and no story.
crictl and kubectl reveal layers, not causes. Provenance stopped being a single-box problem and became a join across the kernel, the runtime, and the control plane. The 2018 ritual could not carry it.
2024 – witr makes provenance one question
witr, short for why is this running, is the obvious-in-hindsight answer. Open source, a single binary with an interactive TUI, and more than 20,000 stars on GitHub. Point it at a process, a port, a container, or a file, and it returns the chain that produced it: the supervisor, the service unit, the shell, the timer, the runtime. Plain text for humans, JSON for scripts, and a dashboard for staring at a screen.
It explains where a running thing came from, how it started, and what chain of systems is responsible for it existing right now.
The tool is deliberately narrow. Origin and behavior are different questions, and witr answers the first. It runs on Linux, macOS, and FreeBSD, with Windows support still experimental. That scope is the point: one question, one answer, no speculation.
Where we are now
witr does not retire ps, lsof, or auditd. Each answers a different question, and all of them remain valid. The change is that provenance became a product instead of a ritual. When a problem spans the kernel, the runtime, and the pipeline, the tool that wins owns the join and drops everything else. witr is the first widely adopted answer to that shape.
Next time something is running and you do not know why, skip the forest of ps. Ask the question and let the tool walk the chain. That is the whole upgrade: one command, one trace, no invented story.