seccomp.md (14098B)
1 --- 2 title: "seccomp" 3 section: "Linux" 4 sectionSlug: "linux-hardening" 5 sourcePath: "src/linux-hardening/containers-namespaces/container-security/protections/seccomp.md" 6 sourceUrl: "https://github.com/HackTricks-wiki/hacktricks/blob/188de82beb54e70956b2952367a0af91d26758b8/src/linux-hardening/containers-namespaces/container-security/protections/seccomp.md" 7 sha: "188de82beb54e70956b2952367a0af91d26758b8" 8 isIndex: false 9 modified: true 10 license: "CC-BY-NC-4.0" 11 --- 12 13 # seccomp 14 15 ## Overview 16 17 **seccomp** is the mechanism that lets the kernel apply a filter to the syscalls a process may invoke. In containerized environments, seccomp is normally used in filter mode so that the process is not simply marked "restricted" in a vague sense, but is instead subject to a concrete syscall policy. This matters because many container breakouts require reaching very specific kernel interfaces. If the process cannot successfully invoke the relevant syscalls, a large class of attacks disappears before any namespace or capability nuance even becomes relevant. 18 19 The key mental model is simple: namespaces decide **what the process can see**, capabilities decide **which privileged actions the process is nominally allowed to attempt**, and seccomp decides **whether the kernel will even accept the syscall entry point for the attempted action**. This is why seccomp frequently prevents attacks that would otherwise look possible based on capabilities alone. 20 21 ## Security Impact 22 23 A lot of dangerous kernel surface is reachable only through a relatively small set of syscalls. Examples that repeatedly matter in container hardening include `mount`, `unshare`, `clone` or `clone3` with particular flags, `bpf`, `ptrace`, `keyctl`, and `perf_event_open`. An attacker who can reach those syscalls may be able to create new namespaces, manipulate kernel subsystems, or interact with attack surface that a normal application container does not need at all. 24 25 This is why default runtime seccomp profiles are so important. They are not merely "extra defense". In many environments they are the difference between a container that can exercise a broad portion of kernel functionality and one that is constrained to a syscall surface closer to what the application genuinely needs. 26 27 ## Modes And Filter Construction 28 29 seccomp historically had a strict mode in which only a tiny syscall set remained available, but the mode relevant to modern container runtimes is seccomp filter mode, often called **seccomp-bpf**. In this model, the kernel evaluates a filter program that decides whether a syscall should be allowed, denied with an errno, trapped, logged, or kill the process.<sup>[[1]](#references)</sup> Container runtimes use this mechanism because it is expressive enough to block broad classes of dangerous syscalls while still allowing normal application behavior. 30 31 Two low-level examples are useful because they make the mechanism concrete rather than magical. Strict mode demonstrates the old "only a minimal syscall set survives" model: 32 33 ```c 34 #include <fcntl.h> 35 #include <linux/seccomp.h> 36 #include <stdio.h> 37 #include <string.h> 38 #include <sys/prctl.h> 39 #include <unistd.h> 40 41 int main(void) { 42 int output = open("output.txt", O_WRONLY); 43 const char *val = "test"; 44 prctl(PR_SET_SECCOMP, SECCOMP_MODE_STRICT); 45 write(output, val, strlen(val) + 1); 46 open("output.txt", O_RDONLY); 47 } 48 ``` 49 50 The final `open` causes the process to be killed because it is not part of strict mode's minimal set. 51 52 A libseccomp filter example shows the modern policy model more clearly: 53 54 ```c 55 #include <errno.h> 56 #include <seccomp.h> 57 #include <stdio.h> 58 #include <unistd.h> 59 60 int main(void) { 61 scmp_filter_ctx ctx = seccomp_init(SCMP_ACT_KILL); 62 seccomp_rule_add(ctx, SCMP_ACT_ALLOW, SCMP_SYS(exit_group), 0); 63 seccomp_rule_add(ctx, SCMP_ACT_ERRNO(EBADF), SCMP_SYS(getpid), 0); 64 seccomp_rule_add(ctx, SCMP_ACT_ALLOW, SCMP_SYS(brk), 0); 65 seccomp_rule_add(ctx, SCMP_ACT_ALLOW, SCMP_SYS(write), 2, 66 SCMP_A0(SCMP_CMP_EQ, 1), 67 SCMP_A2(SCMP_CMP_LE, 512)); 68 seccomp_rule_add(ctx, SCMP_ACT_ERRNO(EBADF), SCMP_SYS(write), 1, 69 SCMP_A0(SCMP_CMP_NE, 1)); 70 seccomp_load(ctx); 71 seccomp_release(ctx); 72 printf("pid=%d\n", getpid()); 73 } 74 ``` 75 76 This style of policy is what most readers should picture when they think about runtime seccomp profiles. 77 78 ## Lab 79 80 A simple way to confirm that seccomp is active in a container is: 81 82 ```bash 83 docker run --rm debian:stable-slim sh -c 'grep Seccomp /proc/self/status' 84 docker run --rm --security-opt seccomp=unconfined debian:stable-slim sh -c 'grep Seccomp /proc/self/status' 85 ``` 86 87 You can also try an operation that default profiles commonly restrict: 88 89 ```bash 90 docker run --rm debian:stable-slim sh -c 'apt-get update >/dev/null 2>&1 && apt-get install -y util-linux >/dev/null 2>&1 && unshare -Ur true' 91 ``` 92 93 If the container is running under a normal default seccomp profile, `unshare`-style operations are often blocked. This is a useful demonstration because it shows that even if the userspace tool exists inside the image, the kernel path it needs may still be unavailable. 94 If the container is running under a normal default seccomp profile, `unshare`-style operations are often blocked even when the userspace tool exists inside the image. 95 96 To inspect the process status more generally, run: 97 98 ```bash 99 grep -E 'Seccomp|NoNewPrivs' /proc/self/status 100 ``` 101 102 ## Runtime Usage 103 104 Docker supports both default and custom seccomp profiles and allows administrators to disable them with `--security-opt seccomp=unconfined`.<sup>[[2]](#references)</sup> Podman has similar support and often pairs seccomp with rootless execution in a very sensible default posture. Kubernetes exposes seccomp through workload configuration, where `RuntimeDefault` is usually the sane baseline and `Unconfined` should be treated as an exception requiring justification rather than as a convenience toggle.<sup>[[3]](#references)</sup> 105 106 In containerd and CRI-O based environments, the exact path is more layered, but the principle is the same: the higher-level engine or orchestrator decides what should happen, and the runtime eventually installs the resulting seccomp policy for the container process. The outcome still depends on the final runtime configuration that reaches the kernel. 107 108 ### Custom Policy Example 109 110 Docker and similar engines can load a custom seccomp profile from JSON. A minimal example that denies `chmod` while allowing everything else looks like this: 111 112 ```json 113 { 114 "defaultAction": "SCMP_ACT_ALLOW", 115 "syscalls": [ 116 { 117 "name": "chmod", 118 "action": "SCMP_ACT_ERRNO" 119 } 120 ] 121 } 122 ``` 123 124 Applied with: 125 126 ```bash 127 docker run --rm -it --security-opt seccomp=/path/to/profile.json busybox chmod 400 /etc/hosts 128 ``` 129 130 The command fails with `Operation not permitted`, demonstrating that the restriction comes from the syscall policy rather than from ordinary file permissions alone. In real hardening, allowlists are generally stronger than permissive defaults with a small blacklist. 131 132 ## Misconfigurations 133 134 The bluntest mistake is to set seccomp to **unconfined** because an application failed under the default policy. This is common during troubleshooting and very dangerous as a permanent fix. Once the filter is gone, many syscall-based breakout primitives become reachable again, especially when powerful capabilities or host namespace sharing are also present. 135 136 Another frequent problem is the use of a **custom permissive profile** that was copied from some blog or internal workaround without being reviewed carefully. Teams sometimes retain almost all dangerous syscalls simply because the profile was built around "stop the app from breaking" rather than "grant only what the app actually needs". A third misconception is to assume seccomp is less important for non-root containers. In reality, plenty of kernel attack surface remains relevant even when the process is not UID 0. 137 138 ## Abuse 139 140 If seccomp is absent or badly weakened, an attacker may be able to invoke namespace-creation syscalls, expand the reachable kernel attack surface through `bpf` or `perf_event_open`, abuse `keyctl`, or combine those syscall paths with dangerous capabilities such as `CAP_SYS_ADMIN`. In many real attacks, seccomp is not the only missing control, but its absence shortens the exploit path dramatically because it removes one of the few defenses that can stop a risky syscall before the rest of the privilege model even comes into play. 141 142 The most useful practical test is to try the exact syscall families that default profiles usually block. If they suddenly work, the container posture has changed a lot: 143 144 ```bash 145 grep Seccomp /proc/self/status 146 unshare -Ur true 2>/dev/null && echo "unshare works" 147 unshare -m true 2>/dev/null && echo "mount namespace creation works" 148 ``` 149 150 If `CAP_SYS_ADMIN` or another strong capability is present, test whether seccomp is the only missing barrier before mount-based abuse: 151 152 ```bash 153 capsh --print | grep cap_sys_admin 154 mkdir -p /tmp/m 155 mount -t tmpfs tmpfs /tmp/m 2>/dev/null && echo "tmpfs mount works" 156 mount -t proc proc /tmp/m 2>/dev/null && echo "proc mount works" 157 ``` 158 159 On some targets, the immediate value is not full escape but information gathering and kernel attack-surface expansion. These commands help determine whether especially sensitive syscall paths are reachable: 160 161 ```bash 162 which unshare nsenter strace 2>/dev/null 163 strace -e bpf,perf_event_open,keyctl true 2>&1 | tail 164 ``` 165 166 If seccomp is absent and the container is also privileged in other ways, that is when it makes sense to pivot into the more specific breakout techniques already documented in the legacy container-escape pages. 167 168 ### Full Example: seccomp Was The Only Thing Blocking `unshare` 169 170 On many targets, the practical effect of removing seccomp is that namespace-creation or mount syscalls suddenly start working. If the container also has `CAP_SYS_ADMIN`, the following sequence may become possible: 171 172 ```bash 173 grep Seccomp /proc/self/status 174 capsh --print | grep cap_sys_admin 175 mkdir -p /tmp/nsroot 176 unshare -m sh -c ' 177 mount -t tmpfs tmpfs /tmp/nsroot && 178 mkdir -p /tmp/nsroot/proc && 179 mount -t proc proc /tmp/nsroot/proc && 180 mount | grep /tmp/nsroot 181 ' 182 ``` 183 184 By itself this is not yet a host escape, but it demonstrates that seccomp was the barrier preventing mount-related exploitation. 185 186 ### Full Example: seccomp Disabled + cgroup v1 `release_agent` 187 188 If seccomp is disabled and the container can mount cgroup v1 hierarchies, the `release_agent` technique from the cgroups section becomes reachable: 189 190 ```bash 191 grep Seccomp /proc/self/status 192 mount | grep cgroup 193 unshare -UrCm sh -c ' 194 mkdir /tmp/c 195 mount -t cgroup -o memory none /tmp/c 196 echo 1 > /tmp/c/notify_on_release 197 echo /proc/self/exe > /tmp/c/release_agent 198 (sleep 1; echo 0 > /tmp/c/cgroup.procs) & 199 while true; do sleep 1; done 200 ' 201 ``` 202 203 This is not a seccomp-only exploit. The point is that once seccomp is unconfined, syscall-heavy breakout chains that were previously blocked may start working exactly as written. 204 205 ## Checks 206 207 The purpose of these checks is to establish whether seccomp is active at all, whether `no_new_privs` accompanies it, and whether the runtime configuration shows seccomp being disabled explicitly. 208 209 ```bash 210 grep Seccomp /proc/self/status # Current seccomp mode from the kernel 211 cat /proc/self/status | grep NoNewPrivs # Whether exec-time privilege gain is also blocked 212 docker inspect <container> | jq '.[0].HostConfig.SecurityOpt' # Runtime security options, including seccomp overrides 213 ``` 214 215 What is interesting here: 216 217 - A non-zero `Seccomp` value means filtering is active; `0` usually means no seccomp protection. 218 - If the runtime security options include `seccomp=unconfined`, the workload has lost one of its most useful syscall-level defenses. 219 - `NoNewPrivs` is not seccomp itself, but seeing both together usually indicates a more careful hardening posture than seeing neither. 220 221 If a container already has suspicious mounts, broad capabilities, or shared host namespaces, and seccomp is also unconfined, that combination should be treated as a major escalation signal. The container may still not be trivially breakable, but the number of kernel entry points available to the attacker has increased sharply. 222 223 ## Runtime Defaults 224 225 | Runtime / platform | Default state | Default behavior | Common manual weakening | 226 | --- | --- | --- | --- | 227 | Docker Engine | Usually enabled by default | Uses Docker's built-in default seccomp profile unless overridden | `--security-opt seccomp=unconfined`, `--security-opt seccomp=/path/profile.json`, `--privileged` | 228 | Podman | Usually enabled by default | Applies the runtime default seccomp profile unless overridden | `--security-opt seccomp=unconfined`, `--security-opt seccomp=profile.json`, `--seccomp-policy=image`, `--privileged` | 229 | Kubernetes | **Not guaranteed by default** | If `securityContext.seccompProfile` is unset, the default is `Unconfined` unless the kubelet enables `--seccomp-default`; `RuntimeDefault` or `Localhost` must otherwise be set explicitly | `securityContext.seccompProfile.type: Unconfined`, leaving seccomp unset on clusters without `seccompDefault`, `privileged: true` | 230 | containerd / CRI-O under Kubernetes | Follows Kubernetes node and Pod settings | Runtime profile is used when Kubernetes asks for `RuntimeDefault` or when kubelet seccomp defaulting is enabled | Same as Kubernetes row; direct CRI/OCI configuration can also omit seccomp entirely | 231 232 The Kubernetes behavior is the one that most often surprises operators. In many clusters, seccomp is still absent unless the Pod requests it or the kubelet is configured to default to `RuntimeDefault`.<sup>[[3]](#references)</sup> 233 234 ## References 235 236 - [1] [Linux kernel documentation: Seccomp BPF (SECure COMPuting with filters)](https://docs.kernel.org/userspace-api/seccomp_filter.html) 237 - [2] [Docker Docs: Seccomp security profiles for Docker](https://docs.docker.com/engine/security/seccomp/) 238 - [3] [Kubernetes Docs: Restrict a Container's Syscalls with seccomp](https://kubernetes.io/docs/tutorials/security/seccomp/)