daemon-sec-cheatsheet

The cheatsheet vault for operators: AD, enumeration, exploitation, priv-esc, web, DFIR
git clone https://git.daemon-sec.xyz/daemon-sec-cheatsheet.git
Log | Files | Refs | README | LICENSE

seccomp.md (14098B)


      1 ---
      2 title: "seccomp"
      3 section: "Linux"
      4 sectionSlug: "linux-hardening"
      5 sourcePath: "src/linux-hardening/containers-namespaces/container-security/protections/seccomp.md"
      6 sourceUrl: "https://github.com/HackTricks-wiki/hacktricks/blob/188de82beb54e70956b2952367a0af91d26758b8/src/linux-hardening/containers-namespaces/container-security/protections/seccomp.md"
      7 sha: "188de82beb54e70956b2952367a0af91d26758b8"
      8 isIndex: false
      9 modified: true
     10 license: "CC-BY-NC-4.0"
     11 ---
     12 
     13 # seccomp
     14 
     15 ## Overview
     16 
     17 **seccomp** is the mechanism that lets the kernel apply a filter to the syscalls a process may invoke. In containerized environments, seccomp is normally used in filter mode so that the process is not simply marked "restricted" in a vague sense, but is instead subject to a concrete syscall policy. This matters because many container breakouts require reaching very specific kernel interfaces. If the process cannot successfully invoke the relevant syscalls, a large class of attacks disappears before any namespace or capability nuance even becomes relevant.
     18 
     19 The key mental model is simple: namespaces decide **what the process can see**, capabilities decide **which privileged actions the process is nominally allowed to attempt**, and seccomp decides **whether the kernel will even accept the syscall entry point for the attempted action**. This is why seccomp frequently prevents attacks that would otherwise look possible based on capabilities alone.
     20 
     21 ## Security Impact
     22 
     23 A lot of dangerous kernel surface is reachable only through a relatively small set of syscalls. Examples that repeatedly matter in container hardening include `mount`, `unshare`, `clone` or `clone3` with particular flags, `bpf`, `ptrace`, `keyctl`, and `perf_event_open`. An attacker who can reach those syscalls may be able to create new namespaces, manipulate kernel subsystems, or interact with attack surface that a normal application container does not need at all.
     24 
     25 This is why default runtime seccomp profiles are so important. They are not merely "extra defense". In many environments they are the difference between a container that can exercise a broad portion of kernel functionality and one that is constrained to a syscall surface closer to what the application genuinely needs.
     26 
     27 ## Modes And Filter Construction
     28 
     29 seccomp historically had a strict mode in which only a tiny syscall set remained available, but the mode relevant to modern container runtimes is seccomp filter mode, often called **seccomp-bpf**. In this model, the kernel evaluates a filter program that decides whether a syscall should be allowed, denied with an errno, trapped, logged, or kill the process.<sup>[[1]](#references)</sup> Container runtimes use this mechanism because it is expressive enough to block broad classes of dangerous syscalls while still allowing normal application behavior.
     30 
     31 Two low-level examples are useful because they make the mechanism concrete rather than magical. Strict mode demonstrates the old "only a minimal syscall set survives" model:
     32 
     33 ```c
     34 #include <fcntl.h>
     35 #include <linux/seccomp.h>
     36 #include <stdio.h>
     37 #include <string.h>
     38 #include <sys/prctl.h>
     39 #include <unistd.h>
     40 
     41 int main(void) {
     42   int output = open("output.txt", O_WRONLY);
     43   const char *val = "test";
     44   prctl(PR_SET_SECCOMP, SECCOMP_MODE_STRICT);
     45   write(output, val, strlen(val) + 1);
     46   open("output.txt", O_RDONLY);
     47 }
     48 ```
     49 
     50 The final `open` causes the process to be killed because it is not part of strict mode's minimal set.
     51 
     52 A libseccomp filter example shows the modern policy model more clearly:
     53 
     54 ```c
     55 #include <errno.h>
     56 #include <seccomp.h>
     57 #include <stdio.h>
     58 #include <unistd.h>
     59 
     60 int main(void) {
     61   scmp_filter_ctx ctx = seccomp_init(SCMP_ACT_KILL);
     62   seccomp_rule_add(ctx, SCMP_ACT_ALLOW, SCMP_SYS(exit_group), 0);
     63   seccomp_rule_add(ctx, SCMP_ACT_ERRNO(EBADF), SCMP_SYS(getpid), 0);
     64   seccomp_rule_add(ctx, SCMP_ACT_ALLOW, SCMP_SYS(brk), 0);
     65   seccomp_rule_add(ctx, SCMP_ACT_ALLOW, SCMP_SYS(write), 2,
     66     SCMP_A0(SCMP_CMP_EQ, 1),
     67     SCMP_A2(SCMP_CMP_LE, 512));
     68   seccomp_rule_add(ctx, SCMP_ACT_ERRNO(EBADF), SCMP_SYS(write), 1,
     69     SCMP_A0(SCMP_CMP_NE, 1));
     70   seccomp_load(ctx);
     71   seccomp_release(ctx);
     72   printf("pid=%d\n", getpid());
     73 }
     74 ```
     75 
     76 This style of policy is what most readers should picture when they think about runtime seccomp profiles.
     77 
     78 ## Lab
     79 
     80 A simple way to confirm that seccomp is active in a container is:
     81 
     82 ```bash
     83 docker run --rm debian:stable-slim sh -c 'grep Seccomp /proc/self/status'
     84 docker run --rm --security-opt seccomp=unconfined debian:stable-slim sh -c 'grep Seccomp /proc/self/status'
     85 ```
     86 
     87 You can also try an operation that default profiles commonly restrict:
     88 
     89 ```bash
     90 docker run --rm debian:stable-slim sh -c 'apt-get update >/dev/null 2>&1 && apt-get install -y util-linux >/dev/null 2>&1 && unshare -Ur true'
     91 ```
     92 
     93 If the container is running under a normal default seccomp profile, `unshare`-style operations are often blocked. This is a useful demonstration because it shows that even if the userspace tool exists inside the image, the kernel path it needs may still be unavailable.
     94 If the container is running under a normal default seccomp profile, `unshare`-style operations are often blocked even when the userspace tool exists inside the image.
     95 
     96 To inspect the process status more generally, run:
     97 
     98 ```bash
     99 grep -E 'Seccomp|NoNewPrivs' /proc/self/status
    100 ```
    101 
    102 ## Runtime Usage
    103 
    104 Docker supports both default and custom seccomp profiles and allows administrators to disable them with `--security-opt seccomp=unconfined`.<sup>[[2]](#references)</sup> Podman has similar support and often pairs seccomp with rootless execution in a very sensible default posture. Kubernetes exposes seccomp through workload configuration, where `RuntimeDefault` is usually the sane baseline and `Unconfined` should be treated as an exception requiring justification rather than as a convenience toggle.<sup>[[3]](#references)</sup>
    105 
    106 In containerd and CRI-O based environments, the exact path is more layered, but the principle is the same: the higher-level engine or orchestrator decides what should happen, and the runtime eventually installs the resulting seccomp policy for the container process. The outcome still depends on the final runtime configuration that reaches the kernel.
    107 
    108 ### Custom Policy Example
    109 
    110 Docker and similar engines can load a custom seccomp profile from JSON. A minimal example that denies `chmod` while allowing everything else looks like this:
    111 
    112 ```json
    113 {
    114   "defaultAction": "SCMP_ACT_ALLOW",
    115   "syscalls": [
    116     {
    117       "name": "chmod",
    118       "action": "SCMP_ACT_ERRNO"
    119     }
    120   ]
    121 }
    122 ```
    123 
    124 Applied with:
    125 
    126 ```bash
    127 docker run --rm -it --security-opt seccomp=/path/to/profile.json busybox chmod 400 /etc/hosts
    128 ```
    129 
    130 The command fails with `Operation not permitted`, demonstrating that the restriction comes from the syscall policy rather than from ordinary file permissions alone. In real hardening, allowlists are generally stronger than permissive defaults with a small blacklist.
    131 
    132 ## Misconfigurations
    133 
    134 The bluntest mistake is to set seccomp to **unconfined** because an application failed under the default policy. This is common during troubleshooting and very dangerous as a permanent fix. Once the filter is gone, many syscall-based breakout primitives become reachable again, especially when powerful capabilities or host namespace sharing are also present.
    135 
    136 Another frequent problem is the use of a **custom permissive profile** that was copied from some blog or internal workaround without being reviewed carefully. Teams sometimes retain almost all dangerous syscalls simply because the profile was built around "stop the app from breaking" rather than "grant only what the app actually needs". A third misconception is to assume seccomp is less important for non-root containers. In reality, plenty of kernel attack surface remains relevant even when the process is not UID 0.
    137 
    138 ## Abuse
    139 
    140 If seccomp is absent or badly weakened, an attacker may be able to invoke namespace-creation syscalls, expand the reachable kernel attack surface through `bpf` or `perf_event_open`, abuse `keyctl`, or combine those syscall paths with dangerous capabilities such as `CAP_SYS_ADMIN`. In many real attacks, seccomp is not the only missing control, but its absence shortens the exploit path dramatically because it removes one of the few defenses that can stop a risky syscall before the rest of the privilege model even comes into play.
    141 
    142 The most useful practical test is to try the exact syscall families that default profiles usually block. If they suddenly work, the container posture has changed a lot:
    143 
    144 ```bash
    145 grep Seccomp /proc/self/status
    146 unshare -Ur true 2>/dev/null && echo "unshare works"
    147 unshare -m true 2>/dev/null && echo "mount namespace creation works"
    148 ```
    149 
    150 If `CAP_SYS_ADMIN` or another strong capability is present, test whether seccomp is the only missing barrier before mount-based abuse:
    151 
    152 ```bash
    153 capsh --print | grep cap_sys_admin
    154 mkdir -p /tmp/m
    155 mount -t tmpfs tmpfs /tmp/m 2>/dev/null && echo "tmpfs mount works"
    156 mount -t proc proc /tmp/m 2>/dev/null && echo "proc mount works"
    157 ```
    158 
    159 On some targets, the immediate value is not full escape but information gathering and kernel attack-surface expansion. These commands help determine whether especially sensitive syscall paths are reachable:
    160 
    161 ```bash
    162 which unshare nsenter strace 2>/dev/null
    163 strace -e bpf,perf_event_open,keyctl true 2>&1 | tail
    164 ```
    165 
    166 If seccomp is absent and the container is also privileged in other ways, that is when it makes sense to pivot into the more specific breakout techniques already documented in the legacy container-escape pages.
    167 
    168 ### Full Example: seccomp Was The Only Thing Blocking `unshare`
    169 
    170 On many targets, the practical effect of removing seccomp is that namespace-creation or mount syscalls suddenly start working. If the container also has `CAP_SYS_ADMIN`, the following sequence may become possible:
    171 
    172 ```bash
    173 grep Seccomp /proc/self/status
    174 capsh --print | grep cap_sys_admin
    175 mkdir -p /tmp/nsroot
    176 unshare -m sh -c '
    177   mount -t tmpfs tmpfs /tmp/nsroot &&
    178   mkdir -p /tmp/nsroot/proc &&
    179   mount -t proc proc /tmp/nsroot/proc &&
    180   mount | grep /tmp/nsroot
    181 '
    182 ```
    183 
    184 By itself this is not yet a host escape, but it demonstrates that seccomp was the barrier preventing mount-related exploitation.
    185 
    186 ### Full Example: seccomp Disabled + cgroup v1 `release_agent`
    187 
    188 If seccomp is disabled and the container can mount cgroup v1 hierarchies, the `release_agent` technique from the cgroups section becomes reachable:
    189 
    190 ```bash
    191 grep Seccomp /proc/self/status
    192 mount | grep cgroup
    193 unshare -UrCm sh -c '
    194   mkdir /tmp/c
    195   mount -t cgroup -o memory none /tmp/c
    196   echo 1 > /tmp/c/notify_on_release
    197   echo /proc/self/exe > /tmp/c/release_agent
    198   (sleep 1; echo 0 > /tmp/c/cgroup.procs) &
    199   while true; do sleep 1; done
    200 '
    201 ```
    202 
    203 This is not a seccomp-only exploit. The point is that once seccomp is unconfined, syscall-heavy breakout chains that were previously blocked may start working exactly as written.
    204 
    205 ## Checks
    206 
    207 The purpose of these checks is to establish whether seccomp is active at all, whether `no_new_privs` accompanies it, and whether the runtime configuration shows seccomp being disabled explicitly.
    208 
    209 ```bash
    210 grep Seccomp /proc/self/status                               # Current seccomp mode from the kernel
    211 cat /proc/self/status | grep NoNewPrivs                      # Whether exec-time privilege gain is also blocked
    212 docker inspect <container> | jq '.[0].HostConfig.SecurityOpt'   # Runtime security options, including seccomp overrides
    213 ```
    214 
    215 What is interesting here:
    216 
    217 - A non-zero `Seccomp` value means filtering is active; `0` usually means no seccomp protection.
    218 - If the runtime security options include `seccomp=unconfined`, the workload has lost one of its most useful syscall-level defenses.
    219 - `NoNewPrivs` is not seccomp itself, but seeing both together usually indicates a more careful hardening posture than seeing neither.
    220 
    221 If a container already has suspicious mounts, broad capabilities, or shared host namespaces, and seccomp is also unconfined, that combination should be treated as a major escalation signal. The container may still not be trivially breakable, but the number of kernel entry points available to the attacker has increased sharply.
    222 
    223 ## Runtime Defaults
    224 
    225 | Runtime / platform | Default state | Default behavior | Common manual weakening |
    226 | --- | --- | --- | --- |
    227 | Docker Engine | Usually enabled by default | Uses Docker's built-in default seccomp profile unless overridden | `--security-opt seccomp=unconfined`, `--security-opt seccomp=/path/profile.json`, `--privileged` |
    228 | Podman | Usually enabled by default | Applies the runtime default seccomp profile unless overridden | `--security-opt seccomp=unconfined`, `--security-opt seccomp=profile.json`, `--seccomp-policy=image`, `--privileged` |
    229 | Kubernetes | **Not guaranteed by default** | If `securityContext.seccompProfile` is unset, the default is `Unconfined` unless the kubelet enables `--seccomp-default`; `RuntimeDefault` or `Localhost` must otherwise be set explicitly | `securityContext.seccompProfile.type: Unconfined`, leaving seccomp unset on clusters without `seccompDefault`, `privileged: true` |
    230 | containerd / CRI-O under Kubernetes | Follows Kubernetes node and Pod settings | Runtime profile is used when Kubernetes asks for `RuntimeDefault` or when kubelet seccomp defaulting is enabled | Same as Kubernetes row; direct CRI/OCI configuration can also omit seccomp entirely |
    231 
    232 The Kubernetes behavior is the one that most often surprises operators. In many clusters, seccomp is still absent unless the Pod requests it or the kubelet is configured to default to `RuntimeDefault`.<sup>[[3]](#references)</sup>
    233 
    234 ## References
    235 
    236 - [1] [Linux kernel documentation: Seccomp BPF (SECure COMPuting with filters)](https://docs.kernel.org/userspace-api/seccomp_filter.html)
    237 - [2] [Docker Docs: Seccomp security profiles for Docker](https://docs.docker.com/engine/security/seccomp/)
    238 - [3] [Kubernetes Docs: Restrict a Container's Syscalls with seccomp](https://kubernetes.io/docs/tutorials/security/seccomp/)