daemon-sec-cheatsheet

The cheatsheet vault for operators: AD, enumeration, exploitation, priv-esc, web, DFIR
git clone https://git.daemon-sec.xyz/daemon-sec-cheatsheet.git
Log | Files | Refs | README | LICENSE

assessment-and-hardening.md (11697B)


      1 ---
      2 title: "Assessment And Hardening"
      3 section: "Linux"
      4 sectionSlug: "linux-hardening"
      5 sourcePath: "src/linux-hardening/containers-namespaces/container-security/assessment-and-hardening.md"
      6 sourceUrl: "https://github.com/HackTricks-wiki/hacktricks/blob/188de82beb54e70956b2952367a0af91d26758b8/src/linux-hardening/containers-namespaces/container-security/assessment-and-hardening.md"
      7 sha: "188de82beb54e70956b2952367a0af91d26758b8"
      8 isIndex: false
      9 modified: true
     10 license: "CC-BY-NC-4.0"
     11 ---
     12 
     13 # Assessment And Hardening
     14 
     15 ## Overview
     16 
     17 A good container assessment should answer two parallel questions. First, what can an attacker do from the current workload? Second, which operator choices made that possible? Enumeration tools help with the first question, and hardening guidance helps with the second. Keeping both on one page makes the section more useful as a field reference rather than just a catalog of escape tricks.
     18 
     19 One practical update for modern environments is that many older container writeups quietly assume a **rootful runtime**, **no user namespace isolation**, and often **cgroup v1**. Those assumptions are not safe anymore. Before spending time on old escape primitives, first confirm whether the workload is rootless or userns-remapped, whether the host is using cgroup v2, and whether Kubernetes or the runtime is now applying default seccomp and AppArmor profiles. These details often decide whether a famous breakout still applies.
     20 
     21 ## Enumeration Tools
     22 
     23 A number of tools remain useful for quickly characterizing a container environment:
     24 
     25 - `linpeas` can identify many container indicators, mounted sockets, capability sets, dangerous filesystems, and breakout hints.
     26 - `CDK` focuses specifically on container environments and includes enumeration plus some automated escape checks.
     27 - `amicontained` is lightweight and useful for identifying container restrictions, capabilities, namespace exposure, and likely breakout classes.
     28 - `deepce` is another container-focused enumerator with breakout-oriented checks.
     29 - `grype` is useful when the assessment includes image-package vulnerability review instead of only runtime escape analysis.
     30 - `Tracee` is useful when you need **runtime evidence** rather than static posture alone, especially for suspicious process execution, file access, and container-aware event collection.
     31 - `Inspektor Gadget` is useful in Kubernetes and Linux-host investigations when you need eBPF-backed visibility tied back to pods, containers, namespaces, and other higher-level concepts.
     32 
     33 The value of these tools is speed and coverage, not certainty. They help reveal the rough posture quickly, but the interesting findings still need manual interpretation against the actual runtime, namespace, capability, and mount model.
     34 
     35 ## Hardening Priorities
     36 
     37 The most important hardening principles are conceptually simple even though their implementation varies by platform. Avoid privileged containers. Avoid mounted runtime sockets. Do not give containers writable host paths unless there is a very specific reason. Use user namespaces or rootless execution where feasible. Drop all capabilities and add back only the ones the workload truly needs. Keep seccomp, AppArmor, and SELinux enabled rather than disabling them to fix application compatibility problems. Limit resources so that a compromised container cannot trivially deny service to the host.
     38 
     39 Image and build hygiene matter as much as runtime posture. Use minimal images, rebuild frequently, scan them, require provenance where practical, and keep secrets out of layers. A container running as non-root with a small image and a narrow syscall and capability surface is much easier to defend than a large convenience image running as host-equivalent root with debugging tools preinstalled.
     40 
     41 For Kubernetes, current hardening baselines are more opinionated than many operators still assume. The built-in **Pod Security Standards** treat `restricted` as the "current best practice" profile: `allowPrivilegeEscalation` should be `false`, workloads should run as non-root, seccomp should be explicitly set to `RuntimeDefault` or `Localhost`, and capability sets should be dropped aggressively. During assessment, this matters because a cluster that is only using `warn` or `audit` labels may look hardened on paper while still admitting risky pods in practice.<sup>[[1]](#references)</sup>
     42 
     43 ## Modern Triage Questions
     44 
     45 Before diving into escape-specific pages, answer these quick questions:
     46 
     47 1. Is the workload **rootful**, **rootless**, or **userns-remapped**?
     48 2. Is the node using **cgroup v1** or **cgroup v2**?
     49 3. Are **seccomp** and **AppArmor/SELinux** explicitly configured, or merely inherited when available?
     50 4. In Kubernetes, is the namespace actually **enforcing** `baseline` or `restricted`, or only warning/auditing?
     51 
     52 Useful checks:
     53 
     54 ```bash
     55 id
     56 cat /proc/self/uid_map 2>/dev/null
     57 cat /proc/self/gid_map 2>/dev/null
     58 stat -fc %T /sys/fs/cgroup 2>/dev/null
     59 cat /sys/fs/cgroup/cgroup.controllers 2>/dev/null
     60 grep -E 'Seccomp|NoNewPrivs' /proc/self/status
     61 cat /proc/1/attr/current 2>/dev/null
     62 find /var/run/secrets -maxdepth 3 -type f 2>/dev/null | head
     63 NS=$(cat /var/run/secrets/kubernetes.io/serviceaccount/namespace 2>/dev/null)
     64 kubectl get ns "$NS" -o jsonpath='{.metadata.labels}' 2>/dev/null
     65 kubectl get pod "$HOSTNAME" -n "$NS" -o jsonpath='{.spec.securityContext.supplementalGroupsPolicy}{"\n"}' 2>/dev/null
     66 kubectl get pod "$HOSTNAME" -n "$NS" -o jsonpath='{.spec.securityContext.seccompProfile.type}{"\n"}{.spec.containers[*].securityContext.allowPrivilegeEscalation}{"\n"}{.spec.containers[*].securityContext.capabilities.drop}{"\n"}' 2>/dev/null
     67 ```
     68 
     69 What is interesting here:
     70 
     71 - If `/proc/self/uid_map` shows container root mapped to a **high host UID range**, many older host-root writeups become less relevant because root in the container is no longer host-root equivalent.
     72 - If `/sys/fs/cgroup` is `cgroup2fs`, old **cgroup v1**-specific writeups such as `release_agent` abuse should no longer be your first guess.
     73 - If seccomp and AppArmor are only inherited implicitly, portability can be weaker than defenders expect. In Kubernetes, explicitly setting `RuntimeDefault` is often stronger than silently relying on node defaults.
     74 - If `supplementalGroupsPolicy` is set to `Strict`, the pod should avoid silently inheriting extra group memberships from `/etc/group` inside the image, which makes group-based volume and file access behavior more predictable.
     75 - Namespace labels such as `pod-security.kubernetes.io/enforce=restricted` are worth checking directly. `warn` and `audit` are useful, but they do not stop a risky pod from being created.
     76 
     77 ## Runtime Baseline Triage
     78 
     79 A runtime baseline is the quick pass that tells you whether a container looks like an ordinary isolated workload or like a host-impacting control plane foothold. It should collect enough facts to prioritize the next page to read: runtime socket abuse, host mounts, namespaces, cgroups, capabilities, or image-secret review.
     80 
     81 Useful checks from inside a workload:
     82 
     83 ```bash
     84 id
     85 hostname
     86 cat /proc/1/cgroup 2>/dev/null
     87 cat /proc/self/uid_map 2>/dev/null
     88 grep -E 'CapEff|Seccomp|NoNewPrivs' /proc/self/status
     89 stat -fc %T /sys/fs/cgroup 2>/dev/null
     90 cat /sys/fs/cgroup/memory.max 2>/dev/null
     91 cat /sys/fs/cgroup/pids.max 2>/dev/null
     92 readlink /proc/self/ns/{pid,mnt,net,ipc,cgroup,user} 2>/dev/null
     93 mount
     94 find /run /var/run -maxdepth 3 \( -name docker.sock -o -name containerd.sock -o -name crio.sock -o -name podman.sock \) 2>/dev/null
     95 ```
     96 
     97 Interpretation:
     98 
     99 - Missing or unlimited `memory.max` / `pids.max` points to weak blast-radius controls even without a clean escape.
    100 - A root shell with `NoNewPrivs: 0`, broad capabilities, and permissive seccomp is much more interesting than a narrow non-root workload.
    101 - Runtime sockets and writable host mounts usually outrank kernel exploits because they already expose a management or filesystem control path.
    102 - Shared PID, network, IPC, or cgroup namespaces are not always full escapes by themselves, but they make the next step easier to find.
    103 
    104 ## Resource-Exhaustion Examples
    105 
    106 Resource controls are not glamorous, but they are part of container security because they limit the blast radius of compromise. Without memory, CPU, or PID limits, a simple shell may be enough to degrade the host or neighboring workloads.
    107 
    108 Example host-impacting tests:
    109 
    110 ```bash
    111 stress-ng --vm 1 --vm-bytes 1G --verify -t 5m
    112 docker run -d --name malicious-container -c 512 busybox sh -c 'while true; do :; done'
    113 nc -lvp 4444 >/dev/null & while true; do cat /dev/urandom | nc <target_ip> 4444; done
    114 ```
    115 
    116 These examples are useful because they show that not every dangerous container outcome is a clean "escape". Weak cgroup limits can still turn code execution into real operational impact.
    117 
    118 In Kubernetes-backed environments, also check whether resource controls exist at all before treating DoS as theoretical:
    119 
    120 ```bash
    121 kubectl get pod "$HOSTNAME" -n "$NS" -o jsonpath='{range .spec.containers[*]}{.name}{" cpu="}{.resources.limits.cpu}{" mem="}{.resources.limits.memory}{"\n"}{end}' 2>/dev/null
    122 cat /sys/fs/cgroup/pids.max 2>/dev/null
    123 cat /sys/fs/cgroup/memory.max 2>/dev/null
    124 cat /sys/fs/cgroup/cpu.max 2>/dev/null
    125 ```
    126 
    127 ## Hardening Tooling
    128 
    129 For Docker-centric environments, `docker-bench-security` remains a useful host-side audit baseline because it checks common configuration issues against widely recognized benchmark guidance. Runtime advisories should be reviewed separately because a configuration benchmark cannot detect whether the installed runtime is affected by a newly disclosed vulnerability. <sup>[[2]](#references)</sup>
    130 
    131 ```bash
    132 git clone https://github.com/docker/docker-bench-security.git
    133 cd docker-bench-security
    134 sudo sh docker-bench-security.sh
    135 ```
    136 
    137 The tool is not a substitute for threat modeling, but it is still valuable for finding careless daemon, mount, network, and runtime defaults that accumulate over time.
    138 
    139 For Kubernetes and runtime-heavy environments, pair static checks with runtime visibility:
    140 
    141 - `Tracee` is useful for container-aware runtime detection and quick forensics when you need to confirm what a compromised workload actually touched.
    142 - `Inspektor Gadget` is useful when the assessment needs kernel-level telemetry mapped back to pods, containers, DNS activity, file execution, or network behavior.
    143 
    144 ## Checks
    145 
    146 Use these as quick first-pass commands during assessment:
    147 
    148 ```bash
    149 id
    150 capsh --print 2>/dev/null
    151 grep -E 'Seccomp|NoNewPrivs' /proc/self/status
    152 cat /proc/self/uid_map 2>/dev/null
    153 stat -fc %T /sys/fs/cgroup 2>/dev/null
    154 mount
    155 find / -maxdepth 3 \( -name docker.sock -o -name containerd.sock -o -name crio.sock -o -name podman.sock \) 2>/dev/null
    156 ```
    157 
    158 What is interesting here:
    159 
    160 - A root process with broad capabilities and `Seccomp: 0` deserves immediate attention.
    161 - A root process that also has a **1:1 UID map** is far more interesting than "root" inside a properly isolated user namespace.
    162 - `cgroup2fs` usually means many older **cgroup v1** escape chains are not your best starting point, while missing `memory.max` or `pids.max` still points to weak blast-radius controls.
    163 - Suspicious mounts and runtime sockets often provide a faster path to impact than any kernel exploit.
    164 - The combination of weak runtime posture and weak resource limits usually indicates a generally permissive container environment rather than a single isolated mistake.
    165 
    166 ## References
    167 
    168 - [1] [Kubernetes Pod Security Standards](https://kubernetes.io/docs/concepts/security/pod-security-standards/)
    169 - [2] [Docker Security Advisory: Multiple Vulnerabilities in runc, BuildKit, and Moby](https://docs.docker.com/security/security-announcements/)