Skip to main content

Docker Architecture

The one-sentence version

Docker is a client/server system: the CLI sends REST calls to a long-running daemon, and the daemon delegates the actual container creation to lower-level runtimes that lean on two Linux kernel features — namespaces for isolation and cgroups for resource limits.

Client, daemon, registry

Three moving parts. Nearly every "why did that happen" question traces back to which one was responsible.

ComponentRuns whereResponsibility
docker CLIYour shellFormats requests. Holds no state.
dockerdHost, as rootBuilds images, manages containers, networks, volumes
RegistryRemoteStores and distributes images

The CLI being stateless is why docker ps on a remote host works just by pointing DOCKER_HOST elsewhere — nothing about your containers lives in the client.

It is also the answer to a common security question: the daemon runs as root and its socket is effectively root access. Anyone in the docker group can mount the host filesystem into a container and escape. That is not a bug, it is the consequence of the architecture.

Below the daemon

dockerd does not create containers itself. It hands off down a chain:

The shim is the part worth understanding. One shim process exists per container and owns its stdio and exit code. Because the shim — not the daemon — is the container's parent, you can restart dockerd without killing running containers.

runc is the piece that actually implements the OCI runtime spec. It sets up the namespaces, applies the cgroup limits, then execs your process and exits. There is no Docker process wrapping your app at runtime.

Isolation is a kernel feature

A container is a normal Linux process with a restricted view of the system. There is no container object in the kernel.

Namespaces answer "what can this process see":

NamespaceIsolates
pidProcess IDs — your app sees itself as PID 1
netInterfaces, routes, ports
mntFilesystem mount points
utsHostname
ipcShared memory, semaphores
userUID/GID mapping

cgroups answer "how much can it use" — CPU shares, memory ceilings, block I/O, process counts.

This is why containers start in milliseconds while VMs take tens of seconds: there is no guest kernel to boot and no hardware to emulate. It is also why a Linux container cannot run on a Windows kernel without a Linux VM underneath — the isolation primitives are kernel-specific.

Container vs virtual machine

The tradeoff is isolation strength for density and speed. A VM has a hardware boundary; a container shares the kernel, so a kernel vulnerability is a shared blast radius. That is the reason regulated workloads sometimes still require VMs, or a sandboxed runtime like gVisor or Kata.

Questions

Loading questions…

Learn more