--- # Virtualization, Containers & Docker ### Modern Infrastructure Essentials From physical servers to portable applications --- ## Learning Objectives By the end of this lecture, you will be able to: - Explain why teams virtualize compute resources - Describe virtual machines and hypervisors - Explain how containers differ from virtual machines - Explain how Docker images are built and executed as containers - Describe how Compose models a multi-service application - Identify important security and production tradeoffs --- ## Why Does Infrastructure Change? ### One application. Different environments. Different surprises. --- ## The “Works on My Machine” Problem - A developer's computer has one operating system and set of libraries. - A test server may have different versions and configuration. - Production has real users, data, and stricter limits. - Manual setup makes machines drift apart over time.
Goal: describe an environment once, then reproduce it reliably.
--- ## From Hardware to Applications ```text Application Libraries and runtime Operating system Virtualization or container runtime Physical CPU, memory, storage, network ``` Virtualization and containers help us use hardware efficiently and package workloads predictably. --- ## Part 1: Virtualization ### Many computers, built from one physical machine --- ## What Is Virtualization? Virtualization makes software behave as if it has its own computer, even while sharing a physical host. - **Host:** the physical computer providing resources - **Guest:** a virtual machine running on that host - **Hypervisor:** software that creates and manages virtual machines --- ## Meet the Hypervisor A hypervisor allocates physical resources to virtual machines and separates their execution environments. ```text Physical server |-- Hypervisor |-- VM 1: guest OS + application A |-- VM 2: guest OS + application B |-- VM 3: guest OS + application C ``` Each VM runs its own operating system kernel. --- ## Type 1 and Type 2 Hypervisors | Type | Where it runs | Common context | | --- | --- | --- | | Type 1 | Directly on server hardware | Data centers and cloud platforms | | Type 2 | As an application on a host OS | Learning, testing, desktop use | Examples include ESXi and Hyper-V in server environments, and VirtualBox on desktops. --- ## What’s Inside a VM? - Virtual CPU, memory, disks, and network devices - A complete guest operating system - Application code, runtime, and dependencies **Tradeoff:** flexibility and separation, with more resources and OS maintenance per workload. --- ## Why Virtual Machines Matter Useful when you need: - Different operating systems on one physical server - A strong boundary between workloads or tenants - Legacy applications tied to a particular OS - A full machine image for recovery or testing --- ## Part 2: Containers ### Isolated application processes sharing an operating system kernel --- ## What Is a Container? A container is an isolated process packaged with the files and configuration it needs to run. - Containers share the host operating system kernel. - They isolate processes, filesystems, networks, and resource usage. - They do not each boot a complete guest OS.
A container is not a tiny VM. It is a different isolation model.
--- ## How Does Linux Isolate Containers? - **Namespaces** give processes separate views of resources such as processes, network, and mounts. - **Control groups (cgroups)** account for and limit CPU, memory, and other resources. - **Capabilities and security controls** restrict what processes can do. These are operating-system features used by container tools; containers are not a separate kernel feature. --- ## VMs and Containers Side by Side | | Virtual machine | Container | | --- | --- | --- | | Isolation unit | Guest OS and virtual hardware | Isolated process | | Kernel | Each VM has a guest kernel | Shared with host | | Typical footprint | Larger | Smaller | | Startup | Boot an OS | Start a process | | Good fit | OS flexibility, stronger boundary | Portable app packaging and density | These are typical tradeoffs, not guarantees. Workload and configuration matter. --- ## VMs and Containers Work Together ```text Cloud physical server |-- VM with Linux guest OS |-- Container runtime |-- API container |-- Web container |-- Worker container ``` Containers often run inside VMs. The technologies are complementary. --- ## When to Choose Which? - **Choose a VM** for a different guest OS, legacy OS-level needs, or a stronger isolation boundary. - **Choose containers** for consistent and efficient application packaging and deployment. - **Use both** when container workloads need VM-based infrastructure isolation.
Start with requirements, not the trend of the year.
--- ## Part 3: Docker ### Tools for building, sharing, and running containers --- ## What Is Docker? Docker is a platform and toolchain for working with container images and containers. - **Dockerfile:** instructions for building an image - **Image:** a versioned, read-only package used to create containers - **Container:** a running instance created from an image - **Registry:** a place to store and distribute images - **Docker Engine:** software that builds and runs containers --- ## Image or Container? Use the class analogy: - **Image:** the recipe or master copy - **Container:** a running copy made from that image ```text one image creates many containers app:1.0 app instance A app instance B ``` The image is immutable by design; a container has a writable runtime layer. --- ## Docker’s Main Building Blocks - **Docker CLI:** a client interface for submitting build and lifecycle requests - **Docker Engine:** manages images and containers - **containerd and runc:** components involved in container lifecycle and execution - **Registry:** remote image storage such as Docker Hub or a private registry The CLI is the steering wheel; the engine and runtime components do the work. --- ## The Docker Workflow ```text Source + build definition | image build v OCI image artifact <----> Registry | image retrieval v Runtime configuration + container process ``` The image is the portable artifact. A runtime turns it into an isolated process with a configured filesystem, network, and resource limits. --- ## Image Layers and Tags - Images use layers that can be reused and cached. - A tag is a readable label, such as `my-api:1.4`. - A tag can move; a digest identifies exact image content. - Do not treat `latest` as a reliable version pin. --- docker rm -f web-demo ## Container Lifecycle 1. The runtime retrieves and verifies image content when it is not already available locally. 2. The image layers are assembled into a root filesystem view. 3. The runtime configures namespaces, cgroups, capabilities, mounts, and networking. 4. The configured entrypoint starts as a process inside the container. 5. The process exit ends that container instance; restart policy is a separate runtime decision. A container is not a miniature machine booting an operating system. Its lifecycle is centered on a process. --- ## Docker Engine Request Path ```text Docker client | API request Docker Engine | lifecycle coordination containerd | OCI runtime setup runc | create namespaces, mounts, cgroups Container process ``` The exact implementation can vary by platform and version. On macOS and Windows, Docker Desktop adds a Linux VM layer for Linux containers. --- ## Dockerfile and Image Construction A Dockerfile is a declarative build recipe. Instructions describe a base image, filesystem changes, metadata, and the default process configuration. ```dockerfile FROM nginx:alpine COPY index.html /usr/share/nginx/html/index.html ``` `FROM` selects the base filesystem; `COPY` adds content from the build context. The resulting image is immutable content, not a running service. --- ## Image Layers and Build Cache - Filesystem-changing build instructions typically produce reusable layers. - A cache hit can reuse a prior layer when its inputs and instruction are unchanged. - A change invalidates dependent later layers, so stable dependencies usually precede frequently changing source files. - Multi-stage builds separate build tools from the smaller final runtime image. - `.dockerignore` excludes irrelevant or sensitive files from the build context. Layer caching improves build efficiency; it does not replace dependency pinning or image verification. --- ## Image Format and Distribution - An image manifest references configuration and filesystem layers by content digest. - A tag such as `service:2.1` is a mutable name; a digest identifies exact content. - A multi-platform image index can point to platform-specific variants, such as Linux on ARM64 or AMD64. - Registries store and distribute image content; clients retrieve the manifest and only the required layers. These properties support caching, deduplication, reproducible references, and cross-environment distribution. --- ## Container Networking - Each container can have a network namespace with its own interfaces, routes, and addresses. - A bridge network connects containers on one host; name-based service discovery avoids fixed container IPs. - Published ports create a path from a host interface to a container port, often using address translation and firewall rules. - Host networking shares the host network namespace; `none` provides no configured external network. - Overlay networks connect workloads across hosts when the platform provides that capability. Network isolation is a configuration property, not an automatic guarantee that every service is private. --- ## Container Storage - Image layers are read-only; each container adds a disposable writable layer. - Data in that writable layer is coupled to that container instance and is not a durable database strategy. - **Volumes** are runtime-managed storage locations that can outlive containers. - **Bind mounts** expose a selected host path inside a container; **tmpfs** keeps transient data in memory. - Backups, permissions, consistency, and lifecycle ownership still need explicit design. --- ## Runtime Configuration and Process Model - Environment variables and mounted configuration can vary runtime behavior without rebuilding an image. - Secrets require a dedicated secret-management mechanism; ordinary image layers and environment settings can expose sensitive values. - The main process runs as PID 1 in its process namespace and must handle termination signals correctly. - Health checks report application readiness or liveness; they do not make an unhealthy service self-healing by themselves. - Logs and metrics need to be collected outside the container's disposable filesystem. --- ## Multiple Services with Compose ### Describe an application stack in one YAML file --- ## Compose as an Application Model Compose describes a project as services, networks, volumes, and configuration. - A service defines a workload from an image or build context. - A project network provides service-name discovery between workloads. - Named volumes represent persistent storage resources. - Port publication describes which service ports are reachable from outside the project. Compose is a useful application model; it is not a cluster scheduler or a production orchestration control plane. --- ## A Small Compose Stack ```yaml services: web: image: nginx:alpine ports: - "8080:80" depends_on: - api api: image: hashicorp/http-echo:1.0.0 command: ["-text=Hello from the API"] ``` --- docker compose logs -f docker compose down ## Startup Order Is Not Readiness - A dependency declaration can express startup ordering between services. - A process being started does not mean it is ready to accept requests. - Applications should tolerate dependency delays and reconnect when appropriate. - Health checks provide signals; orchestration policy determines whether and how those signals trigger action. - Persistent volumes have lifecycles distinct from the containers that mount them. These distinctions prevent common race conditions and accidental data loss. --- ## From Compose to Orchestration - A container runtime creates individual container processes on a host. - An orchestrator schedules workloads across a pool of machines and tracks desired versus observed state. - Controllers can replace failed instances, distribute traffic, and coordinate rolling updates. - Service discovery and persistent storage become cluster-level concerns. - Kubernetes is one orchestration platform; it adds abstractions and operational complexity beyond Compose. --- ## Production Thinking ### Containers make deployment easier; operations still matter --- ## Container Security Basics - Use trusted, maintained base images and rebuild regularly. - Use a non-root user when the application supports it. - Do not put passwords, tokens, or private keys in an image. - Set resource limits and grant only the permissions needed. - Scan images and review what is actually included. --- docker logs
docker inspect
docker exec -it
sh ## Observability and Failure Domains - **Logs** explain discrete events; **metrics** describe numeric behavior over time; **traces** connect a request across services. - Container restarts can hide symptoms unless logs and state are externalized. - A container boundary is not necessarily a failure boundary: containers on one host still share the host kernel and resources. - Resource limits constrain consumption, while capacity planning determines whether the host can meet combined demand. - Readiness, liveness, and startup probes answer different operational questions. --- ## Key Takeaways ### A mental model you can carry into the next project --- ## Remember the Layers - A hypervisor virtualizes hardware so guest operating systems can run on a host. - A container packages and isolates an application process while sharing a kernel. - Docker provides tools to build, distribute, and run container images. - Compose coordinates multiple containers for one application. - Images are replaceable; important data belongs in persistent storage. --- ## Choose with Context Use a VM for a complete guest OS or a stronger machine boundary. Use containers to package applications for consistent deployment. Combine them when the architecture calls for both.
Understand the boundary you are using, and design security, data, and operations around it.
--- ## Q&A - When would a VM be the better choice than a container? - What happens to files when a container is removed? - How would you pass configuration without baking it into an image? - Why might a cloud container still run inside a VM?