K8 Forge: Production Kubernetes Architecture for Cloud-Native Orchestration
Platform Scope
K8 Forge is a multi-tier log analytics platform built as a single deployable system. It orchestrates five application Deployments, three StatefulSet-backed data services, and a unified log-forge namespace with HPA, PDB, and NetworkPolicy enforcement.
Production patterns implemented across the stack:
Pod runtime β init containers for schema migration, three-probe health models (startup, readiness, liveness), and preStop hooks for zero-downtime connection draining
Service networking β ClusterIP internal discovery, NodePort external access on ports 30080 and 30081, and LoadBalancer session affinity for the same exposure model used from local development through production ingress
Incident simulation β isolated fault-injection namespace with intentional scheduling and networking failures, mirroring production incident patterns that separate routine kubectl usage from SRE-grade diagnosis
Why This Matters
Kubernetes is not a container runtime β it is a distributed systems control plane. The difference between teams that treat it as "Docker with YAML" and teams that understand declarative reconciliation, endpoint consistency, and probe semantics is measured in incident duration. At Airbnb, the initial Kubernetes migration stalled because engineers deployed Pods without readiness probes β traffic routed to containers still initializing database connections, producing cascading 503 errors that looked like application bugs but were orchestration failures. At Netflix, every service must implement separate liveness and readiness endpoints before reaching production; conflating them causes restart loops during dependency outages that amplify partial failures into full outages.
K8 Forge establishes the orchestration substrate for the log analytics pipeline. Deployment rolling updates with maxUnavailable: 0, ConfigMap and Secret separation, Service selector alignment, and probe-driven traffic management are the same patterns used in GitOps pipelines, service mesh configurations, and multi-cluster failover. Engineers who internalize these fundamentals deploy advanced patterns with architectural intuition rather than copy-paste manifests.
Kubernetes Architecture Deep Dive
Declarative Reconciliation and the Pod Lifecycle
The fundamental insight: Kubernetes never "runs" your application β it continuously reconciles desired state against observed state. When you kubectl apply a Deployment with replicas: 3, the Deployment controller creates a ReplicaSet, which creates Pods, which the scheduler places on nodes, which the kubelet starts. If a Pod dies, the ReplicaSet controller creates a replacement. This reconciliation loop is why kubectl delete pod is a valid debugging technique β the system self-heals to match declared intent.
The anti-pattern is treating Pods as persistent entities. Pods are cattle, not pets. The log-collector Deployment runs three replicas with a PDB requiring minAvailable: 2 β during node maintenance, Kubernetes evicts at most one Pod while maintaining quorum. Without PDB, a rolling node drain could eliminate all replicas simultaneously.
Trade-off: PDBs can block node drains if misconfigured. Setting minAvailable equal to replicas prevents any voluntary disruption β correct for stateful systems, problematic for stateless tiers that should tolerate full replacement.
Workload Controllers and Zero-Downtime Deployments
Deployments manage stateless application tiers through ReplicaSets with configurable update strategies. log-processor uses maxSurge: 1, maxUnavailable: 0 β during updates, Kubernetes creates one new Pod before terminating any old Pod. Combined with readiness probes, the Service endpoints table only includes Pods that pass /health/ready, ensuring load balancers never route to initializing containers.
Revision history enables instant rollback: kubectl rollout undo deployment/log-collector reverts to the previous ReplicaSet template. At scale, teams pin rollback SLAs β Spotify's deployment platform automatically rolls back if error rates exceed thresholds within five minutes of a rollout.
Failure mode: Setting maxUnavailable: 50% on a two-replica Deployment allows both Pods to terminate simultaneously during updates. Always pair surge-based updates with readiness probes and preStop hooks.
Service Discovery and Endpoint Consistency
Services provide stable DNS names (log-collector.log-forge.svc.cluster.local) backed by dynamic EndpointSlices. The Endpoints controller watches Pod labels and Service selectors β when they diverge, the Service exists but routes to zero backends. This is the most common networking failure in production clusters and the centerpiece of the incident simulation namespace.
The platform demonstrates three Service types for api-gateway:
ClusterIP (default): internal-only, resolved via CoreDNS
NodePort (30080): host-accessible for local development on kind
LoadBalancer: cloud-provider integration with session affinity for stateful client connections
NetworkPolicies add egress and ingress rules on top of Service discovery. log-collector-policy permits traffic only from api-gateway on port 8080 and egress only to kafka and redis plus DNS (UDP 53). The anti-pattern is deploying NetworkPolicies without understanding that default-allow becomes default-deny once any policy selects a Pod.
Configuration Management: ConfigMaps and Secrets
Kubernetes separates configuration (ConfigMaps) from credentials (Secrets) at the API level, though both are etcd-stored and require RBAC for access control. The platform uses three injection patterns:
envFrom.configMapRefβ bulk environment variable injectionconfigMapKeyRefβ selective key mappingVolume mounts with
itemsβ file-based configuration (config.yaml)
Secrets use secretKeyRef with RBAC resourceNames restrictions β the log-processor ServiceAccount can read database-credentials but not unrelated secrets. Kustomize overlays patch ConfigMaps for development (LOG_LEVEL: DEBUG) versus production (LOG_LEVEL: WARN) without duplicating base manifests.
Key insight: ConfigMaps are not hot-reloaded by default. Changing a mounted ConfigMap updates files on disk, but applications must watch for changes or use operators like Reloader.
Health Probes and Graceful Lifecycle Management
The three-probe model separates concerns that a single /health endpoint conflates:
Startup probe: "Has initialization completed?" β protects slow-starting containers from liveness kills
Readiness probe: "Can this instance accept traffic?" β checks dependencies (Kafka, Redis)
Liveness probe: "Is the process fundamentally healthy?" β checks internal state only
log-processor startup probe allows 120 seconds for ML model loading (failureThreshold: 12 Γ periodSeconds: 10). Without it, the liveness probe would restart the container during legitimate initialization.
Lifecycle hooks complete the graceful shutdown story. preStop on log-collector calls /shutdown to flush Kafka buffers before SIGTERM. terminationGracePeriodSeconds: 60 gives the hook time to complete. The anti-pattern is omitting preStop β Kubernetes sends SIGTERM immediately, dropping in-flight requests.
Operations Walkthrough
Cluster bootstrap β from k8-forge-platform/, run scripts/setup-cluster.sh to provision the kind cluster with NodePort mappings, then scripts/build.sh and scripts/deploy.sh to build images and apply manifests via Kustomize.
Cluster foundation β kubectl get all,hpa,pdb -n log-forge reveals the complete control plane surface area for the namespace.
Workload runtime β kubectl describe pod -l app=log-processor shows init container db-migration completing before the main container starts.
Deployment controller β kubectl set image deployment/log-collector ... followed by kubectl rollout status demonstrates availability-preserving rolling updates.
Service networking β curl localhost:30080/api/summary via NodePort; kubectl get endpoints validates selector alignment.
Incident simulation β apply broken manifests to log-forge-debug and use scripts/diagnose.sh following the debug hierarchy documented in debug/README.md.
Configuration management β kubectl apply -k k8s/overlays/development patches ConfigMap log levels without touching base manifests.
Health and lifecycle β tests/test_health_probes.sh confirms three-probe configuration on all services.
Architectural decision: single namespace (log-forge) with Kustomize overlays rather than fragmented projects β mirrors how platform teams manage environment divergence without manifest duplication.
Production Considerations
Resource requests and limits on every container prevent noisy-neighbor scheduling failures and enable HPA metrics
Init containers gate application startup on infrastructure readiness (database schema) without embedding migration logic in application images
PDB plus rolling update strategy is the minimum bar for production Deployment configuration
NetworkPolicy DNS egress (UDP 53) is required β policies blocking DNS cause mysterious connection failures to Services that appear healthy
Production Context
These patterns appear identically at large-scale deployments. Spotify's microservices run three-probe health models enforced by deployment gates. Netflix's Zuul-to-Kubernetes migration relied on readiness-gated rollouts to maintain 99.99% availability during container orchestration adoption. Airbnb's current platform requires NetworkPolicy on every namespace before production promotion. K8 Forge compresses these requirements into a single deployable system for internal validation and operations reference.