Lesson 1: Setting Up the Infrastructure

Lesson 4

1. Introduction

Task scheduling is one of those “boring” infrastructure concerns that quietly determines whether a backend behaves like a system or like a pile of endpoints. Real services don’t only react to inbound HTTP. They poll upstream systems, rotate files, refresh caches, emit periodic heartbeats, and enforce time-based policies. When those time-based behaviors are unreliable, everything else degrades: logs pile up, caches go stale, disk fills, and operators lose observability exactly when they need it most.

This integrated Week 1 project demonstrates scheduling as a concrete, inspectable part of a local log processing pipeline. It starts with simple time-driven behaviors (generation rate accounting and periodic stats), then connects them to more “operational” concerns like log parsing and file rotation. In code terms, this system is intentionally small, but it uses the same primitives you’d use in production Spring services: @Scheduled methods, a shared scheduler (ThreadPoolTaskScheduler), and a lightweight registry that tracks task state across runs.

This is written for a mid-level engineer who already knows Spring Boot basics and wants to treat scheduling as a first-class engineering concern. You’ll learn how scheduling interacts with thread pools and failure handling, how to keep periodic jobs observable, and how to integrate scheduled work into a pipeline without turning it into a ball of side effects.

2. From Fundamentals to a Unified System

Day 1 — Environment + a stable place to “run the system”

Day 1 is about getting a repeatable development environment and repository structure so infrastructure can actually be exercised. In this integrated app, that shows up as the ability to run “the system” with one entry point (com.example.week1.Week1IntegratedApplication) and a predictable set of endpoints (see com.example.week1.common.web.HomeController and SystemOverviewController). The engineering concern is reproducibility: scheduled systems are hard to debug when you can’t reliably start, stop, and observe them.

Day 2 — Configurable log generator and time-driven behavior

Day 2 introduces a log generator that can produce sample events at a configurable rate. In the integrated project, the generator runs in com.example.week1.day2.service.Day2LogGenerationService, which uses an internal thread pool for generation and a scheduled method for per-second rate calculations:

  • Day2LogGenerationService.updateMetrics() is @Scheduled(fixedRate = 1000) and also annotated with @TrackedTask(id = "day2.updateMetrics", ...).

This maps directly to a common production pattern: time-window metrics (or control loops) that run regardless of inbound traffic. You can expose a generator as an HTTP API, but its behavior is only correct if the periodic bookkeeping is correct and thread-safe.

Day 3 — File-based collector and periodic reporting

Day 3 adds a log collector that reads local files, maintains offsets, and de-duplicates entries. In the integrated app, com.example.week1.day3.service.Day3LogCollectorService watches directories and periodically emits statistics:

  • Day3LogCollectorService.logStatistics() is @Scheduled(fixedRate = 60000) and @TrackedTask(id = "day3.logStatistics", ...).

The engineering concern here is “background work as a pipeline.” Watching a filesystem isn’t request/response; it’s event-driven plus periodic health/reporting. When you’re building infrastructure, the first step toward reliability is making the background work visible.

Day 4 — Parsing as a service, not a regex in the hot path

Day 4 implements parsing to extract structured data from common log formats. In the integrated project, parsing lives in com.example.week1.day4.service.Day4LogParsingService, and the raw-to-structured transformation is represented by com.example.week1.day4.service.Day4RawLogProcessor.

The scheduling tie-in is subtle but real: parsing is often invoked by scheduled or asynchronous pipelines (batching, periodic scans, backfills). Even when parsing is triggered by file changes or Kafka, the system still needs predictable timing around reporting and policy enforcement.

Day 5 — Storage, rotation policies, and operational “tick” jobs

Day 5 introduces a storage mechanism using flat files with rotation. In the integrated project, rotation is a scheduled policy evaluation:

  • com.example.week1.day5.service.Day5RotationPolicyService.evaluateRotationPolicies() runs on @Scheduled(fixedRate = 300000) and is annotated with @TrackedTask(id = "day5.rotateColdFiles", ...).

This is exactly what a backend does in production: continuously enforce time/size policies. It’s not “feature code,” but it’s the code that prevents outages. Rotation jobs need predictable scheduling, bounded runtime, and strong observability when they fail.

Day 6 — Querying and filtering as an operational interface

Day 6 adds query APIs that filter stored logs with cache behavior and circuit breakers (com.example.week1.day6.service.Day6LogQueryService). This day is about operator ergonomics: scheduled pipelines need query surfaces so you can validate that the system is behaving over time, not just at a single moment.

The engineering concern is feedback loops. Scheduled systems are prone to silent failure. Query endpoints (plus metrics) are how you close the loop.

Day 7 — Integration into a local pipeline

Day 7 is the integration day: connect generation, collection, parsing, and storage into a pipeline that can run locally. In this integrated project, the pipeline is mediated via a dispatch layer (com.example.week1.common.kafka.LogEventDispatchService) that can either publish to Kafka (when enabled) or persist directly (when Kafka is disabled).

The scheduling concern is orchestration without tight coupling. The system can run local-only and still exercise scheduled tasks like metrics rollups and rotation. When Kafka is enabled, the same scheduled tasks still work; they just become part of a larger asynchronous pipeline.

3. Architecture Overview

Component Architecture

Spring Boot Application Context com.example.week1.Week1IntegratedApplication Configuration layer com.example.week1.common.config.SchedulingConfig Defines shared ThreadPoolTaskScheduler com.example.week1.common.config.JacksonConfig com.example.week1.common.config.RedisConfig application.yml (properties for schedules, topics, etc.) Scheduling + tracking layer ThreadPoolTaskScheduler Worker threads: week1-scheduler-* TaskTrackingAspect TaskRegistry TaskState /api/week1/scheduler/tasks @Scheduled task beans (Week 1 days) Day 2 Day2LogGenerationService @Scheduled fixedRate=1000 @TrackedTask id=day2.updateMetrics Day 3 Day3LogCollectorService @Scheduled fixedRate=60000 @TrackedTask id=day3.logStatistics Day 5 Day5RotationPolicyService @Scheduled fixedRate=300000 @TrackedTask id=day5.rotateColdFiles JVM / OS clock Time source triggers schedules ticks configures invokes via pool tracks

At runtime, the application is a single Spring Boot process with a few key internal layers that matter specifically for scheduling:

The outer boundary is the Spring application context (Week1IntegratedApplication). Inside that context, configuration beans provide core infrastructure:

  • com.example.week1.common.config.SchedulingConfig defines a shared ThreadPoolTaskScheduler used by Spring’s scheduling subsystem.

  • com.example.week1.common.config.JacksonConfig and RedisConfig provide serialization and Redis clients that scheduled tasks use indirectly (e.g., rate limiter windows, file offset tracking).

Scheduled task beans live in per-day modules:

  • Day 2 metrics tick: Day2LogGenerationService.updateMetrics()

  • Day 3 periodic collector stats: Day3LogCollectorService.logStatistics()

  • Day 5 rotation policy evaluation: Day5RotationPolicyService.evaluateRotationPolicies()

To keep “background” behavior observable, the project adds a minimal task registry and tracking aspect:

  • com.example.week1.common.scheduling.TaskRegistry stores TaskState objects keyed by a stable task id.

  • com.example.week1.common.scheduling.TaskTrackingAspect wraps methods annotated with @TrackedTask and updates state on start/success/failure.

  • com.example.week1.common.web.TaskRegistryController exposes the registry at /api/week1/scheduler/tasks.

Finally, the external boundary is the JVM/OS clock, which triggers the scheduler’s timing wheel. In production you’d also model this boundary in terms of deployment topology and clock skew; locally it’s enough to treat it as “time ticks happen.”

4. Data / Control Flow

Flowchart

Control flow: startup -> dynamic scheduling -> trigger cycle -> execution results Application startup Initialize scheduler ThreadPoolTaskScheduler (SchedulingConfig) Register @Scheduled methods Day2.updateMetrics / Day3.logStatistics / Day5.evaluateRotationPolicies WAIT for next trigger Clock fires — scheduler selects a worker thread Invoke task method @TrackedTask via TaskTrackingAspect Task success? Log error (+ optional retry) TaskRegistry: FAILED -> WAITING Log success TaskRegistry: SUCCESS -> WAITING No Yes Next cycle

Even though the project ingests and stores logs, the scheduling lifecycle is the piece that makes the system “keep moving” without humans poking it.

When the application starts, Spring initializes beans, then scheduling infrastructure scans for @Scheduled methods. Those methods are scheduled onto the shared scheduler defined in SchedulingConfig and begin waiting for time-based triggers.

On each trigger cycle, the runtime flow looks like this:

Startup → scheduler initialized → @Scheduled methods registered → WAIT for time → clock fires → scheduler picks a worker thread → the scheduled method executes → it emits logs/metrics and may update storage → the method returns → the thread goes back to the pool → the scheduler waits for the next trigger.

The task tracking aspect turns that flow into observable state transitions. When a tracked scheduled method begins, it is marked RUNNING in the TaskRegistry. If the method returns normally it is marked SUCCESS; if an exception propagates it becomes FAILED. In both cases the registry returns the task to WAITING so the next trigger can proceed.

5. State Management

State Machine

Lifecycle Flow: Task Definition State Transitions REGISTERED WAITING RUNNING SUCCESS FAILED DISABLED scheduled registered -> waiting trigger fires returns next cycle throws retry / continue manual disable shutdown

A scheduled job isn’t just “a method that runs sometimes.” It has state that matters operationally: when it last ran, whether it is failing, and whether it is making progress. The integrated project tracks this state in-memory via TaskRegistry and TaskState:

  • TaskState.status records REGISTERED, WAITING, RUNNING, SUCCESS, FAILED, DISABLED.

  • TaskState.lastStartedAt and lastFinishedAt capture timing.

  • lastSuccessAt, lastFailureAt, and lastErrorMessage capture health.

  • totalRuns, totalSuccess, totalFailures capture a small history without needing a full timeseries DB.

The mechanism is intentionally small and transparent: methods annotated with @TrackedTask are intercepted by TaskTrackingAspect. This keeps the task beans themselves focused on their domain logic, while the cross-cutting concerns (timing, success/failure bookkeeping) live in one place.

You can view the current state at runtime via:

  • GET /api/week1/scheduler/tasks (implemented by TaskRegistryController)

This is not meant to replace Prometheus/Grafana; it’s a debugging surface that makes scheduling correctness visible immediately during local development.

7. Key Engineering Insights

Thread safety isn’t optional in scheduled systems because there’s no user request boundary to “hide” behind. Day2LogGenerationService uses AtomicLong/AtomicInteger for counters because the scheduled tick (updateMetrics) is racing with generator threads that increment counters. If those updates weren’t atomic you’d get negative or inconsistent rates, which is the kind of bug that looks like “the system is flaky” rather than “the math is wrong.”

Fixed-rate and fixed-delay scheduling are not interchangeable. This project uses fixedRate for periodic accounting and rotation evaluation. In a production service you’d choose fixedDelay when work duration matters more than wall-clock alignment (e.g., “run 5 minutes after the last successful completion”), and fixedRate when you want alignment to time windows (e.g., “every second, compute a rolling rate”). Even locally, those differences show up when tasks get slow or the system is under load.

Cron expressions are powerful but brittle when you need dynamic behavior. This week’s code uses @Scheduled with explicit intervals, which keeps behavior testable and easy to reason about. In later weeks, adding cron schedules (or Quartz) makes sense when you need calendar-aligned jobs, but you still need the same observability and state tracking; cron does not remove the need for task health.

Testability improves when scheduling is separated from work. The tracking layer in TaskTrackingAspect makes task state independent from the business logic. This means you can unit test the domain work (e.g., Day4LogParsingService.parseLogEntry) without time, and then separately exercise scheduling behavior via integration tests that assert /api/week1/scheduler/tasks transitions.

8. Success Metrics

Modularity

You should be able to point to scheduling-related code without digging through business logic. Concretely:

  • SchedulingConfig contains scheduler configuration only.

  • Task tracking lives only in common/scheduling/*.

  • Day modules (day2, day3, day5) keep scheduled logic in their own services.

Readability

A reader should be able to answer “what tasks are scheduled?” quickly:

  • The scheduled methods are visible via @Scheduled on:

  • Day2LogGenerationService.updateMetrics()

  • Day3LogCollectorService.logStatistics()

  • Day5RotationPolicyService.evaluateRotationPolicies()

  • Each of those methods has a stable id and description via @TrackedTask.

Correctness

Scheduling correctness isn’t “it compiles.” It’s observable behavior:

  • GET /api/week1/scheduler/tasks returns entries for day2.updateMetrics, day3.logStatistics, and day5.rotateColdFiles.

  • Each task transitions through RUNNING → SUCCESS repeatedly under normal operation.

  • When a task fails, lastErrorMessage is populated and totalFailures increases.

Extensibility

Adding a new scheduled task should not require copy/pasting tracking code:

  • A new task is created by adding @Scheduled + @TrackedTask to a method on any bean.

  • The registry and endpoint automatically reflect the new task id.

  • The scheduler pool size and behavior can be tuned in one place (SchedulingConfig).

9. Conclusion

After building and running this integrated project, you should see scheduling as a system concern rather than an annotation. The core idea is that time-based work needs the same discipline as request-driven code: bounded concurrency (ThreadPoolTaskScheduler), clear ownership (per-day modules), and observability (task state plus metrics).

From here, the natural next step is taking scheduling beyond a single JVM. Once you have multiple instances, “every instance runs the cron” becomes a correctness bug. That’s where distributed scheduling patterns come in: leader election, database-backed locks, Quartz clustered mode, or dedicated schedulers. The mechanics change, but the fundamentals you practiced here—explicit task identity, clear state, and careful concurrency—carry straight through.

Questions & Discussion

Leave a Reply

Your email address will not be published. Required fields are marked *

System Design Fundamentals – E-Book

Free download

Free eBook: System Design Fundamentals

Create a free account and download the ebook instantly. Learn the core building blocks — scaling, caching, databases and messaging — the way interviewers expect you to explain them.

Register free & download →

Already a member? Sign in to download · See what’s inside