Day 3: Wrap the Storage Engine in a Maven Build Lifecycle and Automate Testing with JUnit

Lesson 3 60 min

Day 3: Automating the Build and Validating State

In Day 2, you built a custom chained hash map. It is a beautiful, memory-resident structure, but it exists in a vacuum. To turn that code into a system, you must move from "writing scripts" to "engineering software." Today, we wrap your storage engine in the Maven build lifecycle and use JUnit to codify its behavior. This is not just about organizing files; it is about establishing a "test harness" that acts as the final arbiter of truth before your code touches disk.

The Problem: The "Works on My Machine" Tax

State Machine

Day 3 β€” Hash Map State Machine Storage engine state transitions and validation INITIAL Empty Map COLLISION Same Bucket ACTIVE Data Present GET Read Value UPDATED State Changed put(key,value) hash collision traverse chain get(key) return value update state remains valid INVARIANTS size() = number of keys get(key) = correct value Collisions preserve entries JUnit validates every important state transition

Flowchart

Day 3 β€” Maven / JUnit Test Flow Start Write / Update Code HashMap + JUnit tests Maven Compile Build lifecycle Run JUnit Tests Assertions + invariants All Tests Pass? Assertions valid BUILD SUCCESS Safe to continue BUILD FAILED Fix the issue YES NO Deterministic verification prevents regressions

Component Architecture

Day 3 β€” Architecture Components Source Code Java / Engine Maven Build Lifecycle Bytecode target/ AstraEngine Hash Map Logic JUnit Test Runner Assertions Verify Invariants compile package test validate Source β†’ Maven β†’ Bytecode β†’ JUnit β†’ Assertions

In hyperscale environments, the most dangerous code is the code that is "probably correct." When you are building a storage engine, you are dealing with pointer arithmetic, bucket collisions, and memory offsets. If you rely on manual main() method print-debugging, you are one edge case away from a silent data corruption bug that might not manifest for months.

Consider the Amazon S3 "Ghost Writes" or the infamous Chubby lock-service outages. These systems didn't fail because the core logic was complex; they failed because of subtle state transitions that occurred under specific concurrency conditions. If you don't have an automated way to assert that your hash map behaves exactly as expectedβ€”even when keys collide or memory pressure spikesβ€”you aren't building a system; you are building a ticking time bomb.

The Concept: Deterministic Verification

Your hash map must be verifiable. We will use Maven to manage our dependencies and enforce a build lifecycle. We will use JUnit to create a suite of deterministic tests. A deterministic test is one where, given the same input, the outcome is identical every time, regardless of the environment.

The intuition here is the "Contract-First" approach. Before you write the logic for disk persistence tomorrow, you must define the "contract" of your storage engine today. If the engine cannot pass these tests, it is technically broken.

Implementation: Defining the Invariants

You are moving from a single file to a structured project. Your pom.xml will now manage your dependencies, and your tests will verify the core invariants of your hash map.

java
// src/test/java/com/astrakv/core/HashEngineTest.java
@Test
public void testPutAndGetConsistency() {
    AstraEngine engine = new AstraEngine();
    engine.put("key1", "value1");
    // Assert that the state matches expectation exactly
    assertEquals("value1", engine.get("key1"), "Read operation failed to return written value");
}

By enforcing assertEquals, you are creating a guardrail. If you change your hashing algorithm tomorrow and it causes a collision, this test will fail immediately.

The Failure Demo: The "Negative Test"

Today, you will write a "negative test." You will intentionally try to insert a null key or exceed the capacity of your map. A robust system doesn't just work; it fails gracefully. If your system allows a NullPointerException to bubble up to the client, you have failed as an engineer. Your tests should prove that your system throws a specific, documented StorageException instead.

Production Stakes: The Cost of Manual Verification

In a distributed system, manual testing is a luxury you cannot afford. When you are deploying to 500 nodes across three regions, you need a "continuous integration" pipeline that runs these tests on every commit. If you lack this, you will eventually deploy a bug that causes a "Hotspot" (where one node takes all the traffic because its hash-map implementation is slightly skewed), leading to cascading failures across your cluster.

Assignment

  1. Add a new test method testCollisionHandling to your HashEngineTest.java.

  2. Insert 100 keys that are guaranteed to hash to the same bucket (you will need to use a custom Key object with a fixed hashCode).

  3. Verify that the size() method returns 100.

  4. Run mvn test and ensure the build passes.

Solution Hint:
To force a collision, create a class CollisionKey that overrides hashCode() to return a constant (e.g., 42). When you put 100 entries into your map, they will all chain into the same bucket. If your get() method logic is correct, the map will successfully traverse the chain to find the correct value. If it is broken, it will likely return only the last inserted value or throw an exception.

Questions & Discussion

Leave a Reply

Your email address will not be published. Required fields are marked *

System Design Fundamentals – E-Book

Free download

Free eBook: System Design Fundamentals

Create a free account and download the ebook instantly. Learn the core building blocks β€” scaling, caching, databases and messaging β€” the way interviewers expect you to explain them.

Register free & download β†’

Already a member? Sign in to download Β· See what’s inside