Engineering

Designing a Modern Storage Engine from the Ground Up

Pages, blocks, packs, WAL, MVCC, and checksums: the physical organization underneath PLOMID, and why the layout is a contract every data model shares.

Sainath Sapa · Founder & CEO 4 min read
Fixed-size pages grouping into blocks and packs in the PLOMID storage hierarchy
The physical hierarchy: 16 KiB pages pack into 256 KiB blocks, blocks into segments, segments into packs.
On this page
  1. The layout, in numbers
  2. Write-ahead first, pages second
  3. Snapshots, not locks
  4. Checksums everywhere it matters
  5. Seeing it on disk

A storage engine is a promise about two things: your data survives a crash, and it comes back in a useful order. Everything else — page sizes, checksums, write-ahead logs — is the machinery that keeps the promise. This note walks through PLOMID’s physical organization, because the layout is the one contract every data model shares.

The layout, in numbers

The hierarchy is fixed and deliberately boring:

Page, block, pack, and extent layout

  • Page, 16 KiB — the unit of reads and writes: a 48-byte header, the payload, and a 4-byte trailer.
  • Block, 256 KiB — 16 pages, the unit the buffer pool reasons about.
  • Extent, 64 MiB — the allocation granularity in the middle.
  • Pack, 16 GiB target — the large container; on disk, segment files live at devices/D-<id>/packs/SEG-<id>.dat.
  • Columnar chunk, 16 MiB — the unit for encoded, scan-friendly flushes.

Fixed sizes are a feature, not a limitation. Every offset is computable, every boundary is checkable, and recovery never has to guess where a structure starts. There is no manifest file to lose: the device directory listing plus a small registry is the durable inventory.

Write-ahead first, pages second

Nothing reaches a page before it reaches the log. Each record gets a log sequence number from 1 with no gaps; appends rotate through 8 MiB segments without ever splitting a record or rotating mid-transaction. A commit is a flush followed by an fsync — the durable watermark — and only then do pages absorb the change.

Checkpoints bound how much log a restart must replay. They move through a strict lifecycle — build, flush, verify, sync, publish, with an atomic rename at the end — and a background policy (64 MiB of WAL bytes, 16 segments, or 15 minutes, whichever comes first) decides when one is due. Crash recovery is therefore two steps: validate the root and the log contiguity, then replay from the last published checkpoint. The storage documentation covers the knobs; the shape above is what the knobs tune.

Snapshots, not locks

Concurrent readers and writers meet through multiversion concurrency control. Writers stage new row versions; readers hold a snapshot — a watermark plus the set of transactions active when the read began — and see exactly the versions committed before it. Readers never block writers, and writers never disturb an in-flight read.

Old versions are reclaimed against a GC horizon derived from the oldest active snapshot: anything newer than the horizon stays, plus the newest version at or below it as the floor. The isolation level this yields is read committed, which pairs with the transaction semantics in SQL and the Foundation of PLOMID.

Checksums everywhere it matters

Every stored structure carries a CRC checksum — pages, segments, checkpoints, filter blocks. A failed verification surfaces as an explicit corruption error, never as silent wrong data. That sounds obvious, but it is a design decision with teeth: the engine would rather refuse to answer than answer from bytes it cannot vouch for.

In Rust, the page header is the kind of small, exact structure the language is good at:

/// Fixed 48-byte header prefixing every 16 KiB page.
/// Layout is versioned explicitly: old pages must remain readable.
#[repr(C)]
pub struct PageHeader {
    magic: u32,        // identifies the page kind
    version: u16,      // layout version, checked on every read
    flags: u16,        // checksum-present, compressed, …
    page_id: u64,      // position-independent identity
    lsn: u64,          // last log record applied to this page
    checksum: u32,     // crc32c over header + payload
    reserved: [u8; 24],
}

#[repr(C)] pins the layout so the bytes on disk match the struct in memory; the version field means a future layout can still read today’s pages. Nothing here is clever — checksums, versions, and fixed sizes are old ideas, applied consistently.

Seeing it on disk

Starting a node and looking at the data directory shows the layout with no abstraction in the way:

plomid-server --data ./data &
psql -h localhost -p 5432 -U plomid -c "CREATE TABLE t (id int PRIMARY KEY);"
ls ./data/devices/*/packs/ | head
# SEG-000000000001.dat
# SEG-000000000002.dat

The table just created lives in those segment files, addressed through the page → block → pack hierarchy, guarded by the WAL. The same files will one day hold documents, vectors, and graph edges through the same contract — which is why the layout chapter comes before the multi-model direction.

Durability is not a feature to list. It is the floor everything else stands on, poured first.