Security

FactumDB

MySQL Database Forensics Software

Abstract

FactumDB correlates MySQL InnoDB tablespace files with row-based binary logs to reconstruct record histories and transaction timelines that today have to be assembled by hand. It handles evidence read-only, hashes and verifies every file, links log events to physical records, and reconciles the two; reporting evidence gaps and conflicts explicitly instead of hiding them. Every finding is traceable to its raw source. For forensic investigators, incident responders, and auditors.

The problem

MySQL investigators often work from two disconnected artefacts: .ibd files showing a table's current physical state, and binary logs showing the row-level changes that led there. Existing utilities can read each one, but correlating them, mapping positional binlog fields to columns, tracing transaction boundaries, matching records, replaying history, comparing it to the physical rows is done by hand. It's slow, error-prone, and evidence gaps like a purged log tend to fail silently, making an incomplete timeline look complete.

The solution

FactumDB automates that correlation. It hashes and verifies evidence read-only, then runs the standard extraction utilities (ibd2sdi, innochecksum, ibd2sql, mysqlbinlog) through an orchestrated pipeline with full audit logging. A domain engine groups binlog events into transactions, links them to physical records by key, reconstructs record histories, and reconciles the reconstructed state against the .ibd file, returning a rule-based classification (Exact, Strong, Partial, Conflicting, Unresolved, Unsupported) with full provenance for every finding. Built as a Tauri/React desktop app on a Hexagonal Architecture, with a Python sidecar handling extraction and correlation.

In detail

FactumDB reconstructs MySQL activity from acquired evidence InnoDB tablespace files and binary logs, and presents it as a traceable, examiner-reviewable history, rather than raw utility output the analyst has to correlate by hand.

A case runs evidence through hashing and verified working copies, page-integrity checks, schema and row extraction, and binlog decoding. From there the domain engine groups events into transactions using GTIDs and commit markers, links them to physical records by primary key, replays them chronologically, and reconciles the result against the tablespace. Every finding; a transaction, a record history, a reconciliation result, traces back to its source file, log position, and raw output.

The harder problem was the failure cases. A missing binlog file doesn't produce a false mismatch; it produces an explicit coverage gap and an Unresolved classification, because the evidence genuinely can't tell you whether the difference is tampering or a gap. That same discipline shows up in the UI: the transaction timeline is ordered by binlog position, not wall-clock time, since binlog timestamps are commit-time and second-granularity, and coverage gaps render as visible bands rather than disappearing. Everything, the layout, correlation rules is deterministic, since repeatability across runs on the same evidence is a hard requirement.

Evaluated against synthetic MySQL 8.4.x datasets with known ground truth: single and multi-table transactions, composite keys, missing logs, corrupted pages, mismatched snapshot timing, and interleaved concurrent sessions.

Scope for v1: one MySQL 8.4.x release, InnoDB, file-per-table, offline analysis only. Out of scope: live acquisition, redo/undo log parsing, and any attribution of changes to a person or intent, FactumDB shows what the evidence supports, not who did it or why.