Data Platform Architect · Analytics Engineer
I build the layer that makes a data platform answer questions in its own language — semantic models, validated metric definitions, and self-documenting warehouses. Underneath that, I ship end-to-end dbt pipelines and the agent systems that work reliably around them. Daily, that means reconciling messy commercial data until the numbers stop lying.
The hardest problem in analytics isn't storage or compute. It's that nobody can say what the numbers mean — so every question routes through one person, and every metric is re-derived from scratch.
I built a semantic layer over a production platform: 118 tables and
65 pipelines distilled into a 35-concept graph (70 typed relations)
and a 135-metric business tree, then compiled into a single offline HTML file that
opens from file:// with no server and no runtime dependencies.
The full write-up: the failure modes, the layer design, and what the reconciliation surfaced. Sanitized — no client, table, or metric names.
The governing rule: the machine may propose; the domain owner disposes. Every AI-authored artifact is a proposal for review, validated against source, and an owner can reject it. Of 135 proposed metrics, 128 were approved, 6 failed validation, and 1 was rejected by its owner — a metric tree with a 100% approval rate would mean nobody checked.
Two public pipelines, both live, both with CI enforcing every test.
Four keyless data sources → a Kimball star schema across 23 models
(staging → intermediate → marts), with SCD2 snapshots,
model contracts, custom macros, and 99 passing tests in CI. Serves an
interactive Evidence.dev dashboard where an estate manager decides the day's operations.
Live dashboard — weather + FX + holidays + commodity price → dbt → decision page.
Marketing measurement where ad platforms over-claim conversions by hundreds of percent. Attribution triangulation: Multi-Touch Attribution + Media Mix Modeling (adstock & Hill saturation) + incrementality testing (geo-lift, holdouts), reconciled into one tested semantic layer. 15 models, 19 passing tests.
Platform claims 55,947 conversions; the causal estimate is 7,899 — a 608% bias, measured from the model rather than trusted from the ad network.
Most real analytics work is not greenfield. It's three systems that disagree, and a number somebody trusts.
Replaced a team's manual reconciliation with a pipeline that unified three separate systems, matching 18,537 orders (98.5%). Doing so exposed defects the manual process had been silently absorbing:
| Defect surfaced | Scale | Disposition |
|---|---|---|
| Records double-counted across two systems | 12 records, 0.06% overstatement | fixed |
| Partially-paid records counted at full value | 2.1% overstatement | fixed (owner rule) |
| Cancelled records with payment retained | 87 records | escalated |
| Shipping counted as revenue | 926 records | escalated |
| A "trusted" headline figure that moved 92% in one day | flagship metric | escalated, methodology documented |
Each defect ships with a runnable verification query, so any claim can be re-tested instead of taken on trust. The point of the migration wasn't saving hours — it was that the numbers became verifiable.
Research that turned into working systems, with published benchmarks.
An agent sees a page as a token-budgeted accessibility tree with [#N] action
indices, and acts through verified actions. Built on it, ARC Index extracts
each field from a document with an evidence quote and citation, and rejects any value
not proven at the cited row — then batch-fills the form and returns an audit receipt.
| Measured on real cloud browsers | Baseline | This work |
|---|---|---|
| Task success (same model, head-to-head) | 21 / 24 | 24 / 24 |
| Input tokens, 24 runs | 458,546 | 122,034 (3.8× fewer) |
| Input tokens per document appeal | 26,955 | 1,209 (22× fewer) |
| Cost per appeal | $0.02575 | $0.00305 (8.4× cheaper) |
| Wall time per appeal | 89 s | 11.4 s (7.8× faster) |
| Wrong values let through (of 60 fields) | 0 | 0 — equal accuracy |
Raw per-run JSON committed; measurement methodology documented in-repo. Same model on both sides.
A field CRM app for palm-oil enumerators — Streamlit + Bitable API, 53 passing pytest tests. Building the tools the team actually uses is part of the job.
Not autopilot — I build the contracts that make it safe, and the reasoning lives in the repo.
CLAUDE.md /
AGENTS.md) specifying commands, authority, and production risk — e.g. which target
writes to real production, and which source wins when two disagree.The through-line: ungrounded data shouldn't be trusted. In analytics that's a metric with no source; in agent systems it's a form field with no evidence. Both projects above enforce the same rule — propose, validate, let a human dispose.
One method, applied three times: learn by building, let real users correct you, then erase the bottleneck you just found.
I started in industrial engineering, where the job was never to theorize — it was to build something, watch it fail under real conditions, and fix the flow. That's how I moved into data analytics about six years ago, and it's the same reason I'm moving into AI now.
My first AI-assisted build was a team donation app. The problem was ordinary and I'd seen it repeat: contributions announced in a chat group, transfer screenshots scattered across private threads, one person manually reconciling the mess before every deadline. The transparency was fine — what was missing was a clear flow. I volunteered to build it, shipped it, then ran it against real colleagues: ask for feedback, watch where the flow breaks, ship again.
That loop taught me more than any course, and it also exposed my own gap — I was copying code into a chat and pasting it back without understanding Git or CI/CD. So I rebuilt the loop around agentic coding, which let me read a pipeline instead of guessing at it. The projects above are the result of that shift.
Supabase + Cloudflare Pages · donor sign-up, equal-split or custom amounts, combined payments settling several obligations in one transfer, receipt upload, PIC verification, refunds on overshoot · 129 passing tests.