Document extraction & validation
PDFs into clean, structured JSON — then every total re-derived (foot, crossfoot, articulate) and locked with golden-file tests. No OCR for born-digital text.
I turn messy PDFs and documents into clean, validated, structured data — and build the AI systems and SaaS products around them.
I ship real products — then make the results provably correct: tested code does the math, not the model.
What I do
A finance background, an engineer's hands, and a hard rule: tested code proves the answer — the model never gets the last word on a number.
PDFs into clean, structured JSON — then every total re-derived (foot, crossfoot, articulate) and locked with golden-file tests. No OCR for born-digital text.
Multi-agent systems, RAG, and LLM pipelines built behind verification — sealed holdouts and fresh-context critics — so improvements are earned, not hallucinated.
Production B2B products end to end — auth, multi-tenant data, billing, payroll-format exports, audit trails, and dashboards. Real users, real subscriptions.
Selected work
Live SaaS, document intelligence, autonomous AI, and a strength-per-byte ML benchmark. See all six, including a native C++ Pac-Man →
Helps tipped-industry employers track and document the “no tax on tips & overtime” deductions under the One Big Beautiful Bill Act: automatic Treasury Tipped Occupation Code assignment, FLSA overtime, W-2 Box 14 exports (ADP/Gusto/QuickBooks), an audit trail, multi-tenant role-based access, and an analytics dashboard. Subscriptions plus a free trial.
Visit the live productAn autonomous system that improves its own code behind a 7-tier verification gauntlet — parse, unit, property, mutation, benchmark, sealed holdout, fresh-context critic — so improvements are earned, not hallucinated.
Improvements gated by a sealed holdout + fresh-context critic
Inner research loop ported from Udit Goenka’s autoresearch (MIT, after Karpathy); the meta-improvement architecture and verification stack are my own.
Turns a government audit statement (PDF) into clean structured JSON, then re-derives every total to prove the extraction is correct, with a golden-file regression test. The model never does the arithmetic — tested code does.
25 checks, 0 exceptions — inject one wrong figure and it’s caught
Full-stack B2B SaaS that auto-categorizes nonprofit grant spending into 2 CFR 200 budget categories, tracks budget-to-actual per grant, and generates audit-ready compliance reports, with QuickBooks/Xero integration.
Audit-ready compliance reports, budget-to-actual per grant
Connect 4 is a solved game, so perfect play is a known, fixed target. NeuroFour ranks 19 Connect-4 agents by strength per byte under a hard 5M-FLOP-per-move budget, scored against an exact solver.
Zero leads optimality, ladder Elo, and NeuroFour Score among the 19 under-budget agents, at 0 bytes, ahead of all 14 learned-net entries (6 distinct weight files, largest ~24 KB): ladder Elo 754 vs 631 and NeuroFour Score 96.45 vs 74.25 — a score that divides strength by a size penalty — optimality 0.960 vs 0.957 by a single move in 300. Each “vs” is the strongest learned entry on that axis (optimality is a tie); the four entries behind them load just two weight files. The cost: Zero spends 4,999,028 of the board’s 5M-FLOP budget — 99.98% of it, 18th of 19 on FLOPs/move (and 4th on soundness, 15th on latency).
Free-tier API — first move may take ~30–60s to wake.
About
I’m an AI & document-automation engineer with a finance background. I build document-extraction pipelines, AI systems, and full-stack SaaS — with a focus on output that is verifiable, not just plausible.
Over the past year I designed and shipped five production document-automation tools for a public-accounting firm (Millhuff-Stang, CPA), built a live tax-compliance SaaS (OBBBA Tracker), and built a self-improving multi-agent engineering harness. Earlier I built a Retrieval-Augmented Generation (RAG) document system for a law firm and a case-management web application for a legal practice.
More about me, the full stack & credentials
Writing & research
Short, sanitized pieces on verifiable AI. Read them all →
Why deterministic extraction + tie-out validation + golden-file tests beat trusting an LLM with numbers.
02How AgentA keeps an autonomous improvement loop honest with a sealed holdout and a fresh-context critic.
03Shipping a real tax-compliance SaaS: data model, payroll-format exports, and keeping a compliance product trustworthy.
04Building a strength-per-byte Connect-4 benchmark and honoring the answer it returned, not the one I wanted.
Contact
I’m available for contract or full-time AI, automation, and document-intelligence work. The fastest way to reach me is email — no forms, no backend, just say hello.