Available for contract or full-time work

AI Engineer — verifiable LLM systems

I build LLM pipelines on Claude, Gemini and OpenAI — RAG and vector search, structured output, evals, and human-in-the-loop review — for finance, legal, healthcare, insurance and government documents.

The model proposes. Tested code decides. I ship the harness that proves it, not just the feature.

Remote US · US citizen, no sponsorship needed Python · TypeScript · C#/.NET LinkedIn ↗ GitHub ↗ Resume
tie-out checks · 0 exceptions on a published report — inject a wrong figure and it is caught
25
production AI and document tools shipped for a CPA firm, used daily by non-technical staff
5
verification gates that always run in AgentA — three more arm when a task configures them
4
live B2B SaaS I build and operate (OBBBA Tracker)
1

What I do

Three things, done to an audit standard

A finance background, an engineer's hands, and a hard rule: tested code proves the answer — the model never gets the last word on a number.

Document extraction & validation

PDFs into clean, structured JSON — then every total re-derived (foot, crossfoot, articulate) and locked with golden-file tests. No OCR for born-digital text.

AI systems that stay honest

Multi-agent systems, RAG, and LLM pipelines built behind verification — sealed holdouts and fresh-context critics — so improvements are earned, not hallucinated.

Full-stack SaaS

Production B2B products end to end — auth, multi-tenant data, billing, payroll-format exports, audit trails, and dashboards. Real users, real subscriptions.

Selected work

Products and systems I built and shipped

Live SaaS, document intelligence, autonomous AI, and a strength-per-byte a legal RAG system, and an ML benchmark. Read the full write-ups, including a native C++ Pac-Man →

OBBBA Tracker dashboard showing tipped-wage and overtime deduction tracking
LiveB2B SaaS · I build & operate it

OBBBA Tracker — live tax-compliance SaaS

Helps tipped-industry employers track and document the “no tax on tips & overtime” deductions under the One Big Beautiful Bill Act: automatic Treasury Tipped Occupation Code assignment, FLSA overtime, W-2 Box 14 exports (ADP/Gusto/QuickBooks), an audit trail, multi-tenant role-based access, and an analytics dashboard. Subscriptions plus a free trial.

full-stack SaaSauthmulti-tenantpayroll exportsbilling
Visit the live product
Multi-agent system

AgentA — self-improving engineering harness

An autonomous Claude Code system — specialised subagents, skills and hooks — that improves its own code behind a verification gauntlet. Four gates always run: parse, unit tests, a benchmark delta, and a fresh-context critic that judges the diff with no memory of how it was produced. Three more (property tests, mutation testing, a sealed holdout) arm only when a task configures them.

A skipped gate is logged as skipped and never counted as a pass

Inner research loop ported from Udit Goenka’s autoresearch (MIT, after Karpathy); the meta-improvement architecture and verification stack are my own.

multi-agentverificationPython
PDF → validated structured data

Financial-Statement Extractor

Turns a government audit statement (PDF) into clean structured JSON, then re-derives every total to prove the extraction is correct, with a golden-file regression test. The model never does the arithmetic — tested code does.

25 checks, 0 exceptions — inject one wrong figure and it’s caught

Pythonpdfplumberpytestaudit-grade
Federal grant compliance SaaS · working prototype

GrantLedger

Full-stack B2B SaaS that auto-categorizes nonprofit grant spending into 2 CFR 200 budget categories, tracks budget-to-actual per grant, and generates audit-ready compliance reports, with QuickBooks/Xero integration.

Built and running locally, not yet deployed — source is public

Next.jsTypeScriptSupabaseStripeOpenAIaudit-ready
RAG · client work, source private

Legal RAG with page-level citations

A retrieval system that lets attorneys query a legal-text repository and get answers carrying precise page-level citations, so every claim can be checked against the page it came from. Retrieval plus LLM synthesis, built and deployed at a law firm and used in live case work.

An answer without a citation to a real page is not an answer

Built for a client, so the corpus and the code stay private. Happy to walk through the retrieval and citation design in a conversation.

RAGvector searchcitationsPython
ML benchmark · search vs. learning

NeuroFour — Connect 4 strength-per-byte

Connect 4 is a solved game, so perfect play is a known, fixed target. NeuroFour ranks 20 Connect-4 agents by strength per byte, scored against an exact solver. Nineteen fit the hard 5M-FLOP-per-move budget and are eligible for the headline; the twentieth is that solver itself, which needs 50M FLOPs and is listed but over budget.

Zero leads optimality, ladder Elo, and NeuroFour Score among the 19 under-budget agents, at 0 bytes, ahead of all 14 learned-net entries (6 distinct weight files, largest ~24 KB): ladder Elo 754 vs 631 and NeuroFour Score 96.45 vs 74.25 — a score that divides strength by a size penalty — optimality 0.960 vs 0.957 by a single move in 300. Each “vs” is the strongest learned entry on that axis (optimality is a tie); the four entries behind them load just two weight files. The cost: Zero spends 4,999,028 of the board’s 5M-FLOP budget — 99.98% of it, 18th of 19 on FLOPs/move (and 4th on soundness, 15th on latency).

Free-tier API — first move may take ~30–60s to wake.

benchmark designPython / FastAPIReact 19

About

Finance brain, engineer’s hands

I’m an AI & document-automation engineer with a finance background. I build document-extraction pipelines, AI systems, and full-stack SaaS — with a focus on output that is verifiable, not just plausible.

Over the past year I designed and shipped five production document-automation tools for a public-accounting firm (Millhuff-Stang, CPA), built a live tax-compliance SaaS (OBBBA Tracker), and built a self-improving multi-agent engineering harness. Earlier I built a Retrieval-Augmented Generation (RAG) document system for a law firm and a case-management web application for a legal practice.

More about me, the full stack & credentials

PythonC#/.NETTypeScript / React FastAPIpdfplumber / PyMuPDFRAG / vector search Claude / Gemini / OpenAINext.jsDocker
Education
B.B.A., Finance — University of Cincinnati, Lindner College of Business (Cum Laude)
Certifications
CFI — Financial Modeling & Valuation Analyst (FMVA); Business Intelligence & Data Analyst (BIDA)
Based in
Ohio, USA — working remote
Open to
Contract or full-time AI / automation / document-intelligence work

Writing & research

Notes on building systems you can trust

Short, sanitized pieces on verifiable AI. Read them all →

Contact

Let’s build something verifiable

I’m available for contract or full-time AI, automation, and document-intelligence work. The fastest way to reach me is email — no forms, no backend, just say hello.