Production AI at scale
Half a million dollars in cumulative revenue across co-founded AI ventures, with products adopted by 30+ research labs—including Stanford, Harvard, UPenn, Duke, UCLA, Mount Sinai, and Scripps.
AI Product Engineer · Production Agents · Complex Domains

Proof in production
Half a million dollars in cumulative revenue across co-founded AI ventures, with products adopted by 30+ research labs—including Stanford, Harvard, UPenn, Duke, UCLA, Mount Sinai, and Scripps.
Built production agent systems end to end—multi-agent orchestration, hybrid retrieval and reranking, structured validation, collaborative document editing, resilient model routing, and deeply interactive evals.
Turned ambiguous expert workflows into production AI products—from direct user discovery and product definition through architecture, evals, full-stack delivery, and production reliability.
As co-founder and CTO, set technical direction and led three engineers and a product manager while personally architecting and coding the core AI systems. Published first-author AI evaluation research with Springer, co-authored LLM benchmarking with Lenovo and AMD, and wrote a three-part enterprise AI engineering series.
Beyond the work
I am from Philadelphia, PA. I studied at Duke University. I currently live in Denver, CO.
I like playing and designing strategy board games, running, and tennis. My most useless skill is Etch A Sketch.
Selected work
I led product discovery, architecture, and delivery from the first customer conversations through production, and managed a team of three engineers and a product manager. Deployed at industry-leading enterprise customers.
Multi-agent orchestration, behavior-first evals that mirrored real user workflows, and layered content-validation loops combining deterministic checks, LLM judges, cross-output consistency checks, and expert override.
I built the full-stack product and AI systems for a collaborative writing workspace that helped researchers prepare NIH proposals, manage citations, ingest source documents, and work across multiple grants at once.
A Next.js and Node.js application backed by PostgreSQL, with Python/FastAPI agents, citation retrieval and validation, enterprise authentication, document export, and resilient routing across five LLM providers.
I built and shipped RAG, NLP, and transformer products for three Fortune 500 customers, including enterprise knowledge retrieval, executive competitive intelligence, automated feedback routing, and assistive writing for people with ALS.
Customer-facing RAG systems, evaluation pipelines, data workflows, low-latency transformer inference, and production deployment for on-premises enterprise environments.
Editor MCP Public repository · Demo-ready
I built a TypeScript MCP server and Tiptap/Yjs editor where agents make atomic, reviewable document changes instead of pasting text into chat. People can inspect, navigate to, accept, reject, or rewrite every proposed edit.
View source and demoVersioned semantic tools with revision guards and atomic edit groups
Live shared state through Tiptap, ProseMirror, Yjs, and Hocuspocus
Authenticated Streamable HTTP with OAuth and fail-closed validation
First author · TPCTC 2024
Designed and ran an experimental pipeline that generated synthetic question-answer datasets with four open-source LLM configurations, fine-tuned the same base model on each, and tested which dataset signals predicted downstream QA performance.
Found that semantic similarity can indicate domain diversity and that short chain-of-thought answers may flag lower-quality training data.
Co-author · TPCTC 2023
Co-authored a Lenovo and AMD review of how to compare LLM systems across training and inference—not only on accuracy and throughput, but also cost, energy, bias, trust, and sustainability.
Published by Springer in Lecture Notes in Computer Science; the chapter has been cited in subsequent benchmarking work.
Lenovo Press · Three-part engineering series
I co-authored a practical series for enterprise teams building domain-specific RAG systems, from model selection and evaluation through synthetic dataset creation, LoRA fine-tuning, hardware sizing, latency, and production tradeoffs. I also developed a containerized framework that customers could use to implement the approach.
Graduate AI coursework · 2024
Designed a card-shuffle simulator and trained encoder-only PyTorch transformers to test how model depth and increasing randomness affect ace-sequence prediction. Deeper models learned subtle structure in low-noise sequences; performance degraded as shuffle complexity increased.
Graduate AI coursework · 2024
Collected gameplay trajectories and trained PyTorch policies with behavioral cloning and DAgger for two-agent control. The strongest independent-policy pair scored 18 goals across 32 evaluation games and held the same scoring rate against unseen agents.
Contact