Work

Newest first.

2024 to presentSynergyBoatFounding Engineer & CTO

SynergyBoat agent platform SDK

Built the agent platform SDK behind a US EV startup's move to 150+ B2B clients, with ~40% lower per-client engineering cost and extended runway on a $2M seed.

The agent platform dispatch A planner on the left decides and dispatches; it never runs untrusted code. Model-tier routing sits inside the planner: the small tier is the default, and a frontier model is chosen only when a plan exceeds a cost-or-quality budget. The planner dispatches into a sandbox boundary, drawn as a dashed enclosure. Inside it are three workers, pure function, isolated-vm, deno subprocess, which are pure functions or sandboxed routines, so a misbehaving plugin cannot take the process down. Each worker returns over a channel, and every channel carries a circuit breaker, drawn on the return path so the reader sees where a failure is stopped. The dispatch drawn here runs on isolated-vm, and its breaker closes as the result comes back. Under the planner two stores hold state: Redis for hot state and PostgreSQL for the durable audit trail, and either can be swapped without rewriting the planner graph. One edge arrives from outside the platform, labelled MCP: external AI agents reach the same workers without a new adapter. outside the platform external ai agent an outside agent reaches the same workers over mcp mcp sandbox boundary dispatch workers pure function isolated-vm deno subprocess a misbehaving plugin cannot take the process down each worker runs its plugin inside the sandbox, not in the process circuit breaker channels, one breaker on each planner llm-backed decision node never runs untrusted code model tier small the default frontier over budget the planner decides and never runs untrusted code itself pluggable state redis hot state postgresql audit trail state is pluggable: hot state in redis, the audit trail in postgres
part what it does
planner An LLM-backed decision node. It dispatches to workers and never executes untrusted code itself.
model tier Routing sits in the planner. Cheap calls default to smaller models; escalation to a frontier model happens only when a plan exceeds a cost-or-quality budget.
pure function A worker that is a plain function.
isolated-vm A sandboxed routine, run in an isolated VM.
deno subprocess A sandboxed routine, run in a Deno subprocess.
sandbox Workers run inside it, so a misbehaving plugin cannot take the process down.
channel Carries the result back to the planner. Every channel has a circuit breaker on it, drawn on the return path.
redis Hot state.
postgresql The durable audit trail.
state Pluggable. Either store swaps without rewriting the planner graph.
mcp External AI agents reach the same workers over MCP, without a new adapter.
Fig. 1. The agent platform. A planner decides and dispatches; it never runs untrusted code. Workers run inside a sandbox boundary, and every channel back carries a circuit breaker. State is pluggable, Redis for hot and PostgreSQL for the audit trail. Drawn for the agent platform only.

the case study

Nov 2025 to presentSynergyBoatLead engineer, sole

ChargerDojo, conformance testing for EV roaming

A hosted conformance runner for the OCPI roaming protocol that I build and run alone, with paying B2B customers and a mutation harness where every check is either proven able to fail or listed as one no mutation may name.

The ChargerDojo mutation harness A runner on the left calls the sim-peer on the right, which is the reference implementation. The OCPI wire between them is drawn as one lane each way with a tick per message, because the runner calls message by message rather than once. Four copies of the sim-peer hang off a bus, one per spec version. 2.1.1, where mod_cdrs says a CDR cannot change once sent, and this peer reads one back at a new total. 2.2, where status_codes says 2.2 has no hub code 4000, and this peer answers 4000 to a hub error. 2.2.1, where transport_and_format says limit=10 returns ten records, and this peer serves one object more. 2.3.0, where mod_tariffs says tax_included is required, and this peer serves a tariff without it. Each peer runs one mutation at a time. Under them a build gate stamps a tick for every mutant the checks catch, and one hold mark for a check no mutation may name, which is listed rather than proven. At the foot, dojo-connect, the Go connector, dials out through a dashed customer firewall on a WebSocket so an endpoint on a laptop can be reached. ocpi, message by message runner the runner walks a whole conversation, not one request sim-peer reference implementation copies, one per version 2.1.1 mod_cdrs: a CDR cannot change once sent this peer reads one back at a new total 2.2 status_codes: 2.2 has no hub code 4000 this peer answers 4000 to a hub error 2.2.1 transport_and_format: limit=10 returns ten records this peer serves one object more 2.3.0 mod_tariffs: tax_included is required this peer serves a tariff without it each peer runs one mutation at a time four copies of the simulator, each breaking one required rule build gate 2.1.1 2.2 2.2.1 2.3.0 listed every check is proven able to fail, or listed with a reason a tick per mutant caught, one hold mark for the unprovable check outbound websocket customer firewall dojo-connect go endpoint dojo-connect dials out through the customer firewall
version rule broken gate
wire The runner calls the peer over OCPI message by message, one lane each way, not a single request. n/a
2.1.1 mod_cdrs: a CDR cannot change once sent. This peer reads one back at a new total. caught
2.2 status_codes: 2.2 has no hub code 4000. This peer answers 4000 to a hub error. caught
2.2.1 transport_and_format: limit=10 returns ten records. This peer serves one object more. caught
2.3.0 mod_tariffs: tax_included is required. This peer serves a tariff without it. caught
listed A check no mutation may name, listed with its reason. not proven
connector dojo-connect, in Go, dials out through the customer firewall on a WebSocket so an endpoint on a laptop can be reached. n/a
Fig. 1. The mutation harness. The runner sweeps four mutant peers, each a copy of the simulator with exactly one required rule broken. A check no mutant reddens is listed rather than trusted, and the build fails if a check is neither proven nor listed. Drawn for ChargerDojo only.

chargerdojo.comDocsdojo-connect (Go connector)the case study

2026 to presentSynergyBoatFounding Engineer & CTO

Draftly, a multi-agent AI content platform

Content teams needed blogs, carousels, short videos and social copy without a studio behind them. Built Draftly: seven agents take a blog from brief to image plan in a fixed order inside a durable workflow, each step naming the capability tier it needs while the router walks a failover chain up from there, and the run ends in a revise loop that rewrites flagged blocks until the score clears or three rounds are up. Live at draftly.synergyboat.com, with a 13-tool MCP server beside it.

The Draftly tier floors and the failover chain Seven planned agent steps run left to right, and each one asks for a capability tier from 1 to 5 as a floor, set by hand in the plan. The floors, in order: brief-creator 4, research 3, drafter 5, seo 3, editorial-polish 4, humanizer 3, image-planner 4. Below them, the chain the router walks upward from a floor: it keeps every model at or above the floor whose provider key is configured, sorts by tier and then by a hand-set priority, and wraps the result in a failover chain. With the openai, gemini and anthropic keys, a step at minTier 3 resolves to gpt-5.4-mini, then gemini-2.5-pro, then gemini-3.5-flash, then claude-sonnet-4-6, then gpt-5.4, and a step at minTier 4 or 5 resolves to gemini-2.5-pro, then gemini-3.5-flash, then claude-sonnet-4-6, then gpt-5.4. With the gemini key only, minTier 3, 4 and 5 all resolve to the same chain, gemini-2.5-pro, then gemini-3.5-flash. When a step throws, the runner retries it one tier higher, capped at 5. With three keys that reaches a different chain. With the gemini key alone the chains are identical, so the retry re-runs the model that just failed. No cost lane is drawn, because the repository records no cost per stage. each step asks for a tier floor 5 4 3 4 3 5 3 4 3 4 brief- creator research drafter seo editorial- polish humanizer image- planner the floors are hand-set in the plan, one per step a floor is the lowest tier the router may pick, never a price every model at or above it whose provider key is set, lowest tier first with the openai, gemini and anthropic keys minTier 3 gpt-5.4-mini gemini-2.5-pro gemini-3.5-flash claude-sonnet-4-6 gpt-5.4 a step at tier 3 starts on the mini model when an openai key is set minTier 4 and 5 gemini-2.5-pro gemini-3.5-flash claude-sonnet-4-6 gpt-5.4 tiers 4 and 5 already share one chain, even with three keys with the gemini key only minTier 3, 4 and 5 gemini-2.5-pro gemini-3.5-flash one key flattens the ladder: all three floors resolve here escalate one tier on a thrown error, capped at 5 the same chain, so the retry re-runs the model that just failed no cost lane: the repository records no cost per stage
part what it does
brief-creator Asks for tier 4 as a floor.
research Asks for tier 3 as a floor.
drafter Asks for tier 5 as a floor.
seo Asks for tier 3 as a floor.
editorial-polish Asks for tier 4 as a floor.
humanizer Asks for tier 3 as a floor.
image-planner Asks for tier 4 as a floor.
minTier 3 with the openai, gemini and anthropic keys: gpt-5.4-mini, then gemini-2.5-pro, then gemini-3.5-flash, then claude-sonnet-4-6, then gpt-5.4.
minTier 4 and 5 with the openai, gemini and anthropic keys: gemini-2.5-pro, then gemini-3.5-flash, then claude-sonnet-4-6, then gpt-5.4.
minTier 3, 4 and 5 with the gemini key only: gemini-2.5-pro, then gemini-3.5-flash.
escalate A step that throws is retried one tier higher, capped at 5. With three keys that reaches a different chain.
one key With the gemini key only the three chains are identical, so the retry re-runs the model that just failed.
no cost lane Nothing in the routing reads a price, and the repository records no cost per stage.
Fig. 1. The tier floors and the failover chain. Each step asks for a floor, and the router takes the lowest-tier model at or above it whose key is set. Two key sets are drawn: with three keys a tier-3 step starts on a mini model, and with the gemini key alone every floor resolves to the same chain, so escalating a tier re-runs the model that just failed. Drawn for Draftly only.

Livethe case study

2026PersonalSolo project

EquityOS, an AI research co-pilot for public markets

Retail investors get verdicts they cannot audit. Built EquityOS: it turns public market data into plain-English calls and gates any live trade behind a paper-trade track record. Local-first, read-only by default, secrets in the keychain.

Nothing public to open yet.

2026Personal (open source)Solo project

Lens, a consent layer between AI agents and your data

AI agents get all-or-nothing access to your inbox. Building Lens: it gives an agent a scoped, auditable, revocable slice of Gmail by matching intent, not by handing over raw OAuth scopes.

Nothing public to open yet.

2026Personal (open source)Solo project

trustmebro, a completeness benchmark for AI coding agents

AI coding agents call it done long before it is. Building trustmebro: it grades behavioral coverage and uses deterministic static analysis to find the gap between looks-done and is-done, with no LLM judge in the loop.

GitHub

2024 to 2026SynergyBoatFounding Engineer & CTO

HirePulse, an agentic recruiting platform

Manual candidate sourcing does not scale. Built HirePulse: agents source candidates, score them on five factors, and run first-round voice-screening interviews, so a small team can run a full pipeline around the clock. Live at hire.synergyboat.com.

Live

2026SynergyBoatFounding Engineer & CTO

Dexter, an enforcement layer for AI coding agents

AI coding agents cross module boundaries that review misses. Built Dexter: a local daemon that runs four deterministic checks, architecture, convention, pattern and complexity, over a file as it is saved and over a diff an agent sends to its validate_changes MCP tool. Every check is plain code, none calls a model, and a failed architecture check marks the result blocked.

Site

2024 to 2026SynergyBoatFounding Engineer & CTO

MCP toolkit monorepo

Built one shared toolkit for AST editing, token-budgeted context, and query-intelligence, which cut the net-new boilerplate across SynergyBoat's AI products. Every internal agent depends on it.

Nothing public to open yet.

2026SynergyBoatFounding Engineer & CTO

DeepQuery, plain-English access to your databases

Teams sit on databases their non-technical people cannot query. Built DeepQuery for one client, an EV-charging network: a plain-English question passes a confidence gate, two candidate queries go through ten validation rules and an arbiter, and the winner runs read-only under row and time caps, returning rows, a chart and a narrative. Forty-three questions ran against that client's production data in April 2026, reconciled against the reporting they already had, and the reusable half of the pipeline is published as a package.

The DeepQuery gate and validators A question in plain English, Weekly session hours by charger, enters at the top and drops into the confidence gate. The gate is a scoring function with no model call behind it: it weighs six signals, metricMatch at 0.25, schemaCoverage at 0.2, clarity at 0.2, joinConfidence at 0.1, permissionCompleteness at 0.1, profileCoverage at 0.15, and picks one of four actions. The actions and their thresholds: answer when overall >= 0.7, answer-with-caveats when overall >= 0.4, clarify when 0.4 to 0.7 and clarity < 0.5, refuse when overall < 0.4. Refuse ends the run by design. Clarify reaches the wire and ends it by accident: the engine streams a clarification event and the browser handles four event names that do not include it. Schema context for the prompt is chosen rather than summarised: at most six tables through pgvector similarity search, with a keyword fallback at four. The two answering paths continue to candidate generation, where a governed metric formula gives two candidates: profile, compiled from the governed metric formula; and llm, planned by the model, with the profile query as its baseline. Both candidates cross validatePlan's ten rules. Four are structural: table-existence, write-op-detection, json-parse, column-existence. Six are semantic: metric-preservation, grain-preservation, filter-preservation, additivity-check, certification-check, lookup-completeness. An arbiter picks one, and a violation sends a candidate back for up to two repairs. The winner runs read-only as a single statement, with a LIMIT injected when it lacks one, capped at 1,000 rows and 30 seconds. validateResult runs on the rows a second time after execution, and a failure re-enters the repair loop. What comes back over SSE is rows, a chart spec, a narrative and a provenance record. question, in plain english Weekly session hours by charger one of the 43 questions recorded against the client's data confidence gate a scoring function, no model call metricMatch 0.25 schemaCoverage 0.2 clarity 0.2 joinConfidence 0.1 permissionCompleteness 0.1 profileCoverage 0.15 six weighted signals decide the action, before any model is called answer overall >= 0.7 the plain path, and the one all 43 recorded runs took answer-with- caveats overall >= 0.4 the same path, with what the gate could not confirm attached clarify 0.4 to 0.7 and clarity < 0.5 stops at the wire one hardcoded question; the browser drops the event and shows nothing refuse overall < 0.4 stops here nothing is generated and nothing is run where the semantic profile defines the metric, two candidates profile compiled from the governed metric formula compiled from the governed metric formula llm planned by the model, with the profile query as its baseline planned by the model, with the profile query as its baseline validatePlan ten rules, and both candidates take all of them structural semantic table-existence write-op-detection json-parse column-existence metric-preservation grain-preservation filter-preservation additivity-check certification-check lookup-completeness four structural rules and six semantic ones, before anything runs arbitrate picks the candidate that runs, or sends it back the arbiter decides which candidate runs, and it can decide none does up to 2 repairs execution read-only, one statement 1,000 rows, 30 s a LIMIT is injected when the query forgot one, and comments are rejected validateResult the rows are checked after the query ran two of the 43 recorded runs failed here and took a repair over sse: rows, a chart spec, a narrative, a provenance record
part what it does
question Weekly session hours by charger, typed in plain English.
confidence gate A scoring function with no model call. It weighs metricMatch at 0.25, schemaCoverage at 0.2, clarity at 0.2, joinConfidence at 0.1, permissionCompleteness at 0.1, profileCoverage at 0.15.
answer overall >= 0.7. The plain path, and the one all 43 recorded runs took.
answer-with-caveats overall >= 0.4. The same path, with what the gate could not confirm attached.
clarify 0.4 to 0.7 and clarity < 0.5. One hardcoded question; the browser drops the event and shows nothing.
refuse overall < 0.4. Nothing is generated and nothing is run.
schema context Chosen rather than summarised: at most six tables through pgvector similarity search, and four through the keyword fallback.
profile A candidate query, compiled from the governed metric formula.
llm A candidate query, planned by the model, with the profile query as its baseline.
validatePlan Ten rules, and both candidates take all of them. Structural: table-existence, write-op-detection, json-parse, column-existence. Semantic: metric-preservation, grain-preservation, filter-preservation, additivity-check, certification-check, lookup-completeness.
arbitrate Picks the candidate that runs. A violation goes back for up to two repairs.
execution Read-only, a single statement, comments rejected, a LIMIT injected when the query lacks one, capped at 1,000 rows and 30 seconds.
validateResult The rows are validated a second time after execution, and a failure re-enters the repair loop.
over sse Rows, a chart spec, a narrative and a provenance record, streamed back as the run ends.
Fig. 1. The gate and the validators. A question is scored on six weighted signals before any model is called, and the score picks one of four actions; two of them end the run there. What survives is generated twice, put through ten rules, picked by an arbiter, run read-only under a row and a time cap, and validated again once the rows are back. Drawn for DeepQuery only.

the case study

2021 to 2023Zuddl (YC S20)Software Engineer II to Senior Software Engineer

Zuddl Webinars v1

Zuddl ran live events and needed a webinar product beside them, built from scratch to a public release date. Shipped Webinar Platform v1 by mid-2023 with a new design system: dynamic access control with session and zone permissions, and the live event experience across stage, polls, A/V and networking. Covered it with controller and repository-level integration tests. Joined as Software Engineer II in 2021, moved to senior in 2023, and stayed through the $13.35M Series A.

Nothing public to open yet.

2018 to 2020FnplusCo-founder

Geekfolio, a proof-of-work profile for engineers

A resume asserts and offers no proof for what it claims, which is the complaint the 2020 argument page opens on. Designed Geekfolio at Fnplus, the fourth of five tools in its learning product: it would read an engineer's informal-learning trail from across the web through APIs, run the company's own assessments at Geeks Garage sessions and hackathon-like events, and place people on what they had actually done rather than on what they had written down. It stayed a design.

The 2020 argumentFnplus product page