SynergyBoat agent platform SDK
2024 to present · SynergyBoat · published
Built the agent platform SDK behind a US EV startup's move to 150+ B2B clients, with ~40% lower per-client engineering cost and extended runway on a $2M seed.
The problem
A US EV startup needed to onboard 150+ B2B clients without scaling linearly in engineering headcount. The manual parts of client onboarding, catalog ingestion, and daily operations were eating every new hire. The hard part of onboarding 150 clients was orchestration. The question was whether a single platform could handle the variable work of 150 clients without one integration at a time.
The constraint the team named: stay on a $2M seed budget, avoid vendor lock-in, and produce something a non-founding engineer could extend in a year.
What shipped
The SDK is SynergyBoat’s @synergyboat/agents-sdk, and I wrote most of it.
It is an agent platform SDK (TypeScript) with MCP protocol support, pluggable state
stores (Redis and PostgreSQL), sandboxed code execution via isolated-vm and Deno
subprocesses, and inter-agent communication with circuit breakers. Client
workflows now run as planner-and-worker compositions on top of the SDK.
Outcomes over twelve months: 150+ B2B clients live, ~40% reduction in per-client engineering cost, runway extended on the $2M seed, and the team positioned for Series A.
How it works
The SDK centers on three primitives (Figure 1): planners (LLM-backed decision nodes), workers (pure functions or sandboxed routines), and channels (the communication surface, with circuit breakers on every wire). Planners never execute untrusted code directly. They dispatch to workers. Workers are sandboxed so a misbehaving plugin can’t take down the process.
State is pluggable. Redis for hot state, PostgreSQL for durable audit trails. Clients can swap either without rewriting the planner graph. MCP support means the same workers are reachable from external AI agents without writing a new adapter each time.
Model-tier routing sits in the planner: cheap calls default to smaller models, escalation to frontier models happens only when a plan exceeds a cost-or-quality budget. That is inference spend, not the per-client engineering cost above.
The hard part of onboarding 150 clients was orchestration.
What I’d do differently
The first version of the planner logged too much and explained too little. Most of what it logged was noise. The second version kept only the events that would reproduce a bad decision; that made the hard debugging tractable. I should have designed for that from day one.
The MCP layer was also bolted on late. If I were to rebuild, MCP shape would be the first thing the planner spoke natively. Too many use cases depend on it to treat it as an adapter.
| part | what it does |
|---|---|
| planner | An LLM-backed decision node. It dispatches to workers and never executes untrusted code itself. |
| model tier | Routing sits in the planner. Cheap calls default to smaller models; escalation to a frontier model happens only when a plan exceeds a cost-or-quality budget. |
| pure function | A worker that is a plain function. |
| isolated-vm | A sandboxed routine, run in an isolated VM. |
| deno subprocess | A sandboxed routine, run in a Deno subprocess. |
| sandbox | Workers run inside it, so a misbehaving plugin cannot take the process down. |
| channel | Carries the result back to the planner. Every channel has a circuit breaker on it, drawn on the return path. |
| redis | Hot state. |
| postgresql | The durable audit trail. |
| state | Pluggable. Either store swaps without rewriting the planner graph. |
| mcp | External AI agents reach the same workers over MCP, without a new adapter. |