The role
This is the build-focused seat. Where our FDEs live inside client environments, you’ll spend most of your time writing the code, agent systems, MCP servers, evaluation harnesses, and the internal tooling that lets a small team deliver like a larger one.
The Claude Agent SDK gives you the same agent loop, tool execution, and context management that powers Claude Code, programmable in Python and TypeScript. Your job is to turn that into systems that survive contact with messy client data and keep running after we’ve gone home.
It’s worth being clear about what this role is not: it’s closer to platform engineering than to ML research. You won’t be training models. You’ll be designing subagent architectures, writing tool definitions, managing context across long-running tasks, and building the eval suites that tell us whether any of it works.
How we build, so you can decide whether it’s how you build: the deterministic core ships first and gets proven in isolation; evals stand up alongside it as the change gate, not after it; the model layer adds judgment without owning decisions; guardrails fail closed; and nothing irreversible happens without a human approving it. Solve reliably first, innovate second.
What you’ll do
- Build production agent systems with the Claude Agent SDK: subagent orchestration, hooks, permission design, and context management
- Write MCP servers connecting Claude to client systems: Workspace, BigQuery, CRMs, internal APIs, and third-party SaaS
- Build and maintain evaluation suites and regression tests for agent behavior: golden datasets built from client history, code graders for anything with a right answer (thresholds, math, required structure), model graders for quality and grounding, and pass bars wired to run on every change. This is a first-class part of the job, not an afterthought
- Implement the guardrail layer: retrieval allowlists that treat fetched content as data rather than instructions, fail-closed output screening, and stakes-by-confidence routing to human review
- Deploy and operate agents on Vertex AI inside client Google Cloud projects, with attention to cost, latency, and token spend, including right-sizing model tier to task rather than defaulting everything to the biggest model
- Develop our internal accelerators: the reusable components, reference implementations, and framework docs that make the fifth client engagement faster than the first. A library already exists; you’ll be its heaviest contributor
- Handle the observability layer: logging every gate decision so incidents can be reconstructed, tracing, cost attribution, and failure alerting
- Write the internal documentation and reference implementations the rest of the team builds from
What we’re looking for
- 3+ years of production software engineering in Python or TypeScript
- Demonstrated experience building with LLMs in production: tool use, structured outputs, prompt caching, streaming, and retrieval
- Experience with Vertex AI or other Google Cloud GenAI and infrastructure
- Some experience with the Claude Agent SDK or comparable agent frameworks, and with MCP
- A real point of view on evaluating non-deterministic systems. If your answer to “how do you know it works” is “we tried it,” this isn’t a fit. We want to hear about datasets, graders, and pass bars
- The discipline to keep rules and arithmetic out of the model; you treat domain policy as provided data and code, never as something Claude is assumed to know
- Willingness to join a client call. You won’t live in them, but you won’t hide from them either
How to apply
Send a resume and links to code.
Email your application to [email protected].