The work
Agentic software governance is a first-class design problem here, not a prompt you tune at the end. You own the layer that makes agents dependable: how they retrieve, how they’re judged, how they recover. Your work is what lets everyone else’s tooling — internal and client-facing — be trusted in production.
What you’ll do
- Design agentic workflows, retrieval, and model orchestration that hold up outside the happy path.
- Build the eval and critic systems that decide whether an agent’s output is right — the measurement that turns “it usually works” into “we can ship it.”
- Instrument everything: traces, scorecards, and regression suites that make model behavior observable and improvable.
- Choose models, tools, and context strategies pragmatically, and revisit them as the frontier moves under you.
- Partner with the tooling engineers so your systems land in software people actually use.
Who you are
- Strong engineering fundamentals plus real, hands-on work with LLMs, agents, retrieval, or evals — you’ve shipped, not just experimented.
- Imaginative and rigorous in equal measure: you invent approaches and then build the harness that proves them.
- Skeptical of benchmarks that don’t reflect the job; you design measurement that does.
- A well-rounded builder who can move from research spike to production system without a handoff.
- Curious, opinionated, kind.
What we offer
- Competitive cash + meaningful equity.
- Small team, high agency, long horizons, and problems at the actual frontier.
- The chance to define how reliable agentic software gets built.