AI / LLM Engineer
Agents, evals, tool-use orchestration. Build the layer that turns a 45-min walkthrough into a working OS.
About the role
The hardest engineering problem on the team is turning a customer's 45-minute description of their business into a reliable, durable set of automations. This is your problem to own. You'll build the agent runtime, the eval suite, the tool-use graph, and the prompts that make all of it work in production.
What you'll do
- Own the agent runtime: planning, tool-use, retries, escalation.
- Build evals that catch regressions before customers do.
- Get prompts to a place where the agent is honest about its uncertainty and asks for help when it should.
- Work alongside customers in week one to see what actually breaks.
Who you are
- Have shipped LLM-powered features to real customers, not just demos.
- Comfortable with the limits of current models — and excited about pushing them.
- Strong Python or TypeScript (we use both).
- Care about evals more than benchmarks.
Nice to have
- Open-source contributions to LangChain, LlamaIndex, OpenInference, or similar.
- Have run a prod LLM service with >100k req/day.
- Have read the Lily Mahdavi paper on tool-use evals.
Compensation
€100–140k base + equity. Top of EU market band.
Not the right one? Try these.
Senior Full-stack Engineer
TypeScript + Postgres + LLM agents. You'll own a whole vertical of the OS — design, ship, support — from week one.
Apply →Design Engineer
Bridge between design and shipped UI. The OS surface is a craft job — every micro-interaction matters.
Apply →Product Designer
Translate messy ops into clear flows. You'll own how each business sees their own OS.
Apply →