← All open roles
EngineeringEU · Remote

AI / LLM Engineer

Agents, evals, tool-use orchestration. Build the layer that turns a 45-min walkthrough into a working OS.

About the role

The hardest engineering problem on the team is turning a customer's 45-minute description of their business into a reliable, durable set of automations. This is your problem to own. You'll build the agent runtime, the eval suite, the tool-use graph, and the prompts that make all of it work in production.

What you'll do

  • Own the agent runtime: planning, tool-use, retries, escalation.
  • Build evals that catch regressions before customers do.
  • Get prompts to a place where the agent is honest about its uncertainty and asks for help when it should.
  • Work alongside customers in week one to see what actually breaks.

Who you are

  • Have shipped LLM-powered features to real customers, not just demos.
  • Comfortable with the limits of current models — and excited about pushing them.
  • Strong Python or TypeScript (we use both).
  • Care about evals more than benchmarks.

Nice to have

  • Open-source contributions to LangChain, LlamaIndex, OpenInference, or similar.
  • Have run a prod LLM service with >100k req/day.
  • Have read the Lily Mahdavi paper on tool-use evals.

Compensation

€100–140k base + equity. Top of EU market band.