ABOUT THE COMPANY
Rainmaker is built to give lawyers business intelligence for BD: legal news and deal data, a BD tool for building and managing approaches to prospective clients, plus a live rankings system that lets lawyers credential their work against their peers. The product launches later this year. It is AI-first in a market that has never had a product like it, and you will be shaping it from the start.
ABOUT THE ROLE
We are looking for our first dedicated AI Engineer. You will own the intelligence layer of the product: retrieval over legal news, deals and firm data, the agentic and guided-question flows inside the BD Centre, the extraction that turns unstructured coverage into structured records, and the evaluation infrastructure that tells us whether any of it is actually working.
This is a founding hire. The AI features that exist today were built alongside everything else. Your job is to take them from working to defensible, then build what comes next. You will decide how we do retrieval, how we evaluate, what we run in-house and what we buy, and you will be the person the rest of the engineering team asks when a feature depends on a model behaving predictably.
You will report to the CTO and work closely with our Head of Product and the wider engineering team. You will have meaningful equity.
We care about output rather than theory. Our users are lawyers, and lawyers do not forgive a confident wrong answer. Precision, traceability and knowing when the system should decline to answer matter more here than they do in most consumer products.
RESPONSIBILITIES
Retrieval and knowledge
- Own retrieval end to end: chunking, embedding, indexing, hybrid and re-ranked search over legal news, deal records, firm and lawyer profiles. - Build entity resolution that holds up across sources, so a firm, a lawyer and a deal mean the same thing wherever they appear. - Make grounding and citation a property of the system rather than a prompt instruction. Every claim the product shows a user should be traceable to a source. - Decide where retrieval ends and structured query begins, and stop us reaching for a model where a database would do.
Agentic and generative features
- Build the LLM-backed features in the product, including the guided-question loops in the BD Centre and the generative surfaces around dossiers, feeds and alerts. - Design agent and tool-use flows that fail safely, degrade to something useful and never invent a client, a deal or a quote. - Own prompt architecture as engineering rather than as text: versioned, tested, reviewable and cheap to change.
Evaluation and quality
- Build the eval harness. Define what good looks like for each AI surface, build the datasets, and make regression visible before a release rather than after it. - Instrument quality in production: hallucination rate, retrieval hit rate, refusal behaviour, latency and cost per interaction. - Run the experiments that decide model choice, and be willing to conclude that a smaller or cheaper model is the right answer.
Data and extraction
- Work with the data pipeline that ingests legal news and deal coverage, and own the extraction that turns it into structured, queryable records. - Improve precision and recall on that extraction over time, and know the current numbers for both. - Feed clean data into the rankings engine, and understand enough of its methodology to spot when the inputs are wrong.
Engineering and platform
- Ship production code into a TypeScript and Python stack running on AWS, and take responsibility for it in production. - Own cost, latency and reliability of the AI layer, including caching, batching, fallbacks and rate limits. - Build the internal tooling that lets non-engineers inspect, correct and improve model output without asking you.
Judgement and influence
- Tell Product what is feasible, what is expensive and what is a research project rather than a sprint. - Push back where a feature is being specified as an AI feature when it should not be one. - Set the standard and the practices that the next AI hires will work to.
WHAT WE ARE LOOKING FOR
- 5-8 years in software engineering, with at least two spent building LLM-backed systems that real users depend on. Production experience, not prototypes.
- Deep applied retrieval experience: you have built RAG systems that worked, and you can explain in detail why the first version did not.
- Rigour about evaluation. You have built eval infrastructure and you treat an unmeasured AI feature as an unfinished one.
- Strong engineering fundamentals in Python and TypeScript, comfortable in AWS and in production systems rather than notebooks.
- Fluency across the current model landscape and the tooling around it, with the judgement to pick the boring option when it wins.
- Comfort with ambiguity and with owning a domain alone. You will not have a team to delegate to for some time.
- Directness. You raise problems early, you argue your case on evidence, and you change your mind when the data says so.
PREFERRED
- Experience in legal, financial or another domain where accuracy is a professional obligation rather than a nice-to-have.
- Experience with entity resolution, knowledge graphs or document-heavy extraction at scale.
- Fine-tuning, distillation or model serving where it earned its keep against a hosted API.
- Early-stage or pre-launch experience, and a working understanding of what unit economics mean for an AI product.
- An open-source record, a paper, a side project, or anything else that shows us what you build when nobody is assigning it.
Rainmaker is an equal opportunity employer. We are committed to diversity and inclusivity.