Full Stack with Deep AI coding experience (LATAM / Europe - Remote)

Braintrust

LLM EngineersenioronsitefulltimeTechnology, Information and InternetLLM application layerretrieval architecturehybrid retrievalprompt engineeringtool-useRAGmulti-tenant scopingproduction LLM systemsposted 03 Sep
Unlock apply linkApply links and the original listing are a Pro feature: £4.99/mo or £25 once.
**Sr. AI/LLM Engineer** Reunion Marketing * Altitude Intelligence * Full-time You will own the LLM application layer of Altitude Intelligence — the chat, reports, and analysis engine our teams and automotive dealer clients rely on for performance answers. We are looking for a senior engineer who has operated production LLM systems where mistakes carry real consequences, and who sees beyond the implementation — understanding how what we build affects client outcomes and where it fits in the bigger picture. About This Role Altitude Intelligence is the AI layer of Reunion Marketing's platform. It turns live client performance data and our marketing methodology into answers, reports, and analyses that our teams and automotive dealer clients act on — through three product surfaces. You own the LLM application layer behind all three. In practice, that means the defining questions are yours: * What does the model retrieve, and how? * When is the right tool prompting, tool-use, or retrieval? * Where does the human belong in the loop? * What does "ready for clients" mean — and how do we prove it? It is a senior, hands-on role with unusual visibility: every output of this system either protects a client relationship, wins one, or grows one, and we expect engineering decisions to be made with that in mind. **The opportunity** * Production data at scale. Live performance data across five marketing product lines for hundreds of automotive dealerships, not a demonstration dataset. * Daily users who depend on the output. Strategists, client success, sales, and external clients consume what this system produces as part of their working day. * Meaningful stakes. An incorrect answer does not stay inside a demo; it can reach a client conversation. The engineering standard follows from that. * Genuine ownership. The product direction is set and shipping, and many of the significant architectural decisions are still open for this hire to make and defend. **The perspective we expect** This role calls for more than strong implementation. The engineer we hire will: * Understand where each piece of work fits in the product and what it changes for our clients and teams — retention, expansion, hours returned — and let that understanding shape the engineering decisions. * See the system end to end: how a retrieval choice shows up in a report, how a prompt change lands in a client conversation, how today's shortcut becomes next quarter's incident. * Treat an incorrect number in front of a client as the most expensive defect the system can produce, and design the grounding, evaluation, and review gates accordingly. * Explain technical trade-offs to strategists and executives in their terms, and be willing to defend a delayed release when shipping would put client trust at risk. * Judge their own success by the accuracy, adoption, and time savings the system delivers. What You'll Own * The answer pipeline — from question or scheduled trigger, through data retrieval and reasoning, to a grounded, client-safe artifact: structured outputs, versioned prompts managed as reviewed artifacts, per-product context assembly, and multi-tenant scoping that is never bypassed. * Retrieval architecture — a hybrid retrieval problem spanning relational product data, client history, and our Knowledge Vault, the curated knowledge base of methodology and strategy material that grounds the system's answers. Responsibilities include the ingestion pipeline, hybrid search across embedding and lexical indexes with reranking, relevance evaluation, and selecting the appropriate retrieval method for each use case. * Generation workflows — the multi-step flows behind report and analysis generation, built to be durable, observable, retryable, and cancellable rather than fragile request handlers with hidden state. * Evaluation as an engineering discipline — we run large structured evaluation banks against live data and gate releases on them. You will own and grow that system: offline suites, regression gating in CI, factuality and groundedness checks, cost and latency tracking, and converting reviewer feedback into measurable quality improvement. * Production reliability — hallucinations, retrieval misses, tool-use failures, output drift, cost spikes, model deprecations, and the rollbacks that follow. We expect instrumentation before conjecture, and decisions in writing. * Architectural judgment — prompting versus retrieval versus tool-use versus fine-tuning, human review versus automated release, iteration speed versus client-safe controls. Well-supported positions are expected; unsupported ones are not. What We're Looking For * 6+ years of software engineering experience, including 2+ years shipping production LLM systems that were customer-facing, with real consequences when they failed. Internal demonstrations and proofs of concept do not meet this bar. * Strength in a modern application language used for production LLM systems, with a record of well-structured production systems rather than glue scripts. * Deep, hands-on familiarity with the LLM stack: prompt design and structured outputs; retrieval architecture (vector, lexical, hybrid, reranking, multi-tenant scoping); tool-use and agent design, including when not to use them; evaluation offline, online, and as a regression gate; and a defensible position on fine-tuning, including when to avoid it. * An evaluation suite you designed that caught a genuine regression in production, and the ability to walk through what it caught and what shipped, or did not, because of it. * Production debugging instincts for LLM failures — specific failure modes you have fixed, and a habit of instrumenting before guessing. * A record of owning the quality bar on a multi-team AI product, including incident response, rolling back a prompt change, and defending a delayed release to stakeholders who wanted to ship. * Strong written and architectural communication — a one-page proposal that a non-AI engineer, a strategist, and a CTO can all read and respond to. **Nice to have** * Production experience with Anthropic Claude specifically, including tool-use, prompt caching, and streaming. * Multi-tenant client data and the access-control patterns that come with it. * Domain experience in marketing analytics, SEO/SEM, paid media, local search, or answer-engine visibility. * Experience building a system of record that an organization runs on, rather than a feature inside one.