Senior Data Scientist

HCLTech

LLM EngineerseniorLondon Area, United KingdomonsitefulltimeIT Services and IT ConsultingAgentic AIMicrosoft Agent Framework (MAF)Semantic KernelAutoGenModel Context Protocol (MCP)Retrieval-augmented generation (RAG)LLM-as-judge evaluationPrompt engineeringposted 08 Sep
Unlock apply linkApply links and the original listing are a Pro feature: £4.99/mo or £25 once.
About the role We are hiring a Senior AI Engineer to design and ship the AI systems our lawyers and business teams use every day: agentic systems that carry out multi-step work under supervision, retrieval-augmented generation over dense legal corpora, and the evaluation discipline that tells us any of it is actually working. You will own solutions end to end, from architecture and pipeline design through to a production service that is monitored, evaluated and improved. What you will do * Architect and build production agentic and RAG systems over the firm's documents and know-how: harness and control loops, tool design, retrieval, orchestration, and the evaluation harness that proves it works. * Design the pipelines end to end, from the data sources in SharePoint and the wider Microsoft 365 estate through to the service the user touches, and make the architecture decisions that keep it fast, secure and affordable at firm scale. * Own what you build end to end: architecture, code, deployment, monitoring and the improvements that follow. You write the code and you stay accountable for it once it is live. * Prototype fast. Get something working in front of real users in weeks, not quarters, then decide what to harden, what to scale and what to throw away. * Sit with partners, associates and business services teams, listen past the stated request to the real problem, and come back with a shaped solution and honest trade-offs. * Run structured comparisons of models, frameworks, retrieval strategies and orchestration patterns, and turn them into a recommendation the firm can stand behind. * Set the technical bar for AI work across the team: design reviews, code reviews, evaluation standards and the patterns other engineers reuse. * Mentor engineers and data scientists, review their work generously, and hand them problems that stretch them. * Share what you learn: internal write-ups, brown-bag sessions and presentations to audiences from a two-person product team to a global practice group. Core technical experience * Agentic AI: designing agent harnesses and control loops, tool and function calling, planning and memory, multi-agent patterns, human-in-the-loop checkpoints and guardrails, and keeping cost, latency and failure modes under control in production. * Microsoft Agent Framework (MAF): hands-on experience building with MAF and its Semantic Kernel and AutoGen lineage, and fluent enough in comparable agent SDKs to argue for the right one rather than the familiar one. * Model Context Protocol: building and consuming MCP servers and clients, designing clean tool and resource surfaces, and using MCP to connect agents to firm systems in a standard, reusable and governable way. * Retrieval-augmented generation: chunking and ingestion strategy, embeddings, hybrid keyword and vector retrieval, reranking, grounding and citation, and knowing why a given retrieval design fails on legal documents. * Evaluation: this is not an afterthought for us. Building eval sets and harnesses, offline and online evaluation, LLM-as-judge and its limits, regression testing across model and prompt changes, and metrics that a business owner recognises as meaningful. * Prompt engineering and context engineering: designing what the model sees, structuring and compressing context, managing long-context and multi-turn state, and systematic iteration rather than trial and error. * Token and cost optimisation: caching, context pruning, routing between models by task, batching, and measuring the cost and latency of a system rather than guessing at it. * Broad, current knowledge of the model landscape: GPT and other frontier models, Claude, and open-source models, with informed views on where each earns its place and how to move between them without rewriting the system. * Fine-tuning small language models: knowing when a fine-tuned small model beats a large general one, and having done the work end to end, from dataset construction through training and evaluation to deployment. * Expert Python, plus production and orchestration as a core strength: clean, tested, typed code, containerised services, CI/CD, workflow orchestration, cloud deployment (Azure preferred), observability, tracing and live evaluation. * System architecture, required for pipeline design: you can draw the whole system on a whiteboard and defend it. Ingestion and indexing pipelines, service boundaries and interfaces, data flow and storage choices, caching, queuing and batch versus real-time trade-offs, failure and recovery paths, and the cost, latency and security envelope the design has to live inside. * SharePoint and the Microsoft 365 data estate: building ingestion pipelines over SharePoint and OneDrive content through the Graph API, and working confidently with libraries, site columns, content types, managed metadata and permission trimming so retrieval respects the firm's access model rather than working around it. * Knowledge graphs: ontology and schema design, entity resolution, building graphs from unstructured text, and graph-aware retrieval that hands agents structured, verifiable context. * Foundations underneath all of it: 12+ years across data science, machine learning and software engineering, classical ML, statistics and experimental design, and the judgement to recognise when a language model is the wrong tool. Background * 12+ years of professional experience across data science, machine learning and software engineering, with a recent track record of AI systems that reached production and stayed there. * A scientist's background, strongly preferred: a PhD or an equivalent research record in machine learning, computer science, computational linguistics, statistics, physics or a related quantitative field.