NLP Research Engineer

Forestreet

AI ResearchLondon Area, United KingdomonsitefulltimeSoftware DevelopmentNLPMachine LearningSentence TransformersFAISSTF-IDFtopic modelingPydanticvector spaceposted
Unlock apply linkApply links and the original listing are a Pro feature: £4.99/mo or £25 once.

Our Mission is to democratise the global research industry by providing unbiased business intelligence through analysis of mass public data, using Machine Learning and NLP techniques. We have recently been acquired by Beroe, but have kept the 'startup feeling' within our team, which is now looking to grow.

Our Company values

Curious, Helpful, Independent, Responsible, Remarkable

About the role

The role sits at the front of our data pipeline, designing and training the NLP components that our ETL team productionizes downstream.

While we use agentic LLMs throughout our stack, this role is rooted in traditional NLP craft : you should be as comfortable reasoning about vector comparison, clustering behavior, and tokenization edge cases as you are prompting a model. You'll work primarily in vector space — building, tuning, and evaluating embedding-based systems that transform raw text into structured, meaningful signal.

Output from this role — models, embedding pipelines, evaluation harnesses, labeled datasets and classifiers etc — becomes the spec that our ETL Engineers build into production services, so clear documentation and handoff discipline matter as much as the research itself.

We recognize the importance of proprietary data, and is in the process of creating an extensive network graph capturing decades of analyst insights as well as public information, allowing Graph RAG into the brains of an all-knowing procurement expert.

We work a 4-day week, with on-call every other Friday (business hours only).

Responsibilities

  • Design, train, and evaluate NLP models and embedding pipelines, using both traditional techniques (TF-IDF, topic modeling, rule-based/statistical NLP) and Sentence Transformers where they earn their complexity;
  • Own the vector space end-to-end: procuring training data, model/architecture choice, dimensionality trade-offs, indexing strategy in FAISS, drift monitoring, and rigorous evaluation using precision/recall, clustering quality, human-in-the-loop review) rather than relying on

*vibes* or leaderboard scores;

  • Right-size every model decision: default to CPU-friendly and lightweight approaches, and build the evidence-based case required before reaching for GPU infrastructure;
  • Build and maintain labeled datasets — including guidelines, annotation QA, and versioning — that are clean and well-documented enough for ETL Engineers to productionize without guesswork;
  • Prototype using agentic LLMs, while maintaining independent and conscious judgment on when an LLM-based approach is the right/wrong tool for the job;
  • Partner with ETL/DevOps Engineers to define handoff contracts: model artifacts, inference expectations, latency/cost budgets, and schema for outputs (via Pydantic) that plug cleanly into the existing pipeline;
  • Mentor junior team members on NLP fundamentals and evaluation rigor, and contribute to code reviews on the research side of the pipeline.

Qualities/skills we are looking for

  • Vector-Native:

You have a deep, intuitive understanding of embedding spaces — how to construct them, index and query them efficiently (e.g. FAISS), and critically evaluate whether they're actually capturing what you think they are.

  • Lean by Instinct:

You find elegance in solving a problem with the smallest model that reliably works, and you treat GPU spend and spin-up time as a cost to justify, not a resource to reach for by default.

  • ML-Skeptical, Not ML-Averse:

You reach for machine learning when it's the right tool, and you're equally willing to argue against it — or against a heavier model — when a simpler statistical or rule-based method will do the job more reliably or cheaply.

  • Rigorous Evaluator:

You don't trust a model until you've stress-tested it. You design evaluation frameworks before you design the model, and you're honest about failure modes.

  • AI Capable, but not Reliant:

You use agentic LLMs and coding assistants to move faster, but your judgment — not the model's confidence — is the final word on quality.

  • Discipline:

You value clean, reproducible research code, appreciate versioned datasets, and document your methodology so others can build on it without reverse-engineering your notebook.

  • Self Driven:

You take ownership of open-ended research questions, and you're comfortable defining the problem, but stay clear headed to back out of a rabbit hole to a find second opinion.

  • Collaborative Handoff Mindset:

You understand that your output is someone else's input — you write for the ETL Engineer who has to productionize your work, not just for yourself.

  • Curious and Rigorous:

You stay current with NLP research, but you evaluate new techniques on evidence, not hype, as well as considering maintenance before recommending they enter the pipeline.

Advice for applicants

:

If you believe you are up for the challenge of being a Forestreeter, we have the following advice before you apply:

  • We do not believe in numeric years of experience -

we care about how you think and build

.

  • Github portfolio > CV –

*we love to see your Github* . It does not need to be perfect – progression is what we want to see.

  • We value expertise outside of what we require above – whether it’s a different industry you came from, or a second programming language – they all contain value you can bring to the table.

Company Benefits

  • 4 day work-week (on-call every other Friday during business hours).
  • Competitive salary.
  • Hybrid working environment with two days per week in a well provisioned company office.
  • Pension plan and flexible benefits.
  • 20 days holiday per year.

Interview Process

  • 1. Discovery call
  • 2. First technical interview
  • 3. Take-home task preceding second technical interview
  • 4. Meet the team and leadership at our office in Farringdon / Chancery Lane
  • 5. Values interview with Co Founders
  • 6. Offer

We are a small team, and we insist on Engineers hiring Engineers : the interviewers are the same people you will be working alongside. However this may mean that we may not have lightning response time in the middle of a deployment!