🤩
*TLDR: You'll join Treeconomy — an earth-tech startup using satellite and ML models to measure forest carbon and help fund nature restoration projects — as an Machine Learning Engineer turning their biomass models (satellite + LiDAR + field data) from individual custom runs into automated, GPU-scaled pipelines on GCP, building out uncertainty quantification and data-requirement studies along the way. UK-based, remote or London-hybrid, full-time, ASAP start.*
🌱 Our Mission
At Treeconomy, we believe in the regenerative power of nature. We’re on a mission to combat climate change, restore ecosystems, and improve livelihoods.
By embracing remote sensing technology and financial innovation, we are revolutionising the forest carbon and natural capital industry and incentivising landowners to grow and conserve flourishing, climate-positive forests.
We have massive goals and are looking for passionate people to help us achieve them.
🚀Our Company
We’re a venture-backed earth-tech startup working at the dynamic intersection of technology, ecology, and finance. We are a small but committed team working to bring traceability, transparency, and trust to nature-based carbon removal and ecosystem restoration, harnessing innovation to make a meaningful difference by driving finance toward worthy projects.
We support rural landowners and natural capital investors with project development, financing, measurement, and monitoring services, connecting the highest quality projects to leading corporate carbon credit buyers and project investors via our innovative Sherwood platform.
✨ The Role
Machine Learning Engineer - Geospatial AI
Team: Science \& Engineering, reporting to Dr Matt Amos (Lead Scientist)
Level: Mid to Senior
Location: Must be UK-based either remote or London-hybrid.
The role
Treeconomy is building a asset-specific geospatial AI models that produceabove-ground biomass maps for afforestation projects, fusing multispectral satellite data with drone LiDAR and field inventory, and reporting a defensible confidence interval alongside every estimate.
Early versions of models and data ingestion already exist in our repos and have been run end-to-end. The job is to take them from one or two manual runs to a hundred-plus automated ones, and to build the experimental machinery that tells us where they're weakest. Your time will split roughly evenly between modelling work and the infrastructure that makes it repeatable at volume.
Who you'll work with
You'll join a small, senior technical team: a science lead, a data scientist, a data engineer and a full-stack engineer. You own the space between the data engineer and the scientists, taking modelling work and making it run reliably, at volume, with the instrumentation needed to learn from it.
What you'll own
* Scaling the run pipeline
* Turn existing model and ingestion code into orchestrated, parameterised pipelines that execute unattended across many sites, dates and configurations on GCP.
* GPU training and batch inference sized for whole-landscape runs. Throughput, cost per run, project size, and turnaround time are project KPIs.
* Experiment tracking and run provenance: every output traceable to its input data version, model version and config. When a number changes, we need to know why.
* The calibration loop where new field and/or LiDAR data arrive, triggering versioned retraining, with batch logs and regression checks against previous model versions.
Modelling and evaluation
* Design and run spatio-temporal ablation studies at scale. How much drone or field data is actually required to reach a given confidence level, and which input streams are carrying the prediction?
* Stress-test the models to find where they break: unseen biomes, persistent cloud cover, sparse ground truth, unusual canopy structures. How much data do we need and how can we turn this into a sampling plan for our customers?
* Implement uncertainty quantification: per-pixel confidence intervals, and propagation of input and model uncertainty through to the reported carbon figure. Evaluation on calibration and probabilistic scoring.
* Interrogate architecture choices : one model across biomes, or specialists per ecosystem?
* Work with learned representations from geospatial foundation models, and help us evaluate candidate alternatives on evidence rather than on hype.
* Build the multi-modal fusion path: co-registering Sentinel-2, VHR optical (Planet/Maxar), Sentinel-1 SAR, GEDI/ICESat-2 spaceborne LiDAR, drone LiDAR and field inventory into coherent training and evaluation sets.
Requirements
Essential:
* Strong Python. Comfortable with the geospatial stack (xarray, rasterio/GDAL, geopandas, dask or equivalent) and a modern deep learning framework (currently, we're in pytorch/jax but are flexible).
* Demonstrable experience training and evaluating ML models on spatial or other large gridded data
* Track record of taking research code to reliable automated execution: orchestration, containerisation, config management, CI, cloud batch/GPU compute. This is the single most important requirement.
* Experience with workflow orchestration and experiment tracking tooling (Airflow/Dagster/Prefect, Vertex AI Pipelines, MLflow/W\&B, or equivalents).
* Good grasp of uncertainty: you understand different uncertainty propagation and quantification principles and you care which one a stakeholder is asking for.
* Rigorous validation instincts. You default to spatially and temporally held-out evaluation because you know random splits leak in geospatial data.
* Software engineering fundamentals: version control, testing, code review, reproducible environments.
Desirable:
* GCP specifically (Vertex AI, Batch, GKE, BigQuery, Cloud Storage) and infrastructure-as-code. This is where we work currently.
* Experience working with pretrained representations or embeddings from foundation models, including their failure modes.
* Familiarity with LiDAR processing (multi-return point clouds, canopy height models) and forest inventory data (DBH, allometry).
* Exposure to carbon MRV or the relevant standards (Verra, Gold Standard, Isometric, IPCC LULUCF guidance).
* Published or otherwise peer-reviewed technical work in remote sensing, ecology or ML.
Characteristics we like
* Treats automation as a research tool
* Quantifies doubt rather than hiding it. Our commercial value is a credible error bar.
* Empirical over enthusiastic. Comfortable stopping an approach that doesn't beat the benchmark.
* Owns the whole path. Equally at home in a notebook, a Dockerfile, a Terraform config and a field-data QA review.
* Writes things down. Model cards, methodology write-ups and QA documentation are deliverables here as our customers need a paper trail.
* Works well with messy ground truth. Field data arrives late, inconsistent, and in the wrong CRS.
* Comfortable in a small team. Five+ technical people.
Success in the first 12 months
* Model runs are automatic.
* Every run is reproducible and traceable to its data and model versions.
* Ablation evidence exists that tells a project developer how much field data they need to buy for a given confidence level.
* Confidence/credible intervals generated programmatically for every biomass estimate, in a form external verifiers accept.
* We know, with evidence, the biggest weaknesses in the current model
🌍 Location \& hours
* UK-based, with core hours 10:00–16:00 for collaboration.
* Fully remote for candidates outside commuting distance of London (with a London visit approximately once every 8 weeks); if you're based near London, you’re welcome to co-work with us in central London 2 days a week.
📄 Contract
* Engagement: Full-time (ML Engineer)
* Time commitment: 5 days/week
* Start date: ASAP
* Gross Salary: Competitive (paid in GBP)
* PTO: 25 + Bank holidays
* Equipment: A company laptop can be provided if needed
🎉 Ready to Apply?
Excited by the chance to build software that measurably helps nature projects get off the ground? We’d love to hear from you — even if your profile isn’t a perfect match.
More AI roles like this, weekly
Roles like this expire in about a week. Get new AI openings across the UK in your inbox, free, unsubscribe any time.