*** Temporary employment role through Oxford UK*
Role:
AWS Data Engineers-AI/ML Specialist x3
Location:
Remote-Europe
Start Date:
ASAP
Length:
Until Dec, 31st initially. Potential extensions
Our client are seeking 3 experienced AWS Data \& AI/ML Specialists to support large-scale cloud transformation engagements. This is a hands-on technical role focused on designing, building, and deploying scalable data, machine learning, and Generative AI solutions on AWS.
Project Background
The customer is modernising its data platform by migrating from a legacy
Amazon Redshift
environment, which has been in place for approximately
8 years**
, to a modern, event-driven AWS data platform.
The existing platform consists of:
* Amazon Redshift clusters
* A relatively simple architecture running on EC2
* SQL-based orchestration with scheduled jobs processing a traditional relational database
The new platform is designed around a modern data lake architecture and will focus on preparing data for downstream data marts using a
Bronze → Silver → Gold
data model.
Scope of the Migration
* Migrate the existing
Silver
and
Gold
data layers onto the new platform.
* Implement an
event-driven architecture
using AWS services.
* AWS Lambda functions will orchestrate processing workflows.
* Lambda triggers
dbt
, which submits
Spark
jobs.
* Spark processes the data through the Bronze, Silver and Gold layers.
* The platform follows a
hub-and-spoke architecture
:
* Centralised AWS account for storing and governing data.
* Decentralised consumer accounts for accessing curated datasets.
Data Platform \& Governance
* AWS-based modern data lake.
* Data governance managed through
AWS Lake Formation
for permissions, data sharing and access control.
* Backend processing built on
Apache Spark
(open source).
* Existing Redshift data warehouse landscape (7-8 warehouses) is being migrated to the new lake architecture.
* Informatica
remains part of the ecosystem.
* New processing framework built around
AWS Glue
.
* Infrastructure provisioning managed through
Terraform
.
Current Project Status
The project is entering a critical delivery phase:
* Development is progressing rapidly.
* Large portions of the solution have been developed but
have not yet been fully tested
.
* The customer expects rapid support as releases move into development and testing.
* The team will need to troubleshoot and resolve issues as they emerge, particularly around dbt and Spark processing.
Scale \& Complexity
This is a large-scale enterprise migration with aggressive delivery timelines.
Key metrics include:
* Approximately
20,000 raw source tables
.
* Large numbers of Silver transformation objects.
* Around
65 reusable Gold-layer fact/data objects
(with additional fact tables still being confirmed).
* High-complexity data transformations requiring migration within a demanding timeframe.
Key Challenges
* Migrating highly complex data objects under significant time pressure.
* Supporting customer releases as they move into development.
* Troubleshooting dbt pipelines and Spark job failures.
* Working within a platform that is largely developed but not yet fully validated through testing.
* Delivering reliable production support during a major modernisation programme.
Platform Approach
* Event-driven processing rather than scheduled SQL jobs.
* Modern data lake architecture replacing multiple legacy Redshift warehouses.
* Strong emphasis on governance rather than a pure Data Mesh implementation.
* Increasing use of automation and an
agentic
approach, while still requiring human intervention for operational support and troubleshooting.
Core Technology Stack
* AWS
* Amazon Redshift (legacy)
* AWS Lambda
* AWS Glue
* AWS Lake Formation
* AWS CloudWatch
* Amazon DynamoDB
* Apache Spark
* dbt
* Informatica
* Terraform
* EC2
More ML Engineer roles like this, weekly
Roles like this expire in about a week. Get new ML Engineer openings across the UK in your inbox, free, unsubscribe any time.