Head of Physical AI

Alcor

London, ENG, GBfulltimeposted
Unlock apply linkApply links and the original listing are a Pro feature: £4.99/mo or £25 once.
Location: London is an absolute preference, then Boston or San Diego Employment: Full-time Reports to: CEO Role type: Hands-on technical leader and team builder Travel: As needed About Miraxis Miraxis builds the data, evaluation, and deployment layer for Physical AI. We work across multimodal robot and human data, annotation and assurance, model evaluation, and the systems that turn physical-world experience into useful robot behavior. We are hardware- and model-agnostic. We care whether a dataset, model, or method produces a measurable improvement on a real task. We will build focused model and evaluation capabilities where they strengthen our data products, prove the value of our data, or solve a clear customer or partner problem. The role The Head of Physical AI establishes and leads Miraxis’s AI research and engineering function. You decide which Physical AI problems we work on, define how we test them, and stay directly involved in the most important technical work. You connect four areas that often sit apart: * Multimodal and embodied data * Transformer-based models and robot policies * Rigorous offline and real-world evaluation * Deployment on physical systems This is a player-coach role. In the first year, at least half of your time is direct technical work: designing models and experiments, reviewing or writing code, inspecting data, debugging training, examining failures, and reviewing robot rollouts. You will also build a small research and engineering team. Add people only when the work requires them. Mandate Turn Miraxis data and technical access into measurable advances in Physical AI systems. You own the answers, and the evidence, to questions such as: * Which model families should we train, adapt, or evaluate? * Which data produces meaningful gains in robot performance? * What coverage, sensors, annotations, and quality controls do the models require? * Which offline measurements predict real-world robot performance? * Where should we adapt existing open models rather than train from scratch? * What should Miraxis build itself, and what should it reuse or obtain through partners? Responsibilities Set the technical direction: Define a focused research and engineering roadmap. Select a small number of high-value bets with clear hypotheses, baselines, milestones, and stop criteria. Decide what we build, adapt, license, or access through partners. Drop work that no longer has a strong case. Likely scope includes vision-language-action models, multimodal Transformers, robot foundation models, action representation, imitation and reinforcement learning, world models, cross-embodiment transfer, data-efficient adaptation, and robot-policy evaluation. Own Transformer research and engineering: Design, adapt, train, and evaluate Transformer-based systems for embodied tasks. Make the architecture and training decisions yourself. Start from strong existing models and baselines. Train from scratch only when evidence supports the cost. Build or review the critical code. Connect models to data strategy: Define the data needed to train and evaluate the selected models. Specify sensors, modalities, annotations, mixtures, and quality controls. Measure the effect of data quality, diversity, and composition on model behavior. Distinguish data volume from data value. You set technical data requirements; you do not run annotation workforces or field collection. Own evaluation: Establish offline and real-world evaluation systems, baselines, held-out conditions, and release gates. Guard against leakage, overfitting, and weak success definitions. Compare offline metrics with real-robot outcomes. A benchmark score alone is not a successful result. The evaluation must support a decision. Move work onto physical systems: Take projects from problem definition through training, hardware integration, and real-world testing. Design safe test plans and staged deployment. Analyse failures across data, perception, model, control, hardware, and environment. Validate material claims on physical systems. Build a small technical team: Recruit a few complementary researchers and engineers. Lead architecture, experiment, code, and failure reviews. Keep the team focused on validated results, not paper count, benchmark theatre, or headcount. Represent the function and own responsible development: Explain the strategy to customers, partners, and investors. Convert customer needs into testable technical requirements. State uncertainty and limits clearly. Define safety constraints, supervision, and rollback paths for physical trials. Require reproducible evidence before making capability or safety claims. What success looks like 90 days: Audit current data, evaluation assets, partnerships, and model opportunities. Define one or two focused research bets with hypotheses, baselines, data, compute, success measures, and stop criteria. Establish a reproducible model baseline. Define an initial evaluation suite with a path to physical validation. Present a practical 12-month roadmap. A strategy document without a working baseline does not count. Six months: Train, adapt, or rigorously evaluate at least one relevant Transformer-based model or robot policy. Confirm or reject at least one important hypothesis. Establish a repeatable data-to-training-to-evaluation workflow. Test on a real robot or through a credible hardware partner. Show representative failures and what they imply. Twelve months: Demonstrate a measurable model or evaluation improvement attributable to Miraxis data, methods, or assurance systems. Close the loop between model failures, data decisions, retraining, and re-evaluation. Deliver at least one model, benchmark, evaluation system, or deployment a customer or partner can use. Build a small team that can run the work without extra process layers. We judge success by the quality of decisions, reproducibility of evidence, real-system results, and value created. Not by team size, paper count, parameter count, or number of experiments. Required qualifications Deep, direct Transformer expertise: This is a hard requirement. You must have personally made material architecture or training decisions in one or more substantial Transformer-based systems, such as vision Transformers, VLMs, VLAs, multimodal foundation models, Decision Transformers, diffusion Transformers, video Transformers, world models, or Transformer-based perception, planning, or control. You should be able to explain, in detail, how you represented inputs and outputs, tokenized data, fused modalities, structured attention and temporal context, chose losses, built the data mixture, distributed and monitored training, diagnosed failures, and what you personally changed and why it affected task performance. Using hosted model APIs, prompting language models, or running an unchanged public training recipe does not meet this requirement. Physical AI or embodied-system experience: You have worked on machine learning for a system that perceives or acts in the physical world: robotics, autonomous vehicles, drones, industrial automation, manipulation, mobile robots, humanoids, wearable or egocentric systems, or similar. At least one substantial project progressed beyond an offline dataset into a real or operational physical system. Simulation-only experience is not sufficient unless you also owned credible sim-to-real transfer and physical validation. Hands-on technical ability: You write and review high-quality Python. You work directly with PyTorch or an equivalent framework. You can read unfamiliar model and training code, design and debug experiments, inspect failures and rollout traces, work with large multimodal datasets, and make informed decisions about compute, memory, latency, and cost. Technical leadership: You have set direction for important technical work: a research or engineering group, a major model or robotics programme, or a cross-functional project you owned. You have hired or mentored researchers or engineers, made architecture and resource decisions, stopped weak work, and taken a result into deployment. A management title is not required. Direct ownership is. Evaluation, communication, and judgment: You design held-out tests, run controlled comparisons, find leakage and confounds, and report results with appropriate uncertainty. You separate evidence from belief, state what remains unknown, and reduce broad research questions to testable work. Strong preference Computer Vision We prioritize candidates who combine deep Transformer expertise with deep computer-vision experience: visual representation learning, VLMs, video and temporal modeling, detection, segmentation, tracking, multi-view and egocentric vision, 3D and spatial reasoning, pose, calibration, localization, sensor fusion, or real-time vision systems. Computer vision is not an absolute gate if the candidate has exceptional Transformer and Physical AI depth. Candidates with both receive priority. Strengthens an application VLA or robot foundation model work. Action tokenization or continuous action generation. Diffusion or flow-matching policies. Imitation learning or RL. World models. Cross-embodiment training. Manipulation. Distributed training and inference optimization. ROS 2. Sim-to-real. Safety-critical systems. Owned research publications or widely used open-source work. Early-stage company experience. Direct work with technical customers or research partners. Education A PhD in machine learning, computer vision, robotics, computer science or a related field is valuable but not required. An equivalent record of original model work, technical leadership, and real-system delivery is sufficient. We care more about what you personally designed, trained, tested, and deployed than about a specific degree or number of years.