Computer Vision Engineer

microTECH Global LTD

London Area, United KingdomfulltimeSoftware Development and Artificial Intelligenceposted 16 Aug
Unlock apply linkApply links and the original listing are a Pro feature: £4.99/mo or £25 once.
Backed by an international research team and abundant computing resources, the center focuses on core research directions including multimodal understanding and generation, vision-language large models, and embodied intelligence. This is a permanent, full-time position located in the tech hub of King's Cross, London. Key Responsibilities: Frontier Technical Breakthroughs -Develop ViT and multimodal large model architectures with improved reasoning and efficiency -Advance multimodal alignment, representation learning, and long-context modeling -Explore scalable training methods for large multimodal models -Optimize model architectures for generalization and performance Data Ecosystem Construction -Process large-scale multimodal data across images, videos, audio, and text -Build pipelines for data cleaning, filtering, annotation, and quality control -Construct and maintain datasets with versioning and reproducibility -Optimize data mixtures and sampling strategies for model training -Improve data quality through feedback-driven curation loops MLLM Systems \& Infrastructure -Build distributed training systems for large-scale multimodal models -Optimize GPU utilization, cluster efficiency, and resource scheduling -Develop open-source training frameworks for scalable model development -Engineer training, inference, and serving infrastructure -Improve scalability, stability, and performance of model systems Business Value Delivery -Integrate multimodal capabilities into assistant and content generation scenarios -Translate research into production and user-facing applications -Collaborate with product and engineering teams to deploy and iterate models Person Specification: Essential Requirements: -Academic Background: Bachelor’s degree or above in Computer Science, Mathematics, Statistics, or related technical disciplines. -Technical Skills: Proficient in Python programming with strong hands-on experience in PyTorch and deep learning frameworks. -Core Competencies: Strong algorithm development and implementation skills, solid mathematical and logical reasoning ability, and excellent cross-functional communication and collaboration skills. -Traits: Self-driven and highly motivated toward advancing artificial intelligence (AI), with strong resilience and the ability to tackle challenging technical problems. Desired: -Strong track record of publications in top-tier AI or computer vision conferences (CVPR, ICCV, ECCV, NeurIPS, ICML, ICLR) -Hands-on experience in large-scale model pre-training or fine-tuning -High-impact open-source projects or internship experience in leading technology companies within CV, NLP, or multimodal domains