Wayve

StaffMLPerformanceEngineer(InferenceOptimisation)

London, England, United Kingdom FULL TIME

Market Sentiment

HIGH DEMAND

Neural analysis suggests this role is
optimal for Staff candidates.

The Brief

“Staff ML Performance Engineer (Inference Optimisation) at Wayve. Skills: ML inference optimisation, edge deployment, performance engineering. Optimising ML inference for edge accelerators and GPUs. Running large transformer-based models efficiently on low-cost, low-power edge devices”

What You'll Achieve.

deliver measurable improvements; ensure performance improvements hold across models, devices, and software releases; accelerating the transition from assisted to automated driving; propels the world forward

Industry & Context.

Problems you'll solve

lean into complex challenges to unlock groundbreaking solutions

What They're Looking For.

Must Have

Proven experience improving performance in production systems with tight constraints (latency, memory, bandwidth, power/thermal, or cost), Proficiency with at least one relevant stack/toolchain (e.g. TensorRT, CUDA, Qualcomm QNN, Triton, OpenCL), Comfort operating at multiple levels of abstraction — from high-level model behaviour down to low-level kernel/runtime execution, Software engineering fundamentals (debugging, profiling, testing, and maintainable code), Clear communicator and collaborative able to align multiple stakeholders on performance trade-offs and priorities

Nice to Have

Exposure to embedded or edge deployment of ML models, including benchmarking on real devices and handling system-level constraints, Experience with NVIDIA and/or Qualcomm SoCs and performance tooling, Python and C++ proficiency, Experience mentoring others and/or driving technical direction in a small, fast-moving team

What You'll Do.

Optimising ML inference for edge accelerators and GPUs, Running large transformer-based models efficiently on low-cost, low-power edge devices, Turning models into production systems that run reliably on in-vehicle compute, Profile and pinpoint bottlenecks across the full inference stack (model graph, compiler/runtime, kernel execution, memory movement) and deliver measurable improvements, Implement and validate optimisations in compilers, runtimes, and/or kernels (e.

operator fusion, scheduling, quantisation-aware performance, custom kernels), Build robust benchmarking and regression testing to ensure performance improvements hold across models, devices, and software releases, Optimise for multiple targets (e.

NVIDIA Orin/Thor, Qualcomm) and work with teams to support these in a maintainable way, Collaborate with model developers to influence architecture and training/deployment decisions that affect on-device performance, Contribute to technical roadmaps and tooling and help raise the standard of performance engineering across the team.

How You'll Work.

Team & Collaboration

Collaborate with model developers to influence architecture and training/deployment decisions; Work with teams to support multiple targets in a maintainable way; Align multiple stakeholders on performance trade-offs and priorities

Communication Scope

Clear communicator; able to align multiple stakeholders on performance trade-offs and priorities

Full Job Description

About us Founded in 2017, Wayve is the leading developer of Embodied AI technology. Our advanced AI software and foundation models enable vehicles to perceive, understand, and navigate any complex environment, enhancing the usability and safety of automated driving systems. Our vision is to create autonomy that propels the world forward. Our intelligent, mapless, and hardware-agnostic AI products are designed for automakers, accelerating the transition from assisted to automated driving. In our fast-paced environment big problems ignite us—we embrace uncertainty, leaning into complex challenges to unlock groundbreaking solutions. We aim high and stay humble in our pursuit of excellence, constantly learning and evolving as we pave the way for a smarter, safer future. At Wayve, your contributions matter. We value diversity, embrace new perspectives, and foster an inclusive work environment; we back each other to deliver impact. Make Wayve the experience that defines your career! The role As a Staff ML Performance Engineer, you’ll play a key role in high-impact projects, optimising ML inference for edge accelerators and GPUs. The focus of this team is to run large transformer-based models efficiently on low-cost, low-power edge devices to enable Wayve’s first driving product. You’ll help set the technical direction for turning these models into production systems that run reliably on in-vehicle compute. This is a hands-on role working across ML systems, compilers, runtimes, kernels, and embedded deployment, contributing to several early-stage, high-impact projects at Wayve. Key responsibilities: Profile and pinpoint bottlenecks across the full inference stack (model graph, compiler/runtime, kernel execution, memory movement) and deliver measurable improvements. Implement and validate optimisations in compilers, runtimes, and/or kernels (e. g. operator fusion, scheduling, quantisation-aware performance, custom kernels). Build robust benchmarking and regression testing t

Free ATS check

Applying for this Staff ML Performance Engineer (Inference Optimisation) role?

Most applicants get filtered before a human reads their resume. See if yours makes the cut.

Should you apply? AI reads your resume vs this job — match score, gaps to address, ATS keywords.

SKILL SIGNAL 38 detected · ranked by frequency

performance engineering ×3

performance optimisation ×3

compiler optimisation ×3

runtime optimisation ×3

kernel optimisation ×3

operator fusion ×3

scheduling ×3

quantisation-aware performance ×3

custom kernels ×3

benchmarking ×3

regression testing ×3

debugging ×3

testing ×3

maintainable code ×3

ML inference optimisation ×2

edge deployment ×2

TensorRT ×2

CUDA ×2

Qualcomm QNN ×2

Triton ×2

OpenCL ×2

NVIDIA Orin/Thor ×2

Qualcomm ×2

ML inference

transformer-based models

edge accelerators

GPUs

ML systems

compilers

runtimes

kernels

embedded deployment

BEHAVIOURAL

collaborativeClear communicator

Role Details

Experience 8–15 yrs

Level Staff

Work Mode hybrid

Type FULL TIME

Category vehicle-sw-engineering

AI-Extracted Insights

Domain Areas

embodied-ai-technologyautomated-driving-systemsself-driving-cars

ANONYMOUS · UNFILTERED

What do employees actually say about Wayve?

Real rants from real employees. Read before you apply.

Read Company Rants →