Graphcore

AI compute

InfrastructureandMLOpsEngineer

Bristol, United Kingdom
Market Sentiment
HIGH DEMAND

Neural analysis suggests this role is
optimal for Mid+ candidates.

The Brief

“Infrastructure and MLOps Engineer at Graphcore. Skills: Infrastructure and MLOps. Scaling and managing infrastructure. Develop essential tools and services”

What You'll Achieve.

Scaling and managing our infrastructure; Enhance the build, test, deployment, and productisation processes of our Machine Learning Software components; Eliminate toil wherever possible

Industry & Context.

AI compute
Problems you'll solve

Eliminate toil wherever possible; Solve the toughest problems

What They're Looking For.

Must Have

Knowledge of Python, Familiarity with cloud services (e. g. AWS), Experience managing or developing in Linux environments, Understanding of CI/CD principles, Experience using Kubernetes (k8s), Experience of one of the following: maintaining machine learning applications, Experience of one of the following: deploying ML orchestration tools (e. g. NV Ray, KFP, SkyPilot), Experience of one of the following: managing ML accelerator hardware (e. g. DCGM)

Nice to Have

Experience with Infrastructure as Code (IaC) tools (e. g. Terraform/OpenTofu), Experience with GitHub Actions, Experience with modern observability tooling (e. g. Prometheus), Experience with Grafana, Knowledge of Go/Java/C++ (or similar language)

What You'll Do.

Scaling and managing infrastructure

Develop essential tools and services

and productisation processes of Machine Learning Software components

Work with High-Performance Computing (HPC) AI platforms

Gain invaluable experience in distributed systems

Manage the CI platform and services

Component integration

Packaging and release systems

and maintain tools and services to support AI research and engineering teams

Deploy and maintain services with Kubernetes and Docker

Manage Cloud Infrastructure using tools such as Terraform

How You'll Work.

Team & Collaboration

Empower broader software team; Operate in squads; Fostering a culture of service ownership and empowerment

Full Job Description

About Graphcore At Graphcore, we’re building the future of AI compute. We’re a team of semiconductor, software and AI experts, with deep experience in creating the complete AI compute stack - from silicon and software to infrastructure at datacenter scale. As part of the SoftBank Group, backed by significant long-term investment, we are delivering key technology into the fast-growing SoftBank AI ecosystem.To meet the vast and exciting AI opportunity, Graphcore is expanding its teams around the world.We are bringing together the brightest minds to solve the toughest problems, in a place where everyone has the opportunity to make an impact on the company, our products and the future of artificial intelligence. Job Summary Join our dynamic Software Infrastructure team and take a pivotal role in scaling and managing our infrastructure. You will develop essential tools and services that empower our broader software team. Your contributions will enhance the build, test, deployment, and productisation processes of our Machine Learning Software components. Work with our High-Performance Computing (HPC) AI platforms and gain invaluable experience in distributed systems The Team The Software Infrastructure team provides critical platforms and services for software development teams across the business. Our responsibilities include managing the CI platform and services, build engineering, component integration, and packaging and release systems. We operate in squads, fostering a culture of service ownership and empowerment for our engineers. We focus on long-term engineering solutions and strive to eliminate toil wherever possible. Responsibilities and Duties Develop, own, and maintain tools and services to support AI research and engineering teams Deploy and maintain services with Kubernetes and Docker Manage our Cloud Infrastructure using tools such as Terraform Candidate Profile Essential: Knowledge of Python Familiarity with cloud services (e. g. AWS) Experience managing o

Free ATS check

Applying for this Infrastructure and MLOps Engineer role?

Most applicants get filtered before a human reads their resume. See if yours makes the cut.

How to Apply on Greenhouse

  • Create a Greenhouse profile before applying — it saves time across multiple applications.
  • Upload your resume as a PDF; the parser handles it better than Word.
  • Answer all knockout questions carefully — wrong answers auto-reject before a human sees you.
  • Enable email notifications to track application status in real time.

ANONYMOUS · UNFILTERED

What do employees actually say about Graphcore?

Real rants from real employees. Read before you apply.

Read Company Rants →