NVIDIA

AI computing

SeniorDeepLearningSoftwareEngineer

$224–357k Santa Clara, California, United States FULL TIME Remote Friendly

Market Sentiment

HIGH DEMAND

Neural analysis suggests this role is
optimal for Senior candidates.

The Brief

“Senior Deep Learning Software Engineer at NVIDIA. Skills: Deep Learning, Inference and Deployment Solutions, GPU Optimization, Software Engineering. design and build automated inference and deployment solution. defining a scalable architecture for DL inference with emphasis on ease-of-use and compute efficiency”

What You'll Achieve.

ensure NVIDIA's inference software solutions (TRT, TRT-LLM, TRT Model Optimizer) can maintain and increase its leadership in the market; help build real-time, cost-effective computing platforms driving our success

Industry & Context.

AI computing

Problems you'll solve

performance analysis; debugging; understanding and debugging end-to-end performance

What They're Looking For.

Must Have

Masters, PhD, or equivalent experience in Computer Science, AI, Applied Math, or related field, 8+ years of relevant work or research experience in Deep Learning, Excellent software design skills, including debugging, performance analysis, and test design, proficiency in Python, PyTorch, and related ML tools, algorithms and programming fundamentals, Good written and verbal communication skills and the ability to work independently and collaboratively in a fast-paced environment

Nice to Have

Contributions to PyTorch, JAX, or other Machine Learning Frameworks, Knowledge of GPU architecture and compilation stack, and capability of understanding and debugging end-to-end performance, Familiarity with NVIDIA's deep learning SDKs such as TensorRT, Prior experience in writing high-performance GPU kernels for machine learning workloads in frameworks such as CUDA, CUTLASS, or Triton

What You'll Do.

design and build automated inference and deployment solution, defining a scalable architecture for DL inference with emphasis on ease-of-use and compute efficiency, developing features in high-level frameworks like PyTorch and JAX, designing and implementing a high-performance execution environment, low-level GPU optimizations, developing custom GPU kernels in CUDA and/or Triton, defining of a modular, scalable platform to seamlessly bridge training and deployment workflows, enabling tight integration of deployment tooling with training frameworks such as Megatron and Nemo, Leverage and build upon the torch 2.

0 ecosystem (TorchDynamo, torch.

export, torch.

compile, etc.

) to analyze and extract standardized model graph representation from arbitrary torch models for our automated deployment solution, Develop support for inference optimization techniques such as speculative decoding and LoRA, Collaborate with teams across NVIDIA to use performant kernel implementations within the automated deployment solution, Analyze and profile GPU kernel-level performance to identify hardware and software optimization opportunities, Continuously innovate on the inference performance to ensure NVIDIA's inference software solutions (TRT, TRT-LLM, TRT Model Optimizer) can maintain and increase its leadership in the market, help build real-time, cost-effective computing platforms.

How You'll Work.

Team & Collaboration

Collaborate with teams across NVIDIA; work collaboratively in a fast-paced environment

Communication Scope

Good written and verbal communication skills

Full Job Description

We are looking for a Senior Deep Learning Software Engineer to design and build our automated inference and deployment solution. As part of the team, you will be instrumental in defining a scalable architecture for DL inference with emphasis on ease-of-use and compute efficiency. Your work will span multiple layers of the DL deployment stack, encompassing developing features in high-level frameworks like PyTorch and JAX, designing and implementing a high-performance execution environment, low-level GPU optimizations and developing custom GPU kernels in CUDA and/or Triton. This is an exceptional opportunity for passionate software engineers straddling the boundaries of research and engineering, with a strong background in both machine learning fundamentals and software architecture & engineering. **What you’ll be doing: ** * Play a pivotal role in defining of a modular, scalable platform to seamlessly bridge training and deployment workflows—enabling tight integration of deployment tooling with training frameworks such as Megatron and Nemo * Leverage and build upon the torch 2.0 ecosystem (TorchDynamo, torch.export, torch.compile, etc...) to analyze and extract standardized model graph representation from arbitrary torch models for our automated deployment solution. * Develop support for inference optimization techniques such as speculative decoding and LoRA. * Collaborate with teams across NVIDIA to use performant kernel implementations within the automated deployment solution. * Analyze and profile GPU kernel-level performance to identify hardware and software optimization opportunities. * Continuously innovate on the inference performance to ensure NVIDIA's inference software solutions (TRT, TRT-LLM, TRT Model Optimizer) can maintain and increase its leadership in the market. **What we need to see: ** * Masters, PhD, or equivalent experience in Computer Science, AI, Applied Math, or related field. * 8+ years of relevant work or research experience in Deep Learning

Free ATS check

Applying for this Senior Deep Learning Software Engineer role?

Most applicants get filtered before a human reads their resume. See if yours makes the cut.

Should you apply? AI reads your resume vs this job — match score, gaps to address, ATS keywords.

SKILL SIGNAL 36 detected · ranked by frequency

Deep Learning ×3

automated inference and deployment solution design ×3

scalable architecture for DL inference ×3

compute efficiency ×3

DL deployment stack development ×3

high-performance execution environment design ×3

low-level GPU optimizations ×3

custom GPU kernel development ×3

modular, scalable platform definition ×3

training and deployment workflow bridging ×3

deployment tooling integration ×3

inference optimization techniques (speculative decoding, LoRA) ×3

performant kernel implementations ×3

GPU kernel-level performance analysis and profiling ×3

software design ×3

debugging ×3

performance analysis ×3

test design ×3

algorithms ×3

programming fundamentals ×3

Inference and Deployment Solutions ×2

GPU Optimization ×2

Software Engineering ×2

PyTorch ×2

JAX ×2

CUDA ×2

Triton ×2

TensorRT ×2

TRT-LLM ×2

TRT Model Optimizer ×2

CUTLASS ×2

TorchDynamo

BEHAVIOURAL

passionatecreativemotivatedlove a challengework independentlywork collaboratively

Role Details

Seniority senior

Experience 8–10 yrs

Level Senior

Work Mode Hybrid

Type FULL TIME

Education Masters, PhD, or equivalent experience

Salary Band 200k+

AI-Extracted Insights

Domain Areas

deep-learningmachine-learning-fundamentalsgpu-architecturecompilation-stack

How to Apply on Workday

Workday has a multi-step form — save your progress after every section.
"Apply With LinkedIn" can fail or lose data; manual entry is more reliable.
Watch for the "Submit for Review" final step — hitting "Save" alone does not submit.
Job requisition numbers are useful when following up with HR by email.

ANONYMOUS · UNFILTERED

What do employees actually say about NVIDIA?

Real rants from real employees. Read before you apply.

Read Company Rants →