Together AI

JuniorTechnicalProgramManagerInfrastructureOperations

$150–175k San Francisco, California, United States
Market Sentiment
HIGH DEMAND

Neural analysis suggests this role is
optimal for Entry candidates.

The Brief

“Junior Technical Program Manager — Infrastructure Operations at Together AI. Skills: process management, cross-functional coordination, vendor/stakeholder management, operational metrics, resource planning, continuous improvement, end-to-end node lifecycle management, datacenter bring-ups, GPU utilization loss identification, dashboard building, tracking process development, workflow improvement, lightweight automation. Own the end-to-end node lifecycle - from failure through repair, return, and”

Industry & Context.

Problems you'll solve

hunt down GPU utilization loss; eliminate gaps in ownership at every handoff; drive resolution

What They're Looking For.

Must Have

Some prior experience in a TPM role, A technical background or demonstrated experience in a highly technical environment, A genuine bias toward action, Resilience in a fast-paced, sometimes chaotic environment, organizational instincts, Ability to zoom out

What You'll Do.

Own the end-to-end node lifecycle - from failure through repair

and re-integration — across provider ticketing

and the state machine that governs each stage

Drive node remediation to resolution with urgency

eliminating gaps in ownership at every handoff

Manage project timelines for new datacenter bring-ups

coordinating across internal teams and external providers to keep milestones on track

Identify and diagnose GPU utilization loss across the fleet

working with engineering leads to drive resolution

Build dashboards and tracking processes that make efficiency gaps visible and ensure they get closed

Continuously improve operational workflows through process improvements and lightweight automation

How You'll Work.

Team & Collaboration

cross-functional coordination; coordinating across internal teams and external providers; working with engineering leads

Process & Methodology

Manage project timelines for new datacenter bring-ups

Full Job Description

About the Role Together AI runs one of the most demanding GPU fleets in the industry. Keeping that fleet healthy - every node online, every GPU performing, every datacenter transition running on schedule - is operationally complex and genuinely high-stakes. We're looking for a Junior TPM to own that operational reality. This is not a coordination or status-reporting role. You will own the end-to-end node lifecycle - from the moment a node goes down through repair, return, and re-integration - and you'll drive the cross-functional work to close every gap as fast as possible. You'll manage datacenter bring-ups, hunt down GPU utilization loss, and build the processes and dashboards that make our fleet operations more visible and accountable over time. The environment moves fast and doesn't always come with a clear playbook. Much of what you'll work on is genuinely novel - you'll be figuring things out alongside engineers who are building at the frontier. If that sounds like an obstacle, this isn't the right role. If it sounds like the best possible way to learn, keep reading. Responsibilities Own the end-to-end node lifecycle - from failure through repair, return, and re-integration — across provider ticketing, internal tooling, and the state machine that governs each stage Drive node remediation to resolution with urgency, eliminating gaps in ownership at every handoff Manage project timelines for new datacenter bring-ups, coordinating across internal teams and external providers to keep milestones on track Identify and diagnose GPU utilization loss across the fleet, working with engineering leads to drive resolution Build dashboards and tracking processes that make efficiency gaps visible and ensure they get closed Continuously improve operational workflows through process improvements and lightweight automation Develop and maintain relationships with external datacenter providers Requirements Some prior experience in a TPM role - we're open to candidates who came in

Free ATS check

Applying for this Junior Technical Program Manager — Infrastructure Operations role?

Most applicants get filtered before a human reads their resume. See if yours makes the cut.

How to Apply on Greenhouse

  • Create a Greenhouse profile before applying — it saves time across multiple applications.
  • Upload your resume as a PDF; the parser handles it better than Word.
  • Answer all knockout questions carefully — wrong answers auto-reject before a human sees you.
  • Enable email notifications to track application status in real time.

ANONYMOUS · UNFILTERED

What do employees actually say about Together AI?

Real rants from real employees. Read before you apply.

Read Company Rants →