Amazon Development Center U.S., Inc.

Technology

SoftwareDevelopmentEngineer,EC2UltraServerAvailability

$70–194k Seattle, Washington, United States FULL TIME
Market Sentiment
HIGH DEMAND

Neural analysis suggests this role is
optimal for Mid candidates.

The Brief

“Software Development Engineer, EC2 UltraServer Availability at Amazon Development Center U.S., Inc.. Skills: Cloud-based repair, Recovery workflows, AI/ML infrastructure. Design repair workflows. Orchestrate repair operations”

What You'll Achieve.

Ensure high availability

Industry & Context.

Technology
Problems you'll solve

Investigate solutions; Troubleshoot workflow failures; Propose solutions

What They're Looking For.

Must Have

3+ years software development experience, 2+ years system design experience, Experience programming one language

Nice to Have

3+ years full SDLC experience

What You'll Do.

Design repair workflows

Orchestrate repair operations

Orchestrate recovery operations

Build cloud-based solutions

Write maintainable code

Develop repair workflows

Develop recovery workflows

Create observable systems

Execute UltraServer workflows

Monitor UltraServer workflows

Troubleshoot workflow failures

Manage network configurations

Handle firmware validation

Perform consistency checks

Collaborate with customers

Convert business needs

How You'll Work.

Team & Collaboration

Cross-functional collaboration; Capacity Management; Hardware Engineering; Datacenter Operations; Downstream teams; Stakeholders

Process & Methodology

Requirements gathering, Design reviews, Feature launches, Continuous improvement

Full Job Description

The Software Development Engineer II will design, build, and maintain cloud-based repair and recovery workflows for NVIDIA GB200 / GB300 UltraServers, orchestrating repair and recovery operations from impairment detection through completed recovery. This role requires expertise in AWS services, system architecture, and cross-functional collaboration with Capacity Management, Hardware Engineering, and Datacenter Operations to manage AI/ML infrastructure. Key job responsibilities The Software Development Engineer (SDE II) on the EC2 UltraServer Availability team is responsible for ensuring high availability of customer GB200 and GB300 UltraServers by orchestrating complex repair and recovery workflows. Following are the core responsibilities System Design & Architecture * Design and architect solutions that are cross-functional to Capacity Management, Hardware Engineering, and Datacenter Operations * Work in environments where the technology strategy is defined but the solution design is not * Build solutions that are stable, logical, testable, and efficient with the ability to independently make trade-off decisions * Investigate and develop design concepts to frame solution sets at an application and product level Software Development * Build cloud-based solutions using AWS native services for scaling infrastructure frameworks * Write high-quality, maintainable code with proper testing and code reviews * Develop and maintain the repair and recovery workflows for GB200 and GB300 UltraServer hosts * Implement automation for diagnostic triage, hardware testing, cable validation, and testing processes * Create observable systems with appropriate metrics and alarming Operational Excellence * Execute and monitor UltraServer workflows for UltraServer repair * Troubleshoot workflow failures and coordinate with downstream teams * Focus on operational excellence by identifying problems and proposing solutions Hardware & Software Integration * Work with hardware and software in

Free ATS check

Applying for this Software Development Engineer, EC2 UltraServer Availability role?

Most applicants get filtered before a human reads their resume. See if yours makes the cut.

ANONYMOUS · UNFILTERED

What do employees actually say about Amazon Development Center U.S., Inc.?

Real rants from real employees. Read before you apply.

Read Company Rants →