Databricks

Data and AI

SiteReliabilityEngineer

Costa Rica
Market Sentiment
HIGH DEMAND

Neural analysis suggests this role is
optimal for Senior candidates.

The Brief

“Site Reliability Engineer at Databricks. Skills: Site Reliability Engineering, Cloud Infrastructure, Automation, Observability, CI/CD. Architect and Automate production-grade infrastructure on cloud platforms (AWS/Azure) using Infrastructure as Code (IaC) tools like Terraform or Pulumi. Optimize system performance, architecture, and scaling”

Industry & Context.

Data and AI
Problems you'll solve

Identify root causes; Implement permanent preventive engineering solutions

Eligibility Requirements

Participate in a shared on-call rotation

What They're Looking For.

Must Have

5+ years of production-level experience, Proficiency in Python (non-negotiable), Expert-level proficiency in Terraform (modules, state management) or Pulumi, Hands-on experience with AWS, Azure, or GCP, Hands-on experience with Kubernetes, Docker, and containerization concepts, Deep understanding of observability pillars (logging, metrics, tracing), Experience with tools such as Datadog, Prometheus, or ELK, Proficiency in running systems using concepts like Kafka or messaging queues, Advanced knowledge of GitHub Actions and GitHub Runners, Ability to take ownership of ambiguous projects, Follow a vision set by tech leads, Execute independently with minimal guidance

Nice to Have

Knowledge of Pulumi, Experience with GCP, Experience with Datadog, Experience with Prometheus, Experience with ELK, Experience with Kafka, Experience with messaging queues, Experience with GitHub Runners

What You'll Do.

Architect and Automate production-grade infrastructure on cloud platforms (AWS/Azure) using Infrastructure as Code (IaC) tools like Terraform or Pulumi

Optimize system performance

Architect robust deployment pipelines (e.g.

managing both hosted and self-hosted runners

Create underlying infrastructure to ensure new internal applications are secure and have logging

metrics and alerts enabled by default

Build internal AI plugins

and automation scripts

Focus on subsequent data usage

incident management workflows

and creating necessary dashboards

Participate in a shared on-call rotation

Lead rapid incident response and technical troubleshooting for production outages

Facilitate blameless post-mortems

Implement permanent preventive engineering solutions

How You'll Work.

Team & Collaboration

Collaborate with Security, Engineering, and Support teams to deliver real business outcomes

Process & Methodology

Ownership of ambiguous projects

Full Job Description

GAQ127R40 Team: IT Infrastructure and Operations About the Role At Databricks Information Technology, we are a product-led organization transforming how we work—from the ease of using our IT services to the applications we develop to scale seamlessly during rapid growth. As a Site Reliability Engineer (SRE), you will bridge the gap between software engineering and systems architecture. You will be a core contributor to the IT Infrastructure team, owning the evolution of core infrastructure and observability platforms. This role requires a strong software engineering mindset and deep technical breadth to deliver high-quality, scalable solutions for "immature" system problems. Your focus will be on building resilient, automated infrastructure that empowers development teams and ensures our cloud environment is cost-optimized, secure, and highly available. The Impact You Will Have Architect and Automate: Design and deploy production-grade infrastructure on cloud platforms (AWS/Azure) using Infrastructure as Code (IaC) tools like Terraform or Pulumi. Reliability and Performance Engineering:Optimize system performance, architecture, and scaling to ensure maximum uptime and minimal latency for critical IT services. CI/CD Excellence: Architect robust deployment pipelines (e.g., GitHub Actions), managing both hosted and self-hosted runners for specialized build requirements. Observable by Default: Create underlying infrastructure to ensure new internal applications are secure and have logging, metrics and alerts enabled by default. Agentic ToolingI: Build internal AI plugins, and automation scripts to streamline developer workflows and enhance operational efficiency. Incident Response: Focus on subsequent data usage, incident management workflows, and creating necessary dashboards to maintain service health. Participate in a shared on-call rotation, leading rapid incident response and technical troubleshooting for production outages.Facilitate blameless post-mortems to iden

Free ATS check

Applying for this Site Reliability Engineer role?

Most applicants get filtered before a human reads their resume. See if yours makes the cut.

ANONYMOUS · UNFILTERED

What do employees actually say about Databricks?

Real rants from real employees. Read before you apply.

Read Company Rants →