Senior DevOps Service Reliability Engineer – DGX Cloud at NVIDIA | Hybrid Hired

About the role

Senior DevOps Engineer for NVIDIA's cloud products ensuring high service reliability and availability. Collaborating with cross-functional teams and handling incident management in a 24/7 follow-the-sun environment.

Responsibilities

partner with other key members including Site Reliability Engineering, Security Operations Center, DevOps teams
help make services capable of providing near 100% availability
decrease frequency and duration of any issue
develop monitors, alarms, and alerts to help make the service more reliable
report directly to a manager in the United States
provide their services 24/7 with a follow-the-sun environment
use alerts and alarms to help prevent issues and incidents when possible
work with developers to develop and implement predictive support or diagnostic routines
perform systems administration tasks, network administration tasks, security incident monitoring
develop runbooks which the entire team will use
update and evolve the runbooks as needed
discover incidents and issues, including initiating the incident management procedure
feedback will help us continually improve our service

Requirements

5+ years of experience administering large-scale production systems
3+ years of experience in high-availability Internet, Cloud, or Data Center environments (Systems Administration, SRE, or NOC)
BS in Computer Science, Engineering, Physics, Mathematics, or equivalent experience
Expert-level knowledge of Linux system administration and automation using Ansible and/or Python
Strong experience with shell scripting, DNS, DHCP, storage systems, and core networking (IP Tables, routing, firewalls)
Experience with at least one workload manager (Slurm preferred) or job scheduling system in a production environment
Strong experience troubleshooting and maintaining large-scale bare-metal infrastructure
Strong cross-team collaboration, documentation, and mentoring skills
Experience improving processes for automation, reliability, and operational excellence
Expertise using monitoring tools and problem ticketing systems
Strong problem-solving, analytical, and troubleshooting abilities

Benefits

equity
benefits

Similar roles

Browse all Devops Engineer jobs

47 minutes ago

GR

Senior DevOps Engineer

greehill

DevOps Engineer for designing and maintaining Azure - based hybrid cloud infrastructure for a company specializing in nature - based smart city solutions. Leading cloud architecture and mentoring engineers as part of a high - impact team.

Hybrid Role

Budapest Hungary Devops Engineer

6 hours ago

II

Senior Infrastructure Analyst – SRE

INFOX Tecnologia da Informação

SRE responsible for ensuring reliability and performance of IT systems at a digital transformation company specializing in public sector efficiency. Collaborating on system health, incident response, and automation tasks.

Hybrid Role

São Paulo Brazil Devops Engineer

6 hours ago

BS

DevOps Senior

Beyond Soluções

DevOps Senior role at Beyond Soluções managing CI/CD for .NET and Kubernetes applications. Collaborating on cloud solutions while fostering a culture of innovation and quality.

Hybrid Role

São Paulo Brazil Devops Engineer

11 hours ago

PA

Senior Software Engineer – Cloud Infrastructure, DevOps

PayPal

Senior Software Engineer at PayPal managing cloud infrastructure and DevOps solutions. Delivering complete SDLC solutions and guiding engineering teams for scalable and reliable services.

Hybrid Role

San Jose United States Devops Engineer

$143,500 - $212,850 per year

20 hours ago

VS

Senior Site Reliability Engineer

VALCE Talent Solutions

Senior Site Reliability Engineer at Diligent leading reliability, automation, and observability across cloud infrastructure. Build tools for incident response and enhance performance in fast - paced environments.

Hybrid Role

Guadalajara Mexico Devops Engineer

23 hours ago

CI

Perception Deployment Engineer

Caterpillar Inc.

Perception Deployment Engineer deploying deep learning models on embedded systems at Caterpillar. Collaborating with cross - functional teams for integration and optimization of perception modules in vehicles.

Onsite Role

Wuxi China Devops Engineer

23 hours ago

AT

Principal Site Reliability Engineer, SRE

AT&T

Principal Site Reliability Engineer at AT&T required to design scalable solutions for critical operations with minimal downtime. Collaborating with teams to monitor and improve system performance in cloud environments.

Onsite Role

Plano United States Devops Engineer

$174,100 - $261,100 per year

yesterday

CO

DevOps Engineer, AI SaaS

Coach4expats

DevOps Engineer managing AI SaaS infrastructure at a high - growth European company. Supporting AI model deployment and ensuring platform security and compliance with multiple systems integration.

Hybrid Role

Europe Devops Engineer

yesterday

LE

Observability & DevOps Tools Engineering Manager

LexisNexis

Engineering Manager leading teams for observability platforms at LexisNexis. Owns operational excellence across software delivery lifecycle in Raleigh, NC.

Hybrid Role

Raleigh United States Devops Engineer

$118,300 - $219,800 per year

yesterday

RO

Reliability Engineer

Roche

Reliability Engineer optimizing site facility infrastructure and utility systems at Roche. Conducting root cause analyses and developing maintenance plans to enhance reliability and efficiency.

Onsite Role

Santa Clara United States Devops Engineer

$76,900 - $142,700 per year