Site Reliability Engineer ensuring reliability and scalability of TEG's global live entertainment platforms. Collaborating with teams to enhance system reliability and prevent outages across ticketing platforms.
Responsibilities
Proactively guard the health, availability, and performance of TEG's critical global production systems.
Engineer and automate robust monitoring and auto-healing solutions to proactively prevent outages and meet service level objectives (SLOs).
Drive Infrastructure-as-Code (IaC) principles for provisioning and deploying our highly available, scalable platforms.
Lead critical incident response efforts, ensuring rapid resolution and restoration of platform stability
Provide technical leadership during major incidents, focusing on swift problem analysis and effective communication to stakeholders.
Transform incidents into progress by conducting deep post-mortems and driving the implementation of strategic preventative measures across various teams.
Build and maintain high-performing, fault-tolerant distributed systems emphasizing resiliency and efficiency.
Elevate operational maturity by continuously improving processes, tooling, and efficiency across the department.
Champion operational excellence and shared responsibility, collaborating with development and other teams to improve processes and tools.
Innovate system design by evaluating and integrating new technologies to enhance reliability, scalability, and security.
Mentor and coach colleagues, elevating the overall reliability engineering capability and maturity of the Technology department
Requirements
Mastery of highly available, fault-tolerant AWS system design and management.
Strong foundation in AWS networking (VPC, Route 53) and security best practices.
Proficiency in key scripting languages (Python, Bash, PowerShell) for automation.
Proven ability to perform effectively under pressure, managing high-volume tasks and meeting tight deadlines
Minimum of 3 years of prior SRE or DevOps experience.
Expert knowledge of fundamental infrastructure concepts (Networking, Containerisation, Virtualisation, DNS)
Working familiarity with key CI/CD and Infrastructure-as-Code tools (e.g., Terraform, Ansible, Jenkins)
DevOps Engineer responsible for building and maintaining scalable AI systems on Azure cloud. Collaborating with teams to ensure operational excellence for enterprise - grade AI solutions.
Junior MLOps Engineer helping to design and maintain AI/ML systems at Bupa. Collaborating with teams to operationalize machine learning models and automate workflows.
DevOps Engineer developing and managing scalable AWS infrastructures for a PropTech startup. Collaborating within a growing tech team to achieve ambitious goals in the legal conveyancing space.
Senior DevOps Engineer leading the design and optimization of cloud infrastructure at Growth Acceleration Partners. Ensuring secure and cost - effective deployments within fast - paced product development environment.
Advanced Dev Ops Engineer optimizing infrastructure solutions for engineering teams at a consulting and technology services company. Ensuring secure and cost - effective deployments in a fast - paced environment.
Entry - level DevOps Engineer at Nokia focusing on building and maintaining CI environment for LTE and 5G solutions. Engage with high - end telecommunication technologies and support development workflows.
AI Security Control Developer/Site Reliability Engineer for RBC's enterprise AI ecosystem. Design, implement, and validate security controls to protect AI systems with 24/7 reliability.
Senior Site Reliability Engineer ensuring scalability and reliability for NGINX systems and SaaS platforms. Collaborating across teams to drive automation and system performance.
Site Reliability Engineer ensuring reliability and performance of data platform services for Veepee. Collaborating on cloud migration, Kubernetes operations, and observability best practices.
Senior Lead Site Reliability Engineer overseeing critical systems stability and incident management. Leading Java applications reliability and supporting a dynamic technology environment.