Site Reliability Engineer at WRITER, ensuring 24/7 availability and performance of AI-powered workflows. Collaborating on scalable infrastructure solutions while impacting enterprise customer trust.
Responsibilities
Use and build AI native approaches for operational tasks and infrastructure management and platforms using Python, Go, or similar languages, significantly reducing manual toil across our production environment
Design and implement scalable, fault-tolerant infrastructure AI solutions on public cloud providers (AWS, GCP, Azure) to support WRITER's rapidly expanding, high-traffic AI platform
Own the reliability, performance, and efficiency of WRITER’s core services, defining and upholding stringent Service Level Objectives (SLOs) and Error Budgets
Own the observability stack for monitoring, logging, and alerting systems to ensure rapid detection of issues across our complex distributed systems
Lead incident response, post-mortems, and root cause analyses, applying learnings to proactively prevent future outages and build a more resilient system architecture
Collaborate closely with product and engineering teams, providing expert guidance on system design for reliability, performance, and scalability from conception through launch
Requirements
A solid 7+ years of experience in Site reliability engineering, DevOps, Production engineering, Cloud platform or a similar role focused on building and operating large-scale, high-availability production systems
Deep expertise with cloud platforms (AWS strongly preferred), containerization technologies like Docker and Kubernetes, and Infrastructure-as-Code tools such as Terraform
Strong proficiency in programming languages such as Python, Java, Go for automation and monitoring
Knowledge of monitoring and logging tools (e.g., Prometheus, Grafana, ELK Stack) to maintain system health and performance
Demonstrated ability to Challenge the status quo, proactively identify systemic weaknesses, and propose innovative solutions to complex reliability problems
Excellent communication, collaboration, and problem-solving skills, with a talent for building strong relationships and Connecting with cross-functional teams
A strong sense of ownership and accountability, eager to Own mission-critical systems and drive them toward peak performance and unparalleled reliability
Benefits
Generous PTO, plus company holidays
Medical, dental, and vision coverage for you and your family
Paid parental leave for all parents (12 weeks)
Fertility and family planning support
Early-detection cancer testing through Galleri
Flexible spending account and dependent FSA options
Health savings account for eligible plans with company contribution
Annual work-life stipends for:
Wellness stipend for gym, massage/chiropractor, personal training, etc.
Learning and development stipend
Company-wide off-sites and team off-sites
Competitive compensation, company stock options and 401k
DevOps/IT Apprentice supporting cloud infrastructure and CI/CD pipelines at tech startup. Involves learning, taking ownership, and growing within the engineering team.
DevOps Engineer at Cloud++ collaborating on infrastructure and CI/CD pipelines across multi - cloud environments. Engaging with development teams to ensure reliable and secure releases.
Fullstack Developer at Zenika engaging in impactful tech projects like B2B platforms and architecture modernization. Collaborating with senior consultants in a quality - focused environment.
Production Engineer in a hybrid role ensuring operational performance of applications for a strategic international project. Focusing on automation and optimization within a technical environment at EOLEN.
Senior/Expert DevOps Engineer for AI project in pharmaceutical sector at GECI International. Involves designing, deploying, and operating autonomous AI agents.
DevOps Engineer Intern at Emeria Technologies focusing on Cloud infrastructure design and support. Involves maintaining CI/CD platforms and collaborating with DevOps teams for optimization.
Intern supporting software development infrastructure including CI/CD and cloud integration at Intel. Collaborating with teams to optimize development and release processes.
DevOps Engineer at NetBrain responsible for AWS cloud infrastructure and automating processes. Collaborate with development and security teams to deliver secure solutions efficiently.
Site Reliability Engineer at BlaBlaCar improving CI/CD and tooling for developer efficiency and autonomy. Collaborating with engineering teams to enhance service reliability and facilitate software development.