Site Reliability Engineer 3

jobgether All jobs
India
3 hour(s) ago Remote
Job Overview
Company jobgether
WorkplaceRemote
Job Typefulltime
CategorySecurity & IT – IT
First Seen 3 hour(s) ago

Job Description

This position is listed on behalf of a partner company who manages all applications and next steps. Our partner is looking for a Site Reliability Engineer 3 based in India. We are seeking a Site Reliability Engineer 3 to help modernize reliability engineering through automation observability and AI-powered operations. This role focuses on building resilient cloud platforms improving system performance and reducing operational complexity through intelligent solutions. You will work at the intersection of SRE practices AIOps automation and artificial intelligence to improve incident response and service reliability. The position offers the opportunity to design scalable operational workflows implement advanced monitoring strategies and influence how production systems are managed. You will collaborate with engineering teams to embed reliability into the software lifecycle while driving measurable improvements in availability and efficiency. This is an ideal opportunity for an experienced reliability engineer who enjoys solving complex infrastructure challenges and leveraging emerging technologies. ➡ Accountabilities ➡ The Site Reliability Engineer 3 will be responsible for ensuring the reliability scalability and performance of production systems while implementing modern AIOps capabilities. The role requires strong technical expertise operational ownership and the ability to use automation and AI-driven approaches to improve engineering efficiency and customer experience. Provide end-to-end reliability ownership for production systems including on-call support incident response postmortems and SLO/SLI management. Build and maintain observability solutions across metrics logs and traces with intelligent monitoring and anomaly detection capabilities. Design and implement AIOps workflows for event ingestion enrichment correlation deduplication and alert noise reduction. Develop AI-assisted operational processes including automated log analysis incident summaries root-cause analysis support and runbook generation. Create automation solutions that reduce operational effort through AI-enhanced runbooks controlled remediation workflows and self-healing capabilities. Implement safeguards approval processes rollback strategies and audit controls for automated operational actions. Own and improve observability platforms including ELK/OpenSearch environments for logging indexing querying dashboards and alerting. Enhance incident response processes through AI-assisted diagnosis timeline reconstruction and operational intelligence. Build and maintain integrations between observability platforms ticketing systems ChatOps tools CMDBs and automation frameworks. Partner with software engineering teams to improve deployment reliability operational readiness and production stability. Drive improvements in system resilience performance scalability and infrastructure efficiency. Maintain high-quality documentation runbooks and knowledge resources to support effective operations. Support capacity planning performance optimization and reliability engineering initiatives. Evaluate AI-generated recommendations critically and ensure appropriate human oversight for high-impact operational decisions. Measure and improve AIOps effectiveness through operational metrics such as alert reduction faster incident resolution and automation adoption. Requirements The ideal candidate is an experienced Site Reliability Engineer with strong cloud infrastructure expertise and hands-on experience applying automation and AI technologies to production environments. You should have a strong problem-solving mindset excellent operational judgment and the ability to collaborate across engineering teams. 6+ years of experience in Site Reliability Engineering AIOps DevOps or production engineering roles within large-scale cloud environments. Strong knowledge of Linux/Unix systems networking distributed systems and cloud platforms such as AWS Azure or Google Cloud. Expert-level experience with ELK/OpenSearch including log ingestion indexing scaling querying dashboards and alerting. Strong understanding of observability practices across logs metrics and distributed tracing. Experience preparing telemetry data for AIOps implementation including metadata enrichment service mapping normalization and correlation. Hands-on experience implementing AIOps solutions such as anomaly detection alert correlation intelligent alerting and automated incident enrichment. Experience integrating AI capabilities into SRE workflows including LLM-based incident analysis operational assistants or automated troubleshooting processes. Understanding of prompt engineering concepts for operational use cases including structured outputs tool usage and retrieval-augmented context. Ability to evaluate AI-generated recommendations critically and identify risks such as hallucinations or incorrect remediation guidance. Familiarity with agent-based AI frameworks or integrations connecting AI systems with operational tools is a plus. Strong knowledge of incident management root-cause analysis SLOs reliability practices and production operations. Experience with Infrastructure as Code tools such as Terraform Ansible or similar technologies. Ability to work independently manage priorities and collaborate effectively in a remote global environment. Relevant certifications in cloud DevOps Kubernetes AI/ML AIOps or observability are preferred. Benefits Competitive compensation package. Fully remote work opportunity from India. Opportunity to work with advanced cloud automation and AI-driven technologies. Exposure to large-scale production environments and complex reliability challenges. Collaboration with globally distributed engineering teams. Professional growth opportunities through innovative technology initiatives. Inclusive and collaborative remote-first work culture. Supportive environment focused on continuous learning knowledge sharing and employee development. Opportunities to contribute to technology solutions that create meaningful real-world impact. ➡ How Jobgether works We use an AI-powered matching process to ensure your application is reviewed quickly objectively and fairly against the role's core requirements. Our system identifies the top-fitting candidates and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews assessments) are managed by their internal team. We appreciate your interest and wish you the best!  Why Apply Through Jobgether?    Data Privacy Notice By submitting your application you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access rectification erasure objection) at any time.     #LI-CL1

Unlock: Sign Up for free / Sign In and use the searches from your home page or the links in the footer.