We're looking for a Senior Site Reliability Engineer (SRE) to join our team in Spain in a remote working mode. In this role you will collaborate with development operations security and quality teams to ensure highly reliable scalable and efficient systems for business-critical applications in the financial domain. You will focus on implementing SRE practices reducing toil through automation and driving operational excellence while meeting strict Service Level Objectives (SLOs). This position offers the opportunity to influence system design for reliability and performance within a global delivery context leveraging modern cloud technologies observability tools and automation frameworks to maintain seamless user experiences. Responsibilities Define and maintain Service Level Objectives (SLOs) SLIs and error budgets for critical services Collaborate with cross-functional teams to embed reliability into application and infrastructure design Automate operational tasks to reduce manual toil and improve service performance Troubleshoot and resolve infrastructure and application incidents quickly and effectively Implement robust monitoring and observability systems to detect and prevent outages Plan capacity and scaling strategies to ensure high availability and resiliency Contribute to incident postmortems and continuous improvement initiatives Support the adoption of SRE best practices across all SDLC stages Requirements Bachelor’s degree in Computer Science Engineering or related field Proven experience working in cloud environments (AWS GCP or Azure) Practical knowledge of SRE principles (SLO/SLI design error budgets postmortems automation) Proficiency in Python or other scripting language for automation tasks Strong understanding of monitoring tools and observability frameworks Experience with Infrastructure-as-Code and CI/CD tools (e.g. Terraform Ansible Jenkins GitLab) Hands-on expertise with containerization and orchestration platforms such as Docker and Kubernetes Nice to have Experience deploying and managing Large Language Models (LLMs) including RAG-based solutions Certifications in Kubernetes AWS/GCP/Azure or related cloud technologies Background in DevOps practices and agile delivery frameworks Familiarity with AI/ML model operations deployment monitoring and optimization in production environments We offer Private health insurance EPAM Employees Stock Purchase Plan 100% paid sick leave Referral Program Professional certification Language courses EPAM is a leading digital transformation services and product engineering company with 61700+ EPAMers in 55+ countries and regions. Since 1993 our multidisciplinary teams have been helping make the future real for our clients and communities around the world. In 2018 we opened an office in Spain that quickly grew to over 1450 EPAMers distributed between the offices in Málaga Madrid and Cáceres as well as remotely across the country. Here you will collaborate with multinational teams contribute to numerous innovative projects and have an opportunity to learn and grow continuously. Why Join EPAM WORK AND LIFE BALANCE. Enjoy more of your personal time with flexible work options 24 working days of annual leave and paid time off for numerous public holidays. CONTINUOUS LEARNING CULTURE. Craft your personal Career Development Plan to align with your learning objectives. Take advantage of internal training mentorship sponsored certifications and LinkedIn courses. CLEAR AND DIFFERENT CAREER PATHS. Grow in engineering or managerial direction to become a People Manager in-depth technical specialist Solution Architect or Project/Delivery Manager. STRONG PROFESSIONAL COMMUNITY. Join a global EPAM community of highly skilled experts and connect with them to solve challenges exchange ideas share expertise and make friends.