Responsible at the expert level for ensuring the reliability scalability performance and operational excellence of critical banking platforms and applications. Serves as a senior individual contributor responsible for designing implementing and improving Site Reliability Engineering (SRE) practices across the software development lifecycle. Works closely with application development infrastructure platform engineering and business teams to enhance system resiliency through automation observability testing and proactive operational management while coaching and influencing others. Responsibilities Design implement and support highly available scalable and resilient applications and cloud infrastructure following enterprise technology standards and SRE best practices .Lead initiatives to improve system reliability availability performance and operational maturity through automation and engineering excellence .Define implement and monitor Service Level Objectives (SLOs) Service Level Indicators (SLIs) and error budgets for critical business services .Develop comprehensive observability strategies leveraging Dynatrace OpenTelemetry (OTel) distributed tracing metrics logging dashboards and alerting solutions .Design and maintain end-to-end monitoring solutions that provide actionable insights into application infrastructure and customer experience health .Analyze production telemetry to proactively identify performance bottlenecks reliability risks and capacity constraints .Lead incident response activities for high-severity production events coordinating cross-functional teams to restore services and minimize customer impact .Perform and facilitate Root Cause Analysis (RCA) activities ensuring corrective and preventive actions are identified prioritized and implemented .Drive operational excellence through automation of repetitive tasks operational workflows deployments recovery procedures and reliability controls .Partner with development teams to build reliable and observable services throughout the Software Development Lifecycle (SDLC) .Design develop and execute automated regression testing strategies to validate application stability reliability and performance following deployments and infrastructure changes .Review test coverage and reliability validation approaches to ensure comprehensive testing and risk mitigation .Create maintain and improve Infrastructure as Code (IaC) solutions using Terraform for cloud infrastructure provisioning configuration management and environment standardization .Support and optimize Microsoft Azure environments including Azure App Services resource management scaling strategies deployment automation and application lifecycle management .Utilize Azure-native tools such as Azure Monitor Application Insights Log Analytics and related services to improve platform visibility and reliability .Drive implementation of performance testing resiliency testing fault tolerance validation and disaster recovery preparedness within assigned domains .Establish operational readiness standards and ensure applications meet reliability scalability observability and supportability requirements before production deployment .Review architectural designs and provide recommendations to improve platform resiliency operational efficiency and cloud optimization .Lead capacity planning performance tuning and workload optimization efforts across production environments .Develop and maintain operational runbooks incident playbooks knowledge articles and standard operating procedures .Serve as a key partner with engineering infrastructure cybersecurity architecture and support teams to identify and implement continuous process improvements spanning organizational boundaries .Communicate system health reliability trends operational risks and remediation strategies to technical and business stakeholders .Present reliability initiatives operational metrics and engineering recommendations at architecture reviews technical forums and leadership meetings .Mentor engineers on observability cloud engineering automation SRE principles and operational best practices .Understand and adhere to the Company's risk and regulatory standards policies and controls in accordance with the Company's Risk Appetite .Identify reliability operational and technology risks requiring escalation to management .Promote an environment that supports a culture of belonging and reflects the Client brand .Maintain Client internal control standards including timely implementation of internal and external audit findings and regulatory requirements as applicable .Complete other related duties as assigned . Mandatory Skills Descriptio nFirst 2 weeks onsite in Buffalo or Wilmington after that remote with occasional travelin g Strong experience in observability and monitoring including hands-on expertise wit hDynatra ceOpenTelemetry (OTe l)Distributed traci ngMetrics collection and analys isCentralized logging and log aggregati onAlerting and dashboard developme ntProven experience designing and executing automated regression testing frameworks and test suites to ensure application and platform stability following deployment s.Strong proficiency in Infrastructure as Code (IaC) using Terrafor m.Experience with CI/CD pipelines deployment automation and operational toolin g.Expert knowledge of production systems monitoring incident management and operational troubleshootin g.Strong understanding of application performance management distributed systems and modern cloud-native architecture s.Cloud & Platform Experti seStrong experience with Microsoft Azure includin gAzure App Servic esResource Grou psAzure networking concep tsScaling and performance optimizati onDeployment and release manageme ntApplication lifecycle manageme ntExperience leveraging Azure-native operational tooling such a sAzure Monit orApplication Insigh tsLog Analyti csAzure dashboards and alerti ngExperience supporting cloud-native and hybrid infrastructure environment s.Reliability & Engineering Practic esDemonstrated experience implementing and operating SRE practices includin gService Level Objectives (SLO s)Service Level Indicators (SLI s)Error budge tsIncident manageme ntProblem manageme ntRoot Cause Analysis (RC A)Reliability automati onAbility to improve system reliability throug hPerformance tuni ngCapacity planni ngObservability-driven insigh tsProactive issue detecti onReliability engineering initiativ esExperience developing automated recovery mechanisms and self-healing solution s.Knowledge of resiliency engineering patterns disaster recovery planning and high-availability architecture s. Nice-to-Have Skills Descripti onExperience supporting large-scale enterprise applications in regulated environmen ts.Strong analytical and troubleshooting skills related to production systems and distributed architectur es.Experience working in Agile and DevOps operating mode ls.Ability to work autonomously and lead complex reliability initiativ es.Strong organizational and time management skil ls.Advanced verbal and written communication skil ls.Experience driving project milestones and delivery commitmen ts.Proven experience leading major incident response and post-incident improvement effor ts.Experience partnering with architecture infrastructure cybersecurity and application development tea ms.Experience with scripting and automation using PowerShell Python Bash or similar technologi es.Industry certifications in Azure Terraform Cloud Engineering or Site Reliability Engineering preferr ed.