This position is listed on behalf of a partner company who manages all applications and next steps. Our partner is looking for a Senior DevOps Engineer based in United States. Join a globally distributed infrastructure team responsible for powering highly available production-critical financial technology. You will design automate and operate cloud infrastructure on Google Cloud while helping engineering teams ship faster and more safely. The role combines cloud architecture Infrastructure-as-Code Kubernetes CI/CD observability networking and reliability engineering. You will take ownership of platform capabilities and build self-service solutions that reduce manual work and improve developer productivity. You will also contribute to incident response security capacity planning and continuous improvements across the infrastructure environment. Working in an async-first international setting you will have significant autonomy and a direct influence on technical decisions and platform evolution. This is an opportunity to solve complex infrastructure challenges at scale while applying strong Platform-as-a-Product and SRE principles. Accountabilities Design and evolve highly available cloud architecture on Google Cloud including networking interconnects IAM and resilient infrastructure topologies using Terraform and GitOps practices. Build and maintain secure CI/CD pipelines for Infrastructure-as-Code incorporating automated planning and testing code review Policy-as-Code guardrails drift detection and controlled rollouts. Develop a Platform-as-a-Product approach by creating self-service capabilities reusable infrastructure patterns and developer-friendly golden paths that enable teams to provision resources efficiently. Operate and improve production Kubernetes/GKE environments including workload deployment with Helm scaling networking security observability and troubleshooting. Strengthen platform observability across metrics logs traces and alerting using technologies such as Prometheus Thanos Grafana Loki Tempo and Alertmanager. Operate infrastructure services and data platforms including PostgreSQL and message brokers while partnering with specialized SRE and database teams on complex operational challenges. Participate in a global Follow-The-Sun on-call rotation responding to alerts supporting incidents leading structured troubleshooting coordinating escalations and driving blameless post-mortems and follow-up actions. Apply SRE principles such as SLIs SLOs error budgets and capacity planning to improve reliability and operational maturity. Design and troubleshoot cloud networking components including VPCs routing load balancing DNS TLS firewalls and interconnects. Partner with security and engineering teams to implement secure-by-default infrastructure least-privilege access and effective infrastructure security practices. Identify opportunities to eliminate manual toil through automation and continuously improve reliability security scalability cost efficiency and developer experience. Mentor engineers and lead infrastructure initiatives that raise engineering standards and improve the overall platform.
5+ years of professional experience in DevOps SRE Platform Engineering Infrastructure Engineering or a closely related discipline with experience operating large-scale highly available production systems. Deep hands-on experience designing and operating cloud architecture on Google Cloud Platform including landing zones networking IAM and high-availability architectures. Strong Terraform and Infrastructure-as-Code expertise including structuring large codebases across multiple environments and applying GitOps and least-privilege principles. Proven experience building CI/CD pipelines for Infrastructure-as-Code including automated plan/apply workflows code review Policy-as-Code drift detection and safe deployment practices. Significant production experience with Kubernetes ideally Google Kubernetes Engine (GKE) and Helm-based workload deployment. Strong understanding of cloud and L3/L4-L7 networking fundamentals including VPCs routing load balancing DNS TLS firewalls and interconnects with the ability to troubleshoot complex connectivity issues. Practical experience with modern observability platforms covering metrics logs traces and alerting particularly Prometheus Thanos Grafana Loki Tempo and Alertmanager. Operator-level knowledge of production data stores and messaging systems such as PostgreSQL RabbitMQ or similar message brokers. Solid understanding of SRE principles including SLOs error budgets capacity planning incident management and root-cause analysis. Strong scripting or programming capabilities in Python Go Shell or equivalent technologies. Demonstrated ability to independently troubleshoot complex production problems and drive issues through resolution. Strong ownership communication documentation and problem-solving skills with the ability to work effectively in a globally distributed and async-first environment. Willingness to participate in a Follow-The-Sun on-call rotation including scheduled availability for urgent infrastructure incidents. Experience with Policy-as-Code and IaC quality tools such as OPA/Conftest Checkov tflint or Atlantis is a plus. Experience managing Terraform state module registries and versioning at scale is a plus. Experience building internal developer platforms or golden paths using tools such as Backstage or Tilt is a plus. Working knowledge of Go Linux Docker/containerd Ansible Chef or Puppet is advantageous. Experience securing containers and Kubernetes/GKE environments is beneficial. Knowledge of advanced Google Cloud security controls regulated environments SOC 2 secrets management or audit logging is advantageous. Familiarity with trading brokerage fintech or other regulated and low-latency environments is a plus. Relevant cloud certifications particularly Google Cloud Professional certifications are valued.
Competitive salary and stock options. Health benefits. One-time USD $500 home-office setup allowance for new hires. USD $150 monthly stipend provided through a company expense card. Fully remote work within the eligible location. Opportunity to work with a globally distributed team across multiple regions. High level of autonomy ownership and influence over infrastructure and platform strategy. Opportunity to work on highly available trading-critical systems and complex cloud infrastructure challenges. A strong focus on developer productivity automation reliability and Platform-as-a-Product practices. Exposure to modern cloud Kubernetes Infrastructure-as-Code observability and SRE technologies. Inclusive environment committed to building a diverse and collaborative workforce. How Jobgether Works We use an AI-powered matching process to ensure your application is reviewed quickly objectively and fairly against the role's core requirements. Our system identifies the top-fitting candidates and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews assessments) are managed by their internal team. We appreciate your interest and wish you the best! Why Apply Through Jobgether? Data
By submitting your application you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access rectification erasure objection) at any time. We may use artificial intelligence (AI) tools to support parts of the hiring process such as reviewing applications analyzing resumes or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed please contact us.
Unlock: Sign Up for free / Sign In and use the searches from your home page or the links in the footer.