Position Announcement At Utah Valley University you'll have the opportunity to build and support the technology solutions that help students faculty and staff succeed every day. In this role you will design implement and maintain reliable scalable systems and automation that improve service delivery strengthen system performance and support the university's ongoing digital transformation efforts. Working with modern cloud infrastructure and site reliability practices you will play a key role in ensuring critical services remain secure available and responsive for the campus community. This position offers a dynamic mix of engineering operations and innovation. You will collaborate with technical teams and university leaders to enhance system reliability develop monitoring and automation solutions support technology upgrades and resolve complex challenges. Ideal for a technology professional who enjoys continuous learning and problem-solving this role provides the chance to make a meaningful impact while working in a collaborative environment that values expertise initiative and service excellence.
Of
Performs day-to-day administration maintenance upgrades and operation of existing and recently developed systems including Virtualization infrastructure on-premises and in the cloud Microsoft Windows System administration and Linux operating systems as well as application systems and technologies. Ensures standard operating procedures runbooks disaster recovery and service catalog definitions are mature and ready for production. Responsible for lifecycle of product set maintenance availability reliability and performance reporting to decision makers and developers. Perform tasks as needed to augment work needed on systems by their engineers to ensure timely achievement of project plans and goals. Advocate contribute recommend and facilitate these ever-improving standards and best practices through successful adoption of change within UVU’s Digital Transformation department. Ensure that the underlying infrastructure is running smoothly and that systems and tools are working as expected. Analyze day-to-day functions and the processes of systems and network management software to ensure they are performing within predetermined specifications. In support of core systems availability and reliability integrate diverse monitoring solutions for emerging and existing IT infrastructure using automation and API tools in on-premises and cloud architectures. Engineer centralized enterprise-wide alerting and key performance indicators that give timely actionable information to subject matter experts stakeholders and leadership. SRE teams conduct post-incident reviews documenting findings and acting on lessons learned. Following the incident resolution the engineer will revisit the issue and determine the cause. Build or optimize the incident lifecycle to bolster the reliability of services. Maintain documentation and runbooks to ensure that teams get information when they need it. Develops operational tools and processes builds reliable systems ensures compliance with operational standards and provides support to operational staff. This can be anything from adjustments to monitoring and alerting to code changes in production. An SRE can be tasked with building a homegrown tool from scratch to help with weaknesses in software delivery or incident response and management. SREs' responsibilities include writing and developing code to automate processes such as analyzing logs testing production environments and responding to any issues. Such automation allows developers and engineers to focus their attention on bug fixes and building new features rather than being burdened by the day-to-day operational requirements needed in their projects. Provides leadership communications development engineering automation and feedback necessary for enterprise planning and architecture. Timely and responsive work is key for providing what went well or what went badly during a change/incident/problem cycle. Participate in after-hours and weekend on-call rotation and provide training to other on-call staff. Provide remote hands for systems and application administrators that need physical and virtual support within on-premises and cloud facilities. Perform other job-related duties as assigned.
/ Licenses / Certifications Graduation from an accredited institution with a bachelor's degree in Information Technology or a related field plus three years of work experience in IT OR a combination of education and experience in a related field totaling seven years. For Example Bachelor's degree in Information Technology or a related field + 3 years of related work experience. OR two years of completed college coursework in Information Technology or a related field + 5 years of related work experience 7 years. OR technical certification program or college coursework equivalent to 1 year + 6 years of related work experience 7 years. OR 7 years of progressively responsible related work experience. Licenses/Certifications Site Reliability Engineering (SRE) Professional Certificate Microsoft Certified Azure Fundamentals Administrator Developer or Associate-level certifications Amazon Web Services (AWS) Cloud Practitioner or Associate-level certifications Information Technology Infrastructure Library (ITIL) certification The Open Group Architecture Framework (TOGAF) certification Docker Certified Associate (DCA) Certified Kubernetes Administrator (CKA) Knowledge /
/ Abilities Knowledge Knowledge of ITIL Change Incident and Problem Management. Knowledge of TCP/IP firewall management and operating system configuration. Proficient and current knowledge of industry trends tools and processes. Knowledge of Agile and iterative development processes (e.g. Scrum and Kanban). Knowledge of automation and containerization technologies such as Docker Kubernetes Ansible Terraform and SaltStack. Knowledge of ITSM platforms such as Jira Service Management ServiceNow or others. Knowledge of Engineering practices availability reliability and scalability as well as disaster recovery Knowledge of various automation tools as they are usually responsible for building and integrating software tools to enhance an organizational system’s reliability and scalability.
Recognize key design implementation and process issues and proactively craft and automate solutions. Skill with system engineering and design for NOC/SOC purposes. Skill with scripting languages such as Perl PowerShell Bash and Python.
with most of the common programming languages including JavaScript HTML5 CSS JQuery JSON and PHP.
with the design implementation and maintenance of Active Directory and/or LDAP directories.
with TCP/IP application network protocols firewall management operating system configuration anti-virus software and relational databases. Practical Experience with various Monitoring solutions such as Prometheus PRTG Site24x7 TestCafe Selenium Splunk New Relic Azure Monitor and AWS CloudWatch. Expertise in the major cloud providers such as Azure AWS and Google Cloud. Experience with alert management/on-call tools such as PagerDuty VictorOps and Opsgenie. Experience with instant communication and team collaboration platforms like MS Teams Slack or Jitsi Proven IT project planning and development skills Abilities Expert ability to read write and interpret technical documentation runbooks procedure manuals and knowledge-base articles pertaining to network systems and application management. Ability to complete Root Cause Analysis (RCA) investigations and write post-incident reports. Ability to improve team practices through code reviews handoffs of work and incidents. Be on an on-call (PagerDuty) rotation to respond to incidents that impact availability and provide support for service engineers with customer incidents. Ability to debug production issues and build monitoring that alerts on symptoms rather than on outages. Ability to turn into repeatable actions and into automation. Ability to conduct and direct research into IT issues and products as required Ability to communicate technical ideas and concepts to a non-technical audience.
Unlock: Sign Up for free / Sign In and use the searches from your home page or the links in the footer.