Loading jobs…
Loading jobs…
Mthree — Charlotte North Carolina
**Looking for local candidates** Want to work in technology in the financial industry? Our client is seeking a highly motivated Site Reliability Engineer to join a dynamic team, Global Banking Technology is building and scaling Site Reliability Engineering (SRE) across a large, highly regulated banking environment. We are seeking a senior SRE practitioner to lead and accelerate transformation from traditional L2 production support toward an SRE operating model.
This role will help define, implement, and embed SRE practices across infrastructure and banking services, enabling measurable reliability outcomes, reduced manual toil, stronger automation, and improved service visibility. The successful candidate will bring proven, hands-on experience implementing SRE in a large corporate bank and will be able to influence operations, engineering, and product partners to institutionalize SRE practices on a scale. About mthree: Since 2010, mthree has been helping clients solve their business and technological challenges.
We are a technology and business consultancy with a global workforce delivering significant business and IT projects in some of the largest financial services organizations worldwide. Core Services: Consulting and Advisory Managed Services Alumni Graduate Program Alumni Pro Program We have a global presence and are experts in delivering exceptional quality to our client base, providing consulting services across Risk, Regulation & Compliance; Vendor Products; Application Support; Application Development; Cyber & Information Security; Data Science and DevOps areas. Our Expert program offers experienced professionals access to top roles in tech, finance, aviation and insurance.
Join us to work on groundbreaking technology projects, from international trading platforms to critical applications for leading airlines. We recruit professionals who are eager to fast-track their careers in technology or operations within prestigious global organizations.
Key Responsibilities
SRE Operating Model and Transformation Lead the design and execution of the SRE adoption approach across Global Banking, including the transition path from traditional L2 support to reliability engineering. Establish practical engagement patterns between SRE, application teams, and platform teams and help teams adopt a consistent way of working. Reliability Measurement and Decisioning Drive adoption of Critical User Journeys, Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets for priority services, ensuring metrics reflect user experience and business outcomes Help teams implement error budget based decisioning that balances reliability, delivery velocity, and operational risk Toil Reduction, Automation, and Engineering Excellence Identify operational toil and lead initiatives to eliminate it through automation, self-healing patterns, runbook automation, and operational tooling improvements Establish and implement a model to partner with engineering teams to build reliability into services through design improvements, improved instrumentation, and resilience patterns Incident and Problem Management Excellence Improve production outcomes through strong incident response practices, including major incident triage support, root cause analysis, post incident reviews, and preventive engineering actions.
Strengthen problem management with a focus on reducing repeat incidents, technical debt risk, and manual intervention. Observability and Tooling Enablement Establish practical observability standards across logs, metrics, traces, dashboards, and alerting to reduce noise, improve signal quality, and shorten time to detect and restore service.