Loading jobs…
Loading jobs…
The Scion Group — London England
About CLS: CLS is the trusted party at the centre of the global FX ecosystem. Utilized by thousands of counterparties, CLS makes FX safer, smoother and more cost effective. Trillions of dollars’ worth of currency flows through our systems each day.
Created by the market for the market, our unrivalled global settlement infrastructure reduces systemic risk and provides standardization for participants in many of the world’s most actively traded currencies.
Requirements
by over 96% on average, so clients can put their capital and resources to better use. CLS products are designed to enable clients to manage risk most effectively across the full FX lifecycle – whether through more efficient processing tools or market intelligence derived from the largest single source of FX executed data available to the market. Our ambition to make a positive difference starts with our people.
Our values underpin everything that we do at CLS and define our working environment: Pivotal purpose Trusted guardian Targeted innovation Facilitate connections Delivering excellence Inclusive culture Job information Functional title – Site Reliability Engineer Department – Technology Corporate level – Assistant Vice President Report to – Vice President Location - London, onsite 2 days per week Job Purpose The role is primarily responsible for developing SRE methodologies and ensuring they are applied to the Cloud hosted environment. In addition, the role will act as a central point of expertise for SRE automation across the Platform Operations team. Essential Job Functions Responsible for driving the implementation of SRE methodologies within the CLS environment, collaborating closely with other infrastructure teams to optimize infrastructure and deployment processes, focusing on automation and operational excellence.
Drives continuous improvement in system observability, alerting, and capacity planning through the definition and implementation of SLA, SLOs SLIs Define and enhance frameworks for Toil identification, analysis remediation to identify opportunities to eliminate or automate remediation of recurring tasks and issues Develops secure high-quality production code, and reviews and debugs code written by others. Build out and enhance GitOps capabilities for use in the Cloud hosted environments using tools such as Terraform and Ansible Automation Platform Provide on-call support and escalation for Cloud Automation related issues ensuring that Production stability is the primary requirement. Ensure risks and stability issues in the cloud hosted environment are understood and addressed where possible through SRE best practices as part of any incident postmortems.
g. AWS / Terraform Minimum Job-Related Experience Required Must have strong technical operational support experience within an infrastructure services team performing on-call duties such as handling tickets, owning incidents investigating their root cause Minimum of 2 years experience applying SRE methodologies within a support team and an understanding of Service Level metrics associated with this. Strong knowledge of at least 1 scripting language, preferably either Python or Ansible.
g. Grafana / Datadog / Dynatrace). Experience of working in a regulated financial services / banking organization.
Excellent troubleshooting, analytical, and communication skills with both business and technical staff. Special Skills/Knowledge Software Development background. Familiar with the ITIL framework.