Loading jobs…
Loading jobs…
Reward Gateway — London, London, City of
Reward Gateway, part of Edenred, is a global leader in
Benefits
and employee engagement. We help businesses attract, engage, and retain top talent through strategic reward, recognition, and well-being solutions. Guided by our shared missions - ‘Making the World a Better Place to Work’ and ‘Enriching Connections, For Good’ - we’re committed to transforming workplaces and improving people’s daily lives.
Our team embodies entrepreneurial spirit, innovation, and respect. We push boundaries, speak up, and stay human, fostering a culture where imagination thrives. This role offers a hybrid work model to be present in our London office twice a week.
Your Role in our Mission: This hands-on role sits at the intersection of operational excellence and engineering craft. You’ll bridge the gap between traditional application support and software engineering by executing scripted remediation, configuration management, feature flag operations, safe, bounded code-level fixes, and runbook automation — all under clearly defined guardrails. The goal is to reduce unnecessary L3 escalations while increasing autonomy, quality, and impact for our Application Operations function.
You’ll apply these practices across our AWS environment ( EKS ), PHP services, and MySQL databases, using Datadog as our observability platform, Kibana for log exploration, and Heap to help quantify and understand customer impact. 5 support for PHP applications running on EKS with MySQL backends, operating within clear guardrails that include configuration changes, feature flag operations, scripted runbooks, and safe, bounded code-level fixes. 5, automate more, and escalate less, increasing the percentage of incidents resolved without L3 involvement and improving MTTR.
Participate in a healthy, sustainable on-call rotation with fair schedules, clear escalation paths, and strong post-incident learning practices. Engineering Practices Within Operations Apply engineering discipline to operational work: use version control, code review, and testing standards for scripts, runbooks, and automation tooling you produce. Develop and maintain automation scripts, runbooks, and playbooks for known issue patterns across workloads, services, and operational scenarios.
Identify and automate repetitive remediation tasks to reduce manual toil and improve MTTR. Observability and Service Readiness Collaborate with peers to ensure the right monitoring signals, dashboards, and alerts exist in Datadog. Tune app-level alerts and dashboards to minimize noise and surface actionable signals.
Use Kibana to interrogate logs and correlate events with Datadog signals during investigations; improve log usefulness by feeding back patterns for better parsing and context. Use Heap to triangulate and quantify customer impact (affected flows, cohorts, and volumes) during incidents and problem investigations; incorporate findings into incident timelines and post-incident reviews. Participate in service onboarding and operability reviews to ensure new and changed services meet defined supportability standards before production.
Contribute to the Service Catalogue with accurate ownership, SLAs/SLOs, runbooks, and escalation paths for supported services. , safe config changes, feature flag toggles, rolling restarts, cache purges, scripted data fixes). Support major incidents by providing technical context, structured diagnostics, Datadog/Kibana evidence, Heap impact analysis, and coordinated remediation alongside the incident commander.
Use structured diagnostics before escalating — attach clear evidence, reproducibility steps, and impact assessments to every L3/SRE handoff.