Loading jobs…
Loading jobs…
Valtech
Why Valtech? We’re the experience innovation company - a trusted partner to the world’s most recognized brands. To our people we offer growth opportunities, a values -driven culture, international careers and the chance to shape the future of experience.
The opportunity At Valtech, you’ll find an environment designed for continuous learning, meaningful impact, and professional growth. Whether you're pioneering new digital solutions, challenging conventional thinking or building the next generation of customer experiences, your work will help transform industries. We are proud of: The work we do and the innovation we drive Our values of share, care a nd dare A workplace culture that fosters creativity, diversity and autonomy Our borderless, global framework, which enables seamless collaboration The role As a Site Reliability Engineer (SRE), you are the bridge between software development and operations.
Benefits
of continuous deployment without losing grip on customer experience. You will work with our multidisciplinary teams in an essential DevOps way of working, where your main responsibility is to keep everyone focused on production while creating the infrastructure to do so.
Responsibilities
Work with teams to define SLIs and SLOs. Create systems for observability. Work with teams to analyze failure scenarios and possible mitigations.
(Assisting to) create runbooks to remediate or prevent failure scenarios. Reduce work that does not add value. Participate and facilitate incident management, including on-call duty.
Qualifications
You are someone with 5 years of experience in the field of software engineering, DevOps engineering, QA engineering and/or cloud engineering, of which at least the last 2 years as a dedicated Site Reliability Engineer. You feel comfortable taking the lead, making decisions, and know how to mobilize and motivate people to set things in motion. In your current role, people come to you for advice on what to look for to determine the robustness of their production environments, advice for reliable deployment procedures, assistance in analysis of failure scenarios, and ideas on how to mitigate or remediate those.
: You are assertive with good communicative skills, capable of taking the lead and coaching a development team to make the right choices. You have experience with incident management in a production environment of a public-facing online service with high business value and preferably high traffic in a 24x7 fashion. You have experience in working in corporate environments.
You have experience programming and scripting. You have at least basic knowledge of serverless services in one or more public cloud providers (AWS, Azure, GCP). You have extensive knowledge of and experience with various monitoring systems, amongst which APM systems such as Datadog, New Relic, Dynatrace, Prometheus, and Grafana.
You have knowledge of and experience with various pipelining tools, such as GitHub, Azure DevOps, GitLab, Jenkins. You have knowledge of and experience with microservices-related technology: Docker, Kubernetes. You have a good conceptual understanding of software architecture and system thinking.
You have worked as an engineer in a DevOps context. You have an excellent command of English (C1 or above). Are familiar with the following technologies: Datadog (or APM equivalent).
Argo CI/CD. Java / Springboot. Kafka.
Kubernetes / EKS. AWS. Have worked within the context of publicly accessible, highly available eCommerce platforms.
Have experience working in an international context with on- and off-shore teams. Commitment to reaching all kinds of people We design experiences that work for all kinds of people - and that starts with our own teams.