Loading jobs…
Loading jobs…
Veeam Software — San Jose
Veeam is the Data and AI Trust Company, specializing in helping organizations ensure their data and AI are fully understood, secured, and resilient to enable the acceleration of safe AI at scale. As the market leader in both data resilience and data security posture management, Veeam is built for the convergence of identity, data, security, and AI risk. Headquartered in Seattle with offices in more than 30 countries, Veeam protects over 550,000 customers worldwide, who trust Veeam to keep their businesses running.
Join us as we go fearlessly forward together, growing, learning, and making a real impact for some of the world’s biggest brands.
About The Role
We are looking for a Site Reliability Engineer (SRE) to join our team and ensure the continuous, reliable operation of company services. This role involves proactive monitoring, incident response, and building resilient observability and escalation practices across our infrastructure.
What You'Ll Do
Ensure monitoring and uninterrupted operation of company services Write and maintain alerting rules and runbooks Perform triage of incoming incidents and initial diagnosis of issues Build and maintain escalation chains for incident response Perform technical incident resolution activities according to runbooks Participate in on-call rotations and post-incident reviews (RCA/postmortems) Continuously improve observability coverage and reduce alert noise/false positives Collaborate with development and infrastructure teams to identify reliability risks and implement preventive measures What You'll Bring Experience with observability tools (Grafana, ELK, VictoriaMetrics) Experience working with Linux Experience working with Kubernetes (k8s) Experience with AWS and Azure cloud platforms Ability to analyze incidents, identify root causes, and propose remediation steps Bonus Skills Experience with Infrastructure as Code (Terraform, Ansible, or similar) Scripting skills (Python, Bash) for automation of operational tasks Understanding of DevOps and CI/CD principles Experience with incident management tools (PagerDuty, Opsgenie, etc.) Effective communication skills and a collaborative approach to teamwork What You'll Get Comprehensive Health Coverage – Fully employer-paid medical, dental, and vision insurance for employees and eligible dependents, including virtual care services Wellbeing Mental Health Support – Access to an Employee Assistance Program (EAP) that includes confidential therapy sessions, plus legal and financial counseling services Financial Protection
Benefits
– Company-provided life and disability insurance to help support employees and their families Paid Time Off Global Recharge Days – Vacation time, statutory holidays, and additional company-wide VeeaMe Days dedicated to rest, wellbeing, and self-care Family-Friendly Leave Programs – Competitive maternity, paternity, adoption, and other leave
that support employees through important life moments Give Back to Your Community – Employees receive paid volunteer time each year through the Veeam Cares program Please note: The position is based in San Jose, Costa Rica. If the applicant is permanently located outside of Costa Rica, Veeam reserves the right to decline the application. All applications must be submitted in English.
#LI-FT #LI-REMOTE Veeam Software is an equal opportunity employer and does not tolerate discrimination in any form on the basis of race, color, religion, gender, age, national origin, citizenship, disability, veteran status or any other classification protected by federal, state or local law. All your information will be kept confidential. Personal data collected during the recruitment process will be processed in accordance with our Recruiting Privacy Notice , which explains how your information is collected, used, and handled in connection with hiring activities.
By applying for this position, you consent to this processing.