Loading jobs…
Loading jobs…
Dragos
At Dragos, the mission is personal. The systems we protect deliver the water you drink, power your home, and keep the hospitals your community depends on running. Those critical infrastructure systems that power our civilization around the world are under attack every day by adversaries.
When those systems fail, people are immediately at risk. We are the global leader in xOT cybersecurity, combining technology, threat intelligence, and expert services. The people here chose this work because they understand what is at stake .
Here, you will find a remote-first mission-driven team across North America, Europe, the Middle East, and APAC built on authenticity, transparency, and trust. If safeguarding the systems that protect your family, friends, and community is the kind of work that matters to you, you are in the right place.
About The Role
We are seeking an experienced Senior Cloud Operations Engineer to join our Cloud Skill Community. This role owns the operational health of the Dragos customer cloud fleet across Azure, AWS, and GCP -- both Dragos-managed and customer-managed environments. You will drive fleet reliability, deployment automation, and day-to-day operations at scale, bringing a strong infrastructure-as-code mindset and a bias toward automation and repeatability.
Responsibilities
Operate, maintain , and improve the Dragos cloud fleet across Azure, AWS, and GCP Own the full customer environment lifecycle -- onboarding, configuration, upgrades, and off boarding Build and maintain Terraform-based infrastructure-as-code for customer environment provisioning and fleet standardization Manage fleet health, drift detection, patching, and version lifecycle management at customer scale Design and enforce multi-tenant isolation patterns -- blast radius containment, RBAC at scale, and cross-account access controls Configure and maintain cloud networking components across all three providers (VPCs/ VNets , peering, transit gateways, DNS, firewalls, load balancers) Manage cloud-to-OT/on-prem connectivity for customer environments Implement and maintain IAM, secrets management, and compliance posture across cloud providers Build and maintain Datadog observability -- monitors, dashboards, log pipelines, and SLOs Participate in an on-call rotation (PagerDuty) for the Dragos cloud fleet -- triage, respond to, and resolve production incidents across customer environments Drive SRE practices: define SLOs, manage error budgets, maintain runbooks, and lead post-incident reviews Support audit and compliance activities including evidence collection for FedRAMP, SOC2, and customer-specific
-- this role participates in a PagerDuty rotation for the customer cloud fleet Experience with SRE practices: SLO definition, error budget management, incident response, and blameless post-mortems Strong scripting skills in Python and/or Bash Familiarity with compliance frameworks (FedRAMP, SOC2, NIST CSF) and audit evidence collection Strong ownership mindset and accountability for production stability Excellent communication, documentation, and collaboration skills
Requirements
Identify and eliminate toil through automation and process improvement
Qualifications
4+ years of hands-on cloud operations experience across one or more of Azure, AWS, or GCP Cybersecurity Experience Strong proficiency with Terraform for infrastructure-as-code (primary IaC tool at Dragos) Experience operating cloud environments at scale -- fleet management, patching, upgrades, drift detection Hands-on experience with multi-tenant cloud architectures and customer-facing environment management Solid knowledge of cloud networking: VPCs/ VNets , peering, transit gateways, DNS, firewalls, and hybrid connectivity Experience with IAM across AWS, Azure Entra ID, and/or GCP -- roles, policies, federation, SSO Proficiency with Datadog for monitoring, alerting, dashboards, and log pipelines Comfort with on-call
Preferred Qualifications
Experience with all three major cl