Loading jobs…
Loading jobs…
Striveworks — Fort Carson Colorado
“In 36 months, agentic AI systems will be an operating reality across major institutions. ” — Dr. Jim Rebesco, Cofounder and CEO, Striveworks The government’s demand for AI is growing far faster than the systems required to support it.
Fewer than 15% of federal AI programs have reached sustained production, despite billions of dollars invested. The models perform in testing, but they degrade in the real world. And when performance drops, trust goes with it.
Striveworks was built to solve that problem. What you’ll build Since 2018, we have delivered the most trusted AI systems operating in real-world use cases—providing a layer of assurance underneath hundreds of deployed models that monitors performance, manages drift, and sustains systems long after they leave the lab. As a DevOps Engineer supporting operations on site at Fort Carson, CO, you are the tactical edge of our engineering team.
You aren’t just maintaining a platform; you play a key role in our technical success.
Requirements
—and tailor our automation and deployment strategies to move the needle for them. You will be responsible for maintaining the end-to-end life cycle of our AI platform across a diverse architectural landscape—including remotely accessible cloud environments and on-premises hardware. You’ll thrive in this role if you enjoy the challenge of debugging complex Kubernetes clusters where connectivity is a luxury, not a given.
You are someone who can navigate the unique constraints of local hardware today and write the automation that ensures seamless, reliable performance across the customer’s hybrid stack tomorrow. What it’s like here We lead with trust, treat each other with respect, and use candor consistently, kindly, and constructively. We care deeply about our work, and we find genuine satisfaction in doing it well.
Above all, we take ownership—because we feel the weight of collective results personally. We are looking for people who share these values and are eager to put them into practice.
, and policies Proficiency with US federal information system security policies, including STIGs, NIST SP 800-171, NIST SP 800-53, CMMC, and ICD 503 Experience with software deployments to on-premises and cloud-based unclassified, CUI, and classified networks within the DOD Experience with DevSecOps/DevOps and CI/CD for the administration and deployment of GPU-enabled servers Experience deploying or maintaining CNCF projects Experience with NAS and SAN technologies Experience with Kubernetes and cloud-native applications and services in DDIL impact environments This position is hybrid/on site at our
What We’Re Looking For
3–5+ years of hands-on experience in software, DevOps, site reliability, or systems engineering Proven technical leadership experience, with the communication skills and professional presence required to manage customer relationships and lead cross-functional incident responses Expertise in deploying and diagnosing microservices within K8s Experience with comprehensive observability solutions using tools such as Prometheus, Grafana, and OpenTelemetry to ensure system reliability and performance visibility Ability to design and execute testing strategies to validate application functionality, diagnose issues, and perform root cause analysis, implementing effective long-term fixes to improve stability and performance Deep proficiency in Terraform, Ansible, or similar tools to manage virtual machines and containerized services Strong scripting and coding skills in Bash and/or Python for building custom automation and tooling Ability to manage and rapidly troubleshoot Linux systems Active Secret (or above) US security clearance and US citizenship The following isn’t required, but we’d love to see it: Experience with DOD networking, tools, infrastructure, security