Loading jobs…
Loading jobs…
CoreWeave — Livingston, California
CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence. Trusted by leading AI labs, startups, and global enterprises, CoreWeave combines superior infrastructure performance with deep technical expertise to accelerate breakthroughs and turn compute into capability.
Founded in 2017, CoreWeave became a publicly traded company (Nasdaq: CRWV) in March 2025. com .
What You’Ll Do
: The Fleet Provisioning Automation (FPA) team is responsible for the automated provisioning and lifecycle management of CoreWeave’s rapidly growing fleet of hardware nodes and node types. The team streamlines and coordinates node bring-up, hardware RMA, data center operations, and platform services into a cohesive, high-reliability engine of fleet management. This group sits at the intersection of hardware, data center operations, and platform engineering, building the software that keeps our global fleet healthy and ready for customer workloads.
About The Role
: As a Senior Software Engineer on the Fleet Provisioning Automation team, you will design and build backend services and APIs that automate provisioning, configuration, and lifecycle operations for CoreWeave’s globally distributed bare metal fleet. You will primarily work in Go to implement gRPC APIs that integrate with Kubernetes, vendor and internal services, and data center tooling. Your work will focus on turning complex, multi-step operational procedures into simple, safe, and auditable automation for internal users operating at hyperscale.
In this role, you will: Design and implement backend services and APIs (primarily gRPC in Go) that orchestrate provisioning, configuration, and lifecycle operations across CoreWeave’s global server fleet. Develop Kubernetes custom resource definitions (CRDs) to automate provisioning and lifecycle management of CoreWeave’s entire server fleet. Model and evolve RPC schemas and data contracts used by other engineering and operations teams to integrate with fleet provisioning workflows.
Build integrations with vendor and internal APIs to make hardware and data center processes robust, transparent, and easy to operate at scale. Collaborate closely with hardware engineering, data center operations, and platform teams to design solutions to problems of scale for multi-site deployment and management of CoreWeave’s global hardware fleet. Create test plans, deployment automation, and supporting tooling that enable safe rollouts and continuous improvement of fleet provisioning services.
Participate in the Fleet Provisioning Automation on-call rotation and help drive incident response and post-incident improvements.
Who You Are
: 5+ years of experience in software or infrastructure engineering building and operating production backend or infrastructure services. Proficiency in Go for building networked services and APIs (gRPC and REST) in production environments. Experience designing and implementing distributed systems that operate reliably at scale, including concurrency, failure handling, and resiliency patterns.
Strong understanding of Linux systems and how to debug issues across processes, networking, and storage. Experience with Kubernetes or similar container orchestration platforms and their APIs (for example, interacting with custom resources, controllers, or operators). Familiarity with CI/CD tooling (such as Argo, Flux, or GitHub Actions) to ship and operate services safely and frequently.
Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent practical experience. Preferred: Experience designing and operating services that automate lifecycle management for large fleets of physical servers or other hardware. Experience with infrastructure automation and configuration management tools (for example, Ansible, Puppet, Chef, or Salt).