Loading jobs…
Loading jobs…
Nuro — View Royal, British Columbia
Who We Are
Nuro believes self-driving vehicles are the most immediate and profound opportunity for AI to drive positive change in the physical world. Safer streets, more time for what matters, and easier access to the world around us, that’s why we’re building a universal autonomy platform: self-driving for all roads and all rides. Founded in 2016, Nuro is a physical AI company developing Level 4 autonomous driving technology for a wide range of vehicles, use cases, and markets.
Powered by the Nuro Driver™, our universal autonomy platform enables the global mobility ecosystem to deploy autonomy at scale, from robotaxis and logistics fleets to personal vehicles. With years of real-world deployment experience and a flexible, partner-led business model, Nuro is working toward a future where millions of autonomous vehicles powered by our technology help make everyday life safer, easier, and more connected. Nuro has raised over $2B in capital from Uber, NVIDIA, Google, Softbank, Fidelity, T.
About The Role
Nuro is seeking a Software Engineer with expertise in large-scale infrastructure, workload orchestration, and data processing to join our ML Infrastructure team . In this role, you will focus on building and evolving the core platform that provides researchers and engineers with seamless access to compute and data resources. You will be responsible for executing the technical strategy for automated resource provisioning, high-performance workload scheduling, and efficient feature management to accelerate the Nuro Driver™ development lifecycle.
About the Work You will build the foundation that powers Nuro’s model development from experimentation to production.
Key Responsibilities
include: Resource Provisioning IaC: Scaling automated infrastructure-as-code (IaC) pipelines to manage thousands of GPU/CPU nodes across diverse environments. Intelligent Scheduling: Designing and optimizing workload orchestration to maximize hardware utilization, minimize job wait times, and handle massive-scale distributed training. Data ETL: Designing robust pipelines for the extraction and transformation of petabyte-scale sensor and telemetry data into ML-ready formats.
Feature Management: Implementing robust feature caching and storage solutions to reduce redundant computations and ensure low-latency access to pre-computed features. Platform Abstraction: Contributing to a unified ML platform that abstracts complex cloud infrastructure for end-users. About You Experience: 3+ years of professional experience in ML Infrastructure, Backend Platform Engineering, or Distributed Systems.
Resource Provisioning: Deep familiarity with modern Infrastructure-as-Code and provisioning tools such as Terraform, Pulumi, or Crossplane. , Kubernetes, KubeRay, Ray, Slurm, or Volcano). Distributed Data Processing: Proficiency in at least one distributed processing framework, such as Apache Spark or Apache Beam, for large-scale data extraction and transformation.
, Feast, Hopsworks, or Redis-based custom caching). Systems Design: A strong understanding of distributed systems, networking, and storage bottlenecks in the context of high-performance computing. , CNCF, Ray, or Kubeflow communities).
, Lustre, Ceph, or specialized NVMe caching) for ML data loading. Knowledge of cost-optimization strategies for large-scale GPU clusters in public clouds (AWS, GCP, or Azure).
Compensation
package.
Base Pay Range
is between $160,360 and $240,540 for the level at which this job has been