Senior/Staff Software Engineer, Infrastructure (ML)
Nimble Robotics
$210,000 - $300,000 / year
Need a reasonable accommodation to apply or interview? Contact us.
JobMinglr uses automated technology to recommend jobs based on profile information and job preferences. Match Score does not determine eligibility for a position, prevent a user from viewing or applying to a job, or make hiring decisions on behalf of an employer.
Description
Nimble is an AI robotics company building the autonomous supply chain to power fast, efficient and economical commerce. We’re training robot AGI to power a proprietary generalist supply chain superhumanoid, the first robot in the world capable of performing thousands of tasks across the supply chain. We’ve raised over $220M at over $1B valuation and formed a strategic alliance with FedEx to build a national network of autonomous warehouses capable of generating many billions in annual revenue. We are a hardcore and obsessed team of the world’s best engineers and operators. If you are obsessed with your craft, enjoy a high-intensity and fast moving high impact environment, are super high agency in getting hard things done and want to be part of building the world’s most legendary robotics company at the most pivotal moment in history, we want to work with you.
We are on a mission to empower and inspire mankind to accomplish legendary feats by inventing robots that liberate us from the menial. We will accomplish this by training robot AGI to invent and build the Autonomous Supply Chain – everything from the inside of factories and warehouses to your front door – powered by generalist superhumanoids.
Our founding team comes from the AI labs at Stanford and Carnegie Mellon and our board of directors include famed robotics and AI legends including Fei-Fei Li (Chief Scientist of AI at Google and Director of Stanford’s AI Lab), Marc Raibert (founder of Boston Dynamics), and Sebastian Thrun (founder of GoogleX, Waymo; Stanford Professor and considered the father of autonomous vehicles).
Let’s be legendary.
About the Role
We’re looking for a Software Engineer to join our ML Infrastructure team. In this role, you’ll help build the training and inference systems that power our general-purpose warehouse robots.
You’ll own training infrastructure end to end: keeping GPUs highly utilized, making runs reproducible, and ensuring every researcher can launch the next experiment with a single command. You’ll work closely with ML and Robotics teams to design, build, and scale the systems that turn our GPU clusters into a reliable, high-throughput platform for model development.
Responsibilities
- Design, develop, and maintain ML training infrastructure that enables the AI team to run training jobs efficiently, manage and iterate experiments quickly.
- Build low-latency inference pipelines for production robotics workloads.
- Develop, tune, and optimize low-level CUDA kernels.
- Design training-platform systems for scalable model training, including high-throughput data ingestion, dataset sharding and sampling for distributed training.
- Participate in and lead design reviews with peers and stakeholders to evaluate technical tradeoffs and select appropriate technologies.
- Review code and provide feedback to uphold best practices around style, correctness, testability, performance, and maintainability.
- Contribute to documentation and educational materials, adapting content as systems and workflows evolve.
- Mentor junior engineers and help raise the technical bar across the team.
Qualifications
- Bachelor’s, Master’s, or PhD in Computer Science or a related field, or equivalent practical experience.
- 4+ years of industry experience in infrastructure, distributed systems, ML systems, robotics, or a related area.
- Experience with programming languages such as Rust, Go, Python, or C++.
- Experience with ML frameworks such as PyTorch or JAX.
- Strong understanding of distributed systems, systems programming fundamentals, memory management, and performance optimization.
- Experience with Kubernetes orchestration, resource scheduling for large distributed jobs, and containerized deployment pipelines.
- Ability to debug and optimize bottlenecks across GPU memory hierarchy, networking fabric, filesystems, and multi-GPU operations.
- Ability to reason from first principles and optimize systems for both memory-bound and compute-bound workloads.
- Strong cross-functional communication skills, ownership, and a growth mindset.
Nice to Have
- Hands-on experience with distributed training frameworks and techniques such as PyTorch DDP/FSDP, DeepSpeed, Megatron, or NCCL.
- Hands-on experience with GPU kernel development.
- Experience with data engineering technologies such as Parquet, Arrow, or similar systems.
Compensation
About this role
Nimble Robotics is building autonomous warehouse robots powered by generalist AI, and this role sits at the core of that effort—you'd own the ML training and inference infrastructure that enables the company's AI teams to develop and deploy models at scale. The work spans the full stack: designing systems to keep GPU clusters highly utilized, building reproducible training pipelines, optimizing low-latency inference for production robots, and developing CUDA kernels to squeeze performance out of hardware. You'd work directly with ML and robotics teams to turn infrastructure challenges into solved problems, then mentor others to raise the bar across the team.
This is a senior-level position that requires deep systems thinking. You'll need 4+ years in infrastructure, distributed systems, or ML systems work, strong fundamentals in systems programming and performance optimization, and hands-on experience with Kubernetes, PyTorch or JAX, and languages like Rust, Go, Python, or C++. The role particularly values people who can reason from first principles about memory hierarchies, networking, and multi-GPU operations—and who can debug bottlenecks across the entire stack. Experience with distributed training frameworks like PyTorch DDP or DeepSpeed, GPU kernel development, or data engineering tools is a plus but not required.
The position is based in San Francisco and requires in-office work. Compensation ranges from $210,000 to $300,000 annually, plus equity. This suits someone who thrives in high-intensity environments, takes ownership of hard problems, and wants to work on infrastructure that directly powers a company backed by AI and robotics leaders like Fei-Fei Li and Marc Raibert.
How this employer is doing
solid
- H1B: Sponsors Visas
- Recently Raised Funding
- Sponsors Visas
In the news
Nimble Robotics raises $65M in pursuit of fully autonomous fulfillment network, FreightWaves
This role's local market on JobMinglr
Pay for this role
$210,000 to $300,000 per year
That is the employer's posted range, and where you land in it is usually decided in one conversation. Before that conversation, run the numbers: analyze an offer for this role.
Before you apply, worth reading
How JobMinglr reads this job
Every listing here is scored against your profile before you apply: skills overlap, experience level, location and work arrangement, each weighted and explained. You see the score and the reasons, not just a list. How the matching works.