Skip to main content
← Back to search
R

Member of Technical Staff — Training

RadixArk

$10,000 - $100,000 / year

Palo Alto, CAOn-siteExpert

Need a reasonable accommodation to apply or interview? Contact us.

JobMinglr uses automated technology to recommend jobs based on profile information and job preferences. Match Score does not determine eligibility for a position, prevent a user from viewing or applying to a job, or make hiring decisions on behalf of an employer.

Description

 

About the Role

As a Member of Technical Staff, Training, you will design, build, and operate the distributed systems behind large-scale model post-training — spanning training, inference, and orchestration, with a focus on the performance, correctness, scalability, and reliability of workloads running across large GPU clusters.
This role suits engineers who move fluidly across modeling recipes, complex infrastructure, and low-level systems, identify bottlenecks in distributed workloads, and translate experimental requirements into robust software.

In This Role, You Will

  • Design, build, and operate distributed training, rollout, and orchestration systems for large-scale LLM and multimodal post-training across multi-GPU, multi-node environments.
  • Profile and optimize performance across the full-stack — model implementation, parallelism strategies, communication libraries, and GPU kernels — to improve throughput, latency, memory efficiency, hardware utilization, and cost.
  • Investigate numerical correctness and low-precision issues in distributed training and inference, including train–inference consistency for reinforcement learning.
  • Improve the reliability of long-running workloads through checkpointing, fault recovery, observability, and operational tooling.
  • Build supporting infrastructure for reinforcement learning and agentic post-training, including asynchronous rollout, trajectory collection, sandboxed execution, evaluation harnesses, and data pipelines.
  • Contribute to open-source training and inference systems, including Miles and SGLang, and partner with researchers to turn experimental requirements into production systems.

Minimum Qualifications

  • 3+ years of experience building or operating distributed machine learning systems, large-scale training infrastructure, or high-performance inference systems.
  • Hands-on experience with post-training systems, training backends, or inference systems for large language models (e.g., Megatron-LM, FSDP, SGLang, TensorRT-LLM, vLLM).
  • Experience in at least two of the following areas:
    • Performance, efficiency, and scalability of multi-GPU, multi-node workloads
    • Numerical correctness or low precision
    • Stability, reliability, or fault tolerance
    • Post-training algorithm recipes and orchestration infrastructure for large training runs
    • Multimodal training or inference, including vision-language models and multimodal generation
    • Agent infrastructure, including sandboxes, harnesses, and eval systems
    • Building and maintaining open-source projects widely adopted in industry and academia

Preferred Qualifications

  • Familiarity with RL algorithms such as PPO, GRPO, and their variants, and experience applying them in large-scale post-training.
  • Experience with modern post-training frameworks (e.g., Miles, slime, AReaL, verl, Prime-RL).
  • Key open-source contributions to training or inference frameworks (e.g., SGLang, vLLM, Megatron-LM).
  • GPU kernel development (e.g., CUDA, Triton, CUTLASS) or communication-layer optimization (e.g., NCCL, RDMA, NVLink/NVSwitch).
  • Experience training or serving models at very large scale (e.g., Mixture-of-Experts models on clusters of thousands of GPUs).
  • Top-tier publications in ML systems or other systems fields.

Even if you don't meet every qualification above, we encourage you to apply — we care most about demonstrated ability to build and reason about large-scale systems.

About RadixArk

RadixArk builds open-source and production infrastructure for large language models and multimodal post-training. Our systems — including Miles, an enterprise-grade reinforcement learning training framework, and SGLang, a widely deployed high-performance LLM inference engine — power distributed post-training across clusters of 10k–100k+ GPUs.

Compensation

Depending on background, skills, and experience, the expected annual salary range for this position is $200,000 to $400,000, plus equity.

Benefits include a 401(k) plan and unlimited PTO.

RadixArk sponsors employment visas (e.g., H-1B, O-1) for eligible candidates.

Equal Opportunity

RadixArk is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more.

About this role

RadixArk is hiring a Member of Technical Staff to design and operate the distributed systems that power large-scale LLM post-training across massive GPU clusters. You'd work across the full stack—from model implementation and parallelism strategies through communication libraries and GPU kernels—profiling and optimizing for throughput, latency, memory efficiency, and cost. The role also involves ensuring numerical correctness in low-precision training, building reliability into long-running workloads through checkpointing and fault recovery, and developing infrastructure for reinforcement learning and agentic post-training, including rollout systems, trajectory collection, and evaluation harnesses.

This position suits engineers with at least three years building or operating distributed ML systems and hands-on experience with post-training backends like Megatron-LM, FSDP, or SGLang. You should be comfortable moving between modeling recipes, complex infrastructure, and low-level systems work, with demonstrated expertise in at least two areas such as multi-GPU performance optimization, numerical correctness, fault tolerance, or post-training orchestration. Experience with RL algorithms, GPU kernel development, or open-source contributions to training frameworks is valued but not required; RadixArk emphasizes demonstrated ability to build and reason about large-scale systems over checking every box.

The role is based in Palo Alto and requires in-office work. Compensation ranges from $200,000 to $400,000 annually plus equity, with 401(k) and unlimited PTO. RadixArk sponsors employment visas for eligible candidates.

How this employer is doing

average

  • Sponsors Visas

This role's local market on JobMinglr

Pay for this role

$10,000 to $100,000 per year

That is the employer's posted range, and where you land in it is usually decided in one conversation. Before that conversation, run the numbers: analyze an offer for this role.

How JobMinglr reads this job

Every listing here is scored against your profile before you apply: skills overlap, experience level, location and work arrangement, each weighted and explained. You see the score and the reasons, not just a list. How the matching works.