Skip to main content
← Back to search
R

Member of Technical Staff — Inference-Core Engine

RadixArk

Salary undisclosed

Palo Alto, CAOn-siteExpert

Need a reasonable accommodation to apply or interview? Contact us.

JobMinglr uses automated technology to recommend jobs based on profile information and job preferences. Match Score does not determine eligibility for a position, prevent a user from viewing or applying to a job, or make hiring decisions on behalf of an employer.

Description

About the Role

RadixArk is seeking a Member of Technical Staff — Inference to push the limits of large-scale AI inference.

You will work on the core systems that serve frontier models at scale, optimizing performance, latency, throughput, and cost across thousands of GPUs. This role sits at the intersection of systems engineering, ML infrastructure, and performance optimization.

Your work will directly shape how state-of-the-art models are deployed and experienced by users worldwide.

This is a deeply technical, high-impact role for engineers who enjoy working close to the hardware–software boundary and solving performance-critical problems at scale.

Requirements

  • 5+ years of experience in systems engineering, ML infrastructure, or performance-critical backend systems

  • Strong expertise in large-scale inference systems for LLMs or generative models

  • Deep understanding of GPU architecture and performance characteristics

  • Experience optimizing latency- and throughput-critical production systems

  • Strong knowledge of distributed systems and networking fundamentals

  • Proficiency in Python, Rust, C++, or Go for production systems

  • Experience profiling and optimizing compute-intensive workloads

  • Strong debugging skills across system layers (model, runtime, kernel, network)

Strong Plus

  • Experience with LLM serving stacks (SGLang, vLLM, TensorRT-LLM, etc.)

  • Open-source contributions in ML or systems infrastructure

  • Familiarity with CUDA, Triton, or custom kernel optimization

  • Experience with batching, KV-cache management, and scheduling strategies

  • Experience running inference at scale (1000+ GPUs)

  • Background in HPC or high-performance systems

Responsibilities

  • Design and build large-scale inference systems for frontier AI models

  • Optimize latency, throughput, and GPU utilization in production inference

  • Develop and improve model serving architectures and runtimes

  • Work on batching, scheduling, and memory management strategies

  • Collaborate with kernel, compiler, and systems teams on performance optimization

  • Debug performance bottlenecks across the stack

  • Drive reliability and scalability of inference infrastructure

  • Build tooling for observability, profiling, and performance analysis

  • Contribute to long-term inference architecture and strategy

About RadixArk

RadixArk is an infrastructure-first company built by engineers who've shipped production AI systems, created SGLang (30K+ GitHub stars, the fastest open LLM serving engine), and developed Miles (our large-scale RL framework). Founded by AI infrastructure veterans from xAI and NVIDIA, we're on a mission to democratize frontier-level AI infrastructure by building world-class open systems for inference and training. Our team has optimized kernels serving billions of tokens daily, designed distributed training systems coordinating 10,000+ GPUs, and contributed to infrastructure that powers leading AI companies and research labs.

Compensation

Depending on background, skills, and experience, the expected annual salary range for this position is $200,000 - $400,000 USD + equity.

Equal Opportunity

RadixArk is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more.

About this role

RadixArk is building inference infrastructure for frontier AI models at scale, and this role puts you at the center of that work. You'll own the core systems that serve large language models across thousands of GPUs, tackling the hard problems of latency, throughput, GPU utilization, and cost optimization. The work spans the full stack—from kernel-level performance tuning to distributed scheduling and memory management—and requires comfort moving fluidly between hardware constraints, runtime behavior, and system-level tradeoffs.

This is a fit for engineers with 5+ years shipping performance-critical systems, particularly those with hands-on experience optimizing LLM inference at scale. You'll need deep knowledge of GPU architecture, distributed systems, and production profiling; fluency in Python, Rust, C++, or Go; and the debugging chops to track down bottlenecks across model, runtime, kernel, and network layers. Prior work with serving stacks like vLLM or TensorRT-LLM, CUDA kernel optimization, or running inference across 1000+ GPUs is valuable. The team was founded by infrastructure veterans from xAI and NVIDIA who created SGLang and built systems serving billions of tokens daily, so you'd be working alongside people with proven depth in this space.

How this employer is doing

average

  • Sponsors Visas

This role's local market on JobMinglr

Pay for this role

The employer didn't post a pay range for this role. That usually means pay is set in negotiation, which favors whoever arrives with numbers. Check ranges on comparable Member of Technical Staff — Inference-Core Engine postings in Palo Alto, and analyze any offer before you accept it.

How JobMinglr reads this job

Every listing here is scored against your profile before you apply: skills overlap, experience level, location and work arrangement, each weighted and explained. You see the score and the reasons, not just a list. How the matching works.