Member of Technical Staff — Inference-Core Engine
RadixArk
Salary undisclosed
Need a reasonable accommodation to apply or interview? Contact us.
JobMinglr uses automated technology to recommend jobs based on profile information and job preferences. Match Score does not determine eligibility for a position, prevent a user from viewing or applying to a job, or make hiring decisions on behalf of an employer.
Description
About the Role
RadixArk is seeking a Member of Technical Staff — Inference to push the limits of large-scale AI inference.
You will work on the core systems that serve frontier models at scale, optimizing performance, latency, throughput, and cost across thousands of GPUs. This role sits at the intersection of systems engineering, ML infrastructure, and performance optimization.
Your work will directly shape how state-of-the-art models are deployed and experienced by users worldwide.
This is a deeply technical, high-impact role for engineers who enjoy working close to the hardware–software boundary and solving performance-critical problems at scale.
Requirements
5+ years of experience in systems engineering, ML infrastructure, or performance-critical backend systems
Strong expertise in large-scale inference systems for LLMs or generative models
Deep understanding of GPU architecture and performance characteristics
Experience optimizing latency- and throughput-critical production systems
Strong knowledge of distributed systems and networking fundamentals
Proficiency in Python, Rust, C++, or Go for production systems
Experience profiling and optimizing compute-intensive workloads
Strong debugging skills across system layers (model, runtime, kernel, network)
Strong Plus
Experience with LLM serving stacks (SGLang, vLLM, TensorRT-LLM, etc.)
Open-source contributions in ML or systems infrastructure
Familiarity with CUDA, Triton, or custom kernel optimization
Experience with batching, KV-cache management, and scheduling strategies
Experience running inference at scale (1000+ GPUs)
Background in HPC or high-performance systems
Responsibilities
Design and build large-scale inference systems for frontier AI models
Optimize latency, throughput, and GPU utilization in production inference
Develop and improve model serving architectures and runtimes
Work on batching, scheduling, and memory management strategies
Collaborate with kernel, compiler, and systems teams on performance optimization
Debug performance bottlenecks across the stack
Drive reliability and scalability of inference infrastructure
Build tooling for observability, profiling, and performance analysis
Contribute to long-term inference architecture and strategy
About RadixArk
RadixArk is an infrastructure-first company built by engineers who've shipped production AI systems, created SGLang (30K+ GitHub stars, the fastest open LLM serving engine), and developed Miles (our large-scale RL framework). Founded by AI infrastructure veterans from xAI and NVIDIA, we're on a mission to democratize frontier-level AI infrastructure by building world-class open systems for inference and training. Our team has optimized kernels serving billions of tokens daily, designed distributed training systems coordinating 10,000+ GPUs, and contributed to infrastructure that powers leading AI companies and research labs.
Compensation
Depending on background, skills, and experience, the expected annual salary range for this position is $200,000 - $400,000 USD + equity.
Equal Opportunity
RadixArk is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more.
About this role
RadixArk is building inference infrastructure for frontier AI models at scale, and this role puts you at the center of that work. You'll own the core systems that serve large language models across thousands of GPUs, tackling the hard problems of latency, throughput, GPU utilization, and cost optimization. The work spans the full stack—from kernel-level performance tuning to distributed scheduling and memory management—and requires comfort moving fluidly between hardware constraints, runtime behavior, and system-level tradeoffs.
This is a fit for engineers with 5+ years shipping performance-critical systems, particularly those with hands-on experience optimizing LLM inference at scale. You'll need deep knowledge of GPU architecture, distributed systems, and production profiling; fluency in Python, Rust, C++, or Go; and the debugging chops to track down bottlenecks across model, runtime, kernel, and network layers. Prior work with serving stacks like vLLM or TensorRT-LLM, CUDA kernel optimization, or running inference across 1000+ GPUs is valuable. The team was founded by infrastructure veterans from xAI and NVIDIA who created SGLang and built systems serving billions of tokens daily, so you'd be working alongside people with proven depth in this space.
How this employer is doing
average
- Sponsors Visas
This role's local market on JobMinglr
Pay for this role
The employer didn't post a pay range for this role. That usually means pay is set in negotiation, which favors whoever arrives with numbers. Check ranges on comparable Member of Technical Staff — Inference-Core Engine postings in Palo Alto, and analyze any offer before you accept it.
Before you apply, worth reading
How JobMinglr reads this job
Every listing here is scored against your profile before you apply: skills overlap, experience level, location and work arrangement, each weighted and explained. You see the score and the reasons, not just a list. How the matching works.