Member of Technical Staff — Inference-Kernel, Compiler & Communication
RadixArk
Salary undisclosed
Need a reasonable accommodation to apply or interview? Contact us.
JobMinglr uses automated technology to recommend jobs based on profile information and job preferences. Match Score does not determine eligibility for a position, prevent a user from viewing or applying to a job, or make hiring decisions on behalf of an employer.
Description
About the Role
RadixArk is seeking a Member of Technical Staff — Kernel / Compiler / Communication to push the limits of performance for frontier AI systems.
You will work at the lowest layers of the stack — kernels, runtimes, compilers, and communication libraries — to unlock maximum efficiency from modern accelerators and interconnects.
This role is critical to scaling training and inference across thousands of GPUs, where microseconds and memory bandwidth matter. Your work will directly shape the performance envelope of next-generation AI systems.
This is a deeply technical role for engineers who enjoy working close to hardware and solving performance problems that most engineers never encounter.
Requirements
5+ years of experience in systems, compiler, or performance engineering
Strong expertise in CUDA or accelerator programming
Deep understanding of GPU architecture and memory hierarchy
Experience writing or optimizing high-performance kernels
Strong background in compilers, runtimes, or code generation
Experience with distributed communication libraries (NCCL, MPI, RCCL, etc.)
Solid knowledge of networking and interconnect technologies
Proficiency in C++ and Python
Strong debugging and profiling skills at system level
Strong Plus
Experience with Triton, TVM, XLA, or MLIR
Experience building compiler passes or IR transformations
Familiarity with NVLink, InfiniBand, or RDMA
Experience optimizing collective communication at scale
Background in HPC or performance-critical systems
Contributions to kernel/compiler/ML systems open source
Experience scaling workloads to 1000+ GPUs
Experience with mixed-precision or quantized kernels
Responsibilities
Design and implement high-performance kernels for AI workloads
Optimize compiler and runtime stacks for ML systems
Improve communication efficiency across large GPU clusters
Reduce latency and increase throughput for distributed workloads
Profile and eliminate system bottlenecks across the stack
Collaborate with training and inference teams on performance optimization
Develop tooling for profiling and performance analysis
Contribute to long-term architecture for performance-critical systems
Push the limits of hardware–software co-design
About RadixArk
RadixArk is an infrastructure-first company built by engineers who've shipped production AI systems, created SGLang (30K+ GitHub stars, the fastest open LLM serving engine), and developed Miles (our large-scale RL framework). Founded by AI infrastructure veterans from xAI and NVIDIA, we're on a mission to democratize frontier-level AI infrastructure by building world-class open systems for inference and training. Our team has optimized kernels serving billions of tokens daily, designed distributed training systems coordinating 10,000+ GPUs, and contributed to infrastructure that powers leading AI companies and research labs.
Compensation
Depending on background, skills, and experience, the expected annual salary range for this position is $200,000 - $400,000 USD + equity.
Equal Opportunity
RadixArk is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more.
About this role
RadixArk is building infrastructure for frontier AI systems and needs engineers who can extract maximum performance from hardware at the lowest levels of the stack. In this role, you'd work on kernels, compilers, runtimes, and communication libraries—the components that determine whether training and inference scale efficiently across thousands of GPUs. Your focus would be on problems where microseconds and memory bandwidth translate directly into system capability: optimizing collective communication, reducing latency in distributed workloads, and eliminating bottlenecks that most engineers never see.
This is a deeply technical position suited to someone with 5+ years in systems or performance engineering who has shipped real CUDA kernels or compiler work. You should have strong GPU architecture knowledge, hands-on experience with high-performance kernel optimization, and familiarity with distributed communication libraries like NCCL or MPI. Proficiency in C++ and Python, combined with serious debugging and profiling skills, is essential. Experience with compiler frameworks like Triton, TVM, or MLIR, or background in HPC and large-scale workload optimization, would be particularly valuable.
The role is based in Palo Alto and requires in-office presence. RadixArk was founded by infrastructure veterans from xAI and NVIDIA who created SGLang and have optimized systems serving billions of tokens daily. Compensation ranges from $200,000 to $400,000 annually plus equity, depending on background and experience.
How this employer is doing
average
- Sponsors Visas
This role's local market on JobMinglr
Pay for this role
The employer didn't post a pay range for this role. That usually means pay is set in negotiation, which favors whoever arrives with numbers. Check ranges on comparable Member of Technical Staff — Inference-Kernel, Compiler & Communication postings in Palo Alto, and analyze any offer before you accept it.
Before you apply, worth reading
How JobMinglr reads this job
Every listing here is scored against your profile before you apply: skills overlap, experience level, location and work arrangement, each weighted and explained. You see the score and the reasons, not just a list. How the matching works.