Member of Technical Staff — Reliability-CI Infrastructure
RadixArk
Salary undisclosed
Need a reasonable accommodation to apply or interview? Contact us.
JobMinglr uses automated technology to recommend jobs based on profile information and job preferences. Match Score does not determine eligibility for a position, prevent a user from viewing or applying to a job, or make hiring decisions on behalf of an employer.
Description
About the Role
RadixArk is hiring a Member of Technical Staff — CI / Infrastructure to own the infrastructure that keeps SGLang moving. Our CI system runs 300+ GPU tests across NVIDIA, AMD, Intel, and Ascend hardware pools, gating every commit to one of the fastest-growing open-source LLM inference engines. When CI is green and fast, 100+ contributors ship with confidence. When it isn't, the entire project stalls. That bottleneck is your problem to solve.
You won't just maintain pipelines — you'll architect them. You'll replace brittle static thresholds with regression-based detection, harden runners against supply-chain attacks from fork PRs, and cut cycle times so contributors get feedback in minutes, not hours. You'll work directly with core maintainers, hardware partners, and the open-source community to keep the system that gates every merge request trustworthy, fast, and secure.
This is not a role for someone who wants to write CI YAML and walk away. It's for an engineer who treats CI infrastructure the way we treat serving infrastructure — as a system worth designing well.
What You’ll Do
- Own CI reliability end-to-end — triage failures, distinguish real regressions from flaky tests and infra issues, keep main green
- Build regression-based CI — replace hardcoded static thresholds with automated baseline comparison (metrics pipeline, durable storage, detection logic)
- Harden runner infrastructure — ephemeral runners, container isolation, security hardening for fork PR execution
- Cut CI time — right-size eval suites, deduplicate server startups, separate PR smoke tests from nightly full runs
- Improve developer experience — faster feedback, clearer failure messages, workflow orchestration
Requirements
- 3+ years operating CI/CD at scale (GitHub Actions, Buildkite, Jenkins, GitLab CI, or similar)
- Deep Linux, Docker, GPU computing knowledge
- Self-hosted runner management experience
- Strong Bash and Python
- Security mindset — CI supply chain risks, fork PR attack vectors, runner hardening
- NVIDIA GPU drivers, CUDA, NCCL, InfiniBand/RDMA experience in CI contexts
- Familiarity with ML inference workloads (model loading, KV cache, quantization)
Nice to Have
- Large open-source project CI experience (100+ contributors)
- AMD ROCm or Intel XPU CI pipelines
What Success Looks Like
- Day 20 — Full CI landscape understood, daily triage taken over, top recurring flaky tests fixed, PR CI time reduced 30%+
- Day 40 — Regression-based checks live on nightly CI, ephemeral runner prototype deployed, runner isolation in place
- Day 60 — Zero flaky tests. Main CI 100% green when no real regression exists
How to Apply
Reach out via Slack or email. CI fix PRs to major open-source projects are worth more than a resume.
About RadixArk
RadixArk is an infrastructure-first company built by engineers who've shipped production AI systems, created SGLang (30K+ GitHub stars, the fastest open LLM serving engine), and developed Miles (our large-scale RL framework). Founded by AI infrastructure veterans from xAI and NVIDIA, we're on a mission to democratize frontier-level AI infrastructure by building world-class open systems for inference and training. Our team has optimized kernels serving billions of tokens daily, designed distributed training systems coordinating 10,000+ GPUs, and contributed to infrastructure that powers leading AI companies and research labs.
Compensation
Depending on background, skills, and experience, the expected annual salary range for this position is $200,000 - $400,000 USD + equity.
Equal Opportunity
RadixArk is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more.
About this role
RadixArk's CI infrastructure gates every commit to SGLang, one of the fastest-growing open-source LLM inference engines, running 300+ GPU tests across NVIDIA, AMD, Intel, and Ascend hardware. This role owns that system end-to-end — not just maintaining pipelines, but architecting them to be fast, reliable, and secure. You'll work directly with core maintainers and hardware partners to keep the infrastructure that 100+ contributors depend on trustworthy and responsive. When CI is slow or flaky, the entire project stalls; when it's green, the team ships with confidence.
The work spans the full stack: triaging failures to distinguish real regressions from infrastructure issues, building regression-based detection to replace static thresholds, hardening ephemeral runners against supply-chain attacks from fork PRs, and cutting cycle times so developers get feedback in minutes. You'll need 3+ years operating CI/CD at scale (GitHub Actions, Buildkite, Jenkins, or similar), deep Linux and Docker expertise, strong Bash and Python, and hands-on experience with GPU computing — NVIDIA drivers, CUDA, NCCL, and ideally some familiarity with ML inference workloads. A security mindset is essential; you'll be thinking about CI supply chain risks and fork PR attack vectors from day one.
This is an in-office role in Palo Alto. Success means understanding the full CI landscape within three weeks, reducing PR cycle time by 30%, deploying regression-based checks and runner isolation within six weeks, and eliminating flaky tests within two months. RadixArk was founded by AI infrastructure veterans from xAI and NVIDIA; the team has optimized kernels serving billions of tokens daily and designed systems coordinating thousands of GPUs. The salary range is $200,000–$400,000 USD plus equity.
How this employer is doing
average
- Sponsors Visas
This role's local market on JobMinglr
Pay for this role
The employer didn't post a pay range for this role. That usually means pay is set in negotiation, which favors whoever arrives with numbers. Check ranges on comparable Member of Technical Staff — Reliability-CI Infrastructure postings in Palo Alto, and analyze any offer before you accept it.
Before you apply, worth reading
How JobMinglr reads this job
Every listing here is scored against your profile before you apply: skills overlap, experience level, location and work arrangement, each weighted and explained. You see the score and the reasons, not just a list. How the matching works.