Member of Technical Staff — Training

radixark

Palo Alto4d ago

Seniority: Staff

About the role

<h2 data-start="322" data-end="343">About the Role</h2> RadixArk is seeking a Member of Technical Staff — Training to build and scale the systems that train frontier AI models. You will work on large-scale distributed training infrastructure for LLMs and generative models, pushing the limits of scale, efficiency, accuracy and reliability across 10k, or 100k+ of GPUs. This role sits at the intersection of ML, systems, and performance engineering. Your work will directly impact how next-generation AI models are trained and scaled. This is a deeply technical, high-impact role for engineers who enjoy solving hard systems problems at extreme scale. <h2 data-start="945" data-end="964">Requirements</h2> <ul data-start="966" data-end="1597"> <li data-start="966" data-end="1067"> 3+ years of experience in ML systems, or large-scale training infrastructure </li> <li data-start="966" data-end="1067"> Experience building or operating large-scale agentic post-training systems. </li> <li>Experience working on training / inference correctness or other precision-related problem</li> <li data-start="1310" data-end="1390"> Experience debugging performance and stability issues in large post-training jobs </li> <li>Experience improving training or inference efficiency.</li> </ul> <h3 data-start="1604" data-end="1623">Strong Plus</h3> <ul data-start="1625" data-end="2061"> <li data-start="1625" data-end="1679"> Experience training 100+ billion-parameter models </li> <li>Experience with train / inference optimization for large-scale RL or other production workload.</li> <li data-start="1680" data-end="1756"> Familiarity with training stacks (e.g. Megatron-LM, FSDP, torchtitan, etc.) and inference stack (e.g. SGLang, vLLM, etc.) </li> <li>Familiarity with post-training framework (e.g. Miles, Slime, veRL, Prime-RL, AReaL, etc.)</li> <li data-start="1757" data-end="1822"> Experience with RDMA, InfiniBand, NVLink, NCCL/RCCL, or high-speed GPU interconnects </li> <li data-start="1879" data-end="1931"> Contributions to ML systems open-source projects </li> <li data-start="1932" data-end="2003"> Experience with checkpointing, fault recovery, and elastic training. </li> <li data-start="1932" data-end="2003">Experience building infrastructure for agentic post-training, such as async rollout pipelines, sandbox, or harness system.</li> </ul> <h2 data-start="2068" data-end="2091">Responsibilities</h2> <ul data-start="2093" data-end="2680"> <li data-start="2093" data-end="2156"> Contribute to open-source large-scale post-training infrastructure Miles, and inference system SGLang. </li> <li data-start="2157" data-end="2218"> Optimize throughput, scalability, and hardware efficiency </li> <li data-start="2219" data-end="2293"> Improve reliability and fault tolerance for long-running training jobs </li> <li data-start="2294" data-end="2352"> Develop training frameworks and infrastructure tooling </li> <li data-start="2353" data-end="2423"> Collaborate with model researchers to support frontier experiments </li> <li data-start="2424" data-end="2481"> Debug and resolve cross-layer performance bottlenecks </li> <li data-start="2482" data-end="2554"> Build observability systems for training performance and reliability </li> <li data-start="2555" data-end="2617"> Drive capacity planning and cluster utilization strategies </li> <li data-start="2618" data-end="2680"> Contribute to long-term training infrastructure architecture </li> </ul> <h2 data-start="2687" data-end="2708">About RadixArk</h2> RadixArk is an infrastructure-first company built by engineers who've shipped production AI systems, created SGLang (20K+ GitHub stars, the fastest open LLM serving engine), and developed Miles (our large-scale RL framework). We're on a mission to democratize frontier-level AI infrastructure by building world-class open systems for inference and training. Our team has optimized kernels serving billions of tokens daily, designed distributed training systems coordinating 10,000+ GPUs, and contributed to infrastructure that powers leading AI companies and research labs. We're backed by well-known infrastructure investors and partner with Nvidia, Google, AWS, and frontier AI labs. Join us in building infrastructure that gives real leverage back to the AI community. <h2 data-start="3265" data-end="3284">Compensation</h2> We offer competitive compensation with meaningful equity, comprehensive benefits, and flexible work arrangements. Compensation depends on location, experience, and level. <h2 data-start="3463" data-end="3487">Equal Opportunity</h2> RadixArk is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more.

Perks & benefits

Equity Compensation

753,000+ hidden jobs like this

radixark and thousands of companies post here first — often days before LinkedIn or Indeed. Your first 5 applications are free; go Pro to apply without limits.

Everything Pro unlocks:

Unlimited applications — free stops at 5
Track every application in one place
Apply straight to the source, one click
Save & organize roles you love
Roles pulled from company boards before the big sites

Weekly

$9.99

$4.99/week

For an active search. Cancel anytime.

Get Weekly

Monthly

$24.99

$12.99/month

The smart pick. Save 35% vs weekly.

Get Monthly

Lifetime

$99

$49.99once

Pay once. Every future feature, forever.

Get Lifetime