
Senior ML Systems Engineer, Inference
Immigration summary
Visa sponsorship
This posting says it cannot offer visa sponsorship for this role.
24 of 24 recent openings here state a sponsorship restriction
View visa evidenceGreen card sponsorship
We didn't find recent certified PERM filings for this employer.
View green card evidenceJob description
Runpod is the AI Developer Cloud. More than one million developers, from indie researchers to teams running frontier models in production, use Runpod to experiment, train, fine-tune, deploy, and scale AI on one platform. The platform has processed more than 20 billion inference requests. We closed a $100M Series A in June 2026. We're at an inflection point for AI infrastructure, and we're building the platform the next generation of developers will depend on.
We're a small, remote-first team. We take ownership seriously, move fast, and ship work that more than a million developers rely on every day. We're looking for people who care deeply, build with urgency, and want to matter at scale.
Learn more in our CEO's funding announcement: https://www.runpod.io/blog/one-million-developers.
We're looking for a ML Systems Engineer, Inference. We want Runpod to be the best place in the world to run LLM inference, meaning the fastest and the most cost-efficient. You'll lead that effort. You'll own LLM serving performance end to end. That means measuring it, understanding it, and improving it across models, hardware generations, and workloads. The work you ship will show up directly in the latency and cost our customers experience. This is a hands-on engineering role for someone who likes finding the real bottleneck and fixing it, then turning that fix into something that runs reliably in production.
Responsibilities
Define how we measure inference performance, including throughput, time to first token, inter-token latency, and cost per token, and build the tooling that makes those measurements rigorous and repeatable.
Profile and diagnose performance problems across the serving stack, from scheduling and memory management down to kernels and interconnect.
Improve serving efficiency for large, state-of-the-art models on single-node and multi-node GPU deployments.
Turn what you learn into production-ready runtimes, configurations, and defaults that customers benefit from automatically.
Work closely with product and infrastructure teams to shape how inference is offered on Runpod.
Keep up with the fast-moving inference ecosystem, including the open-source community, and decide what's worth adopting, what's worth building, and what's worth contributing back.
Trace bottlenecks in the serving engine/runtime and implement fixes when configuration tuning is not enough.
Requirements
5+ years of professional system engineering experience.
Deep, hands-on experience with vLLM, SGLang (or a comparable serving engine) in production or at serious benchmark scale.
Strong software engineering skills in Python. You're comfortable working in large, performance-critical codebases.
A solid understanding of what drives LLM inference performance: batching, memory, parallelism, and the trade-offs between latency and throughput.
Experience with modern inference optimization techniques such as quantization, speculative decoding, or distributed serving.
Rigor in benchmarking and performance analysis, plus comfort with GPU profiling tools.
The ability to explain your results clearly in writing and turn them into decisions.
Preferred
Experience writing or tuning GPU kernels in CUDA or Triton.
Contributions to inference or ML systems projects.
Experience with multi-node GPU systems and high-speed networking.
Experience at a company where inference cost and latency were core business metrics.
What You’ll Receive:
The competitive base pay for this position ranges from ($150,000 - $220,000). This salary range may be inclusive of several career levels at Runpod and will be narrowed during the interview process based on a number of factors, including the candidate’s experience, qualifications, and location
Meaningful equity in a fast-growing company- everyone on the team receives stock options — your impact drives our growth, and you share in the upside.
Generous medical, dental & vision plans
Flexible PTO- take the time you need to recharge
Most roles are remote work first with an inclusive, collaborative teams utilizing slack as the main form of internal communication
Join a passionate team on the cutting edge of AI infrastructure — where culture, learning, and ownership are at the heart of how we scale.
$1,200 Home Office & Equipment Stipend- We set you up for success from day one with gear and support to create your ideal workspace
Runpod is committed to maintaining a workplace free from discrimination and upholding the principles of equality and respect for all individuals. We believe that diversity in all its forms enhances our team. As an equal opportunity employer, Runpod is committed to creating an inclusive workforce at every level. We evaluate qualified applicants without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, age, marital status, protected veteran status, disability status, or any other characteristic protected by law. We welcome every qualified candidate eligible to work in the United States; however, we are currently unable to sponsor employment visas.
Sponsorship evidence
Why Openbound reached the conclusions above.
Visa sponsorship evidence
Current posting
Sponsorship restriction stated
“Unable to sponsor”
Detected directly from this job posting.
Other openings
24 of 24 recent openings at this employer state a sponsorship restriction.
Employer H-1B history
- 2
- recent certified H-1B filings
- 2
- new-hire petitions
- 0
- filings for similar roles
- 1
- so far in FY2026
Filed titles like this role: data scientist · machine learning engineer
More evidence details
- 2 certified H-1B filings for closely related roles, though none for this exact title
- 2 recent certified H-1B filings across the employer
- Still filing this year — 1 filing in FY2026
- 24 other recent postings at this company state a sponsorship restriction
- The supporting filings are for related work, not for this exact role
- No filing activity is recorded for this job's location
Green card sponsorship evidence
Employer PERM history
No recent PERM filings on record
This employer's identity is verified, and no certified PERM case was found for it in FY2023, FY2024, FY2025.
Filing history reflects past employer behavior; it isn't a promise for this opening.
All open roles at RunpodHow Openbound evaluates sponsorship
Visa history uses official U.S. Department of Labor H-1B LCA disclosure data and USCIS H-1B petition history. Green card history uses DOL PERM disclosure data. Each is read for the employer as a whole, for roles like this one, and for this location, weighted toward the most recent fiscal years.
An employer is matched to its filing entities by verified legal name and reviewed aliases; a match is never made on a name resemblance alone. Where no verified entity can be matched, the page says so and draws no conclusion from the absence. 8 filing titles were examined for this employer.
What this posting states outranks history in both directions, and an employer's published policy outranks past filings. Filing history reflects past behavior; it is not a promise of sponsorship for this opening, and none of this is legal advice.