Back to jobs
TikTok10K+ employees

Software Engineer Project Intern (Model Infrastructure) - 2026 Start (BS/MS)

San Jose, CASalary not listedAdded to Openbound 26 days ago

Immigration summary

Visa sponsorship

Highly likelyHigh confidence

This employer sponsors this kind of work repeatedly, and is still filing this year.

1,721 recent H-1B filings · 189 similar-role filings

View visa evidence

Green card sponsorship

Strong historyHigh confidence

This employer has recently sponsored green cards at scale.

218 recent certified PERM filings · 138 similar-role filings

View green card evidence

Job description

About the Team

The TikTok Model Infrastructure team is the core engine powering the world’s most engaged "For You" feed. We focus on the engineering efficiency and architectural evolution of recommendation models at an unprecedented scale. As we lead the industry’s shift toward LLM2Rec and Large Recommendation Models (LRM), our mission is to build ultra-high-performance infrastructure that bridges the gap between massive data scale and extreme algorithmic complexity.
We tackle the industry's most demanding "frontier" challenges: managing Petabyte-scale distributed embedding states, optimizing thousand-node GPU clusters, and perfecting real-time Sparse/Dense streaming. Our work ensures that models with hundreds of billions of dense parameters—on par with the world's largest LLMs—can operate with millisecond-level latency.

We are seeking Software Engineering Interns to join the Model Infra team to redefine the performance boundaries of recommendation systems. In this role, you will focus on the efficiency of the entire model lifecycle. You will work on the convergence of generative AI and recommendation architecture, optimizing everything from the raw throughput of multi-billion parameter dense blocks to the efficient retrieval of sparse features across massive distributed memory fabrics.

As a project intern, you will have the opportunity to engage in impactful short-term projects that provide you with a glimpse of professional real-world experience. You will gain practical skills through on-the-job learning in a fast-paced work environment and develop a deeper understanding of your career interests.

Applications will be reviewed on a rolling basis - we encourage you to apply early.

Responsibilities

  • Engineering Efficiency at Scale: Drive the optimization of training and inference pipelines to maximize hardware utilization (MFU/HFU) for models featuring hundreds of billions of dense parameters.
  • LLM2Rec Infrastructure: Architect specialized systems to support the integration of LLMs into the recommendation stack, focusing on memory-efficient attention mechanisms and advanced KV cache management for long-sequence user modeling.
  • Massive Sparse & Dense Streaming: Build and optimize high-concurrency engines for Petabyte-scale streaming training, handling continuous parameter updates and high-frequency data ingestion without compromising stability.
  • Hardware-Aware Co-Design: Work closely with researchers to design next-generation recommendation architectures optimized for modern GPU/NPU interconnects, ensuring high-bandwidth utilization across the cluster.
  • Distributed State Management: Innovate on how we store and synchronize massive model states across heterogeneous memory hierarchies (HBM, DDR, and NVMe).

Qualifications

Minimum Qualification(s)
- Currently pursuing an Undergraduate/Master in Software Development, Computer Science, Computer Engineering, or a related technical discipline.

- Strong programming skills in C++ and Python.

- Solid understanding of Computer Architecture and the GPU software stack (CUDA, Triton, or NCCL).

- Experience with deep learning frameworks (e.g., PyTorch, TensorFlow) and a desire to "look under the hood" of model execution runtimes.

- A strong interest in solving system-level bottlenecks in large-scale distributed environments.

Preferred Qualification(s)

  • Experience with Transformer-based architectures, 3D parallelism (TP/PP/DP).
  • Deep understanding of the torch.compile stack, including TorchDynamo (graph acquisition) and TorchInductor (lowering).
  • Hands-on experience writing high-performance kernels or optimizing collective communication (e.g., customizing NCCL/UCX).
  • Familiarity with RDMA networking, high-performance storage, or specialized Parameter Server architectures.
  • Success in programming competitions (ACM-ICPC) or contributions to prominent open-source AI infrastructure or high-performance computing projects.

Sponsorship evidence

Why Openbound reached the conclusions above.

Visa sponsorship evidence

Current posting

Silent on sponsorship

Other openings

35 of 4308 recent openings at this employer state a sponsorship restriction.

Employer H-1B history

1,721
recent certified H-1B filings
1,290
new-hire petitions
189
filings for similar roles
1,176
so far in FY2026

Filed titles like this role: software engineer · software engineer - · software development engineer · software engineer graduate

More evidence details
  • 189 certified H-1B filings for this same role, 1,290 new-hire petitions across the employer, and 901 new hires already this year
  • 189 certified H-1B filings for this same role
  • 1,721 recent certified H-1B filings across the employer
  • Still filing this year — 1,176 filings in FY2026
  • 362 USCIS H-1B new-employment approvals, counted separately from LCA filings
  • Strong filing activity in CA
  • 1,276 further USCIS approvals for extensions or transfers
  • 35 other recent postings at this company state a sponsorship restriction
  • The posting says nothing about sponsorship either way

Strong filing activity in CA.

Green card sponsorship evidence

Employer PERM history

218
recent certified PERM filings
138
filings for similar roles
160
filings in this location
Certified PERM filings by fiscal year
2023142
2024270
2025217
2026629YTD

Filing history reflects past employer behavior; it isn't a promise for this opening.

All open roles at TikTok
How Openbound evaluates sponsorship

Visa history uses official U.S. Department of Labor H-1B LCA disclosure data and USCIS H-1B petition history. Green card history uses DOL PERM disclosure data. Each is read for the employer as a whole, for roles like this one, and for this location, weighted toward the most recent fiscal years.

An employer is matched to its filing entities by verified legal name and reviewed aliases; a match is never made on a name resemblance alone. Where no verified entity can be matched, the page says so and draws no conclusion from the absence. 2,134 filing titles were examined for this employer.

What this posting states outranks history in both directions, and an employer's published policy outranks past filings. Filing history reflects past behavior; it is not a promise of sponsorship for this opening, and none of this is legal advice.