Back to jobs
Innodata1K–5K employees

Research Scientist, Video & Multimodal

Remote — US$160,000 – $185,000Posted yesterday

Immigration summary

Visa sponsorship

PossibleLow confidence

There is real but limited sponsorship history here.

2 recent H-1B filings

View visa evidence

Green card sponsorship

UnknownNo recent PERM history

We didn't find recent certified PERM filings for this employer.

View green card evidence

Job description

Innodata (Nasdaq: INOD) is a global data engineering company. We believe that data and Artificial Intelligence (AI) are inextricably linked. Our mission is to enable the responsible advancement of artificial intelligence by providing the data, evaluation frameworks, and human expertise required to build AI systems that can be trusted at scale. We provide a range of transferable solutions, platforms, and services for Generative AI / AI builders and adopters. In every relationship, we honor our 36+ year legacy delivering the highest quality data and outstanding outcomes for our customers.

Scope of the Role:

Video is where multimodal models are weakest and hardest to grade. Temporal reasoning, long-form understanding, grounding events in time, and holding audio, video, and text together do not fall out of image benchmarks — and the evaluations for them are still immature. Closing that gap is gated as much by how we design data and evaluation as by architecture. Innodata builds that data and those evaluations for the customers and frontier labs advancing video and multimodal models, and we are hiring a Research Scientist to own the science behind it.

You will partner directly with the customers and frontier labs building video understanding, video-language, and video-generation models, as interested in the data behind them as in the models themselves. Video spans two model families judged in completely different ways: models that understand video — answering questions, localizing events, grounding language in time — where the question is whether the answer is correct; and models that generate it, where fidelity, temporal coherence, and physical plausibility matter and no automatic metric is settled. You own the evaluation science for both, and knowing when model-based scoring can stand in for a human versus when it can't. Your conclusions shape what our partners measure and collect next.

What You’ll Own:

You will define how Innodata designs, structures, and evaluates video data for video and multimodal models, and you will validate those choices experimentally. Concretely, you will:

  • Translate the requirements of video and multimodal models — video understanding, temporal and event localization, action recognition, long-form video, video-language models, video generation, cross-modal reasoning, and multimodal retrieval and grounding — into concrete data specifications: modalities, annotation schemas, sampling, and evaluation criteria.
  • Build evaluation methodology for video understanding — temporal grounding accuracy, long-context and long-horizon reasoning, and dynamic multi-turn, cross-modal, and retrieval-and-grounding evaluation — clear about when model-based scoring is trustworthy and when a human is needed.
  • Build evaluation methodology for video generation — fidelity, temporal coherence, and physical plausibility, including generative video used as a world model — the regime where automatic metrics are weakest and human judgment matters most.
  • Decide how existing and incoming video should be structured, enriched, and sampled to extract the most model value from it, including from messy, domain-specific footage.
  • Run experiments that prove data decisions matter: fine-tune and evaluate models on Innodata data, with ablations tying specific data choices to measurable improvement.
  • Design adversarial and stumping evaluations that surface where video and multimodal systems fail, and turn those failures into better data.
  • Publish. Turn what you learn into benchmarks, methodology, and papers that advance the field and earn the trust of the customers and frontier labs we partner with.
  • Work with annotation teams, subject-matter experts, and the synthetic-data pipeline to turn specifications into operational collection and labeling plans.

You’ll Thrive in This Role If You Have:

  • Roughly 5+ years of hands-on industry experience in video understanding or multimodal ML. We weight practical experience over formal credentials; a PhD with a compelling, current research agenda can offset the lower end.
  • A Bachelor's degree in computer science, electrical engineering, or a related technical or quantitative field is required; an advanced degree (MS or PhD) in a relevant field is preferred.
  • Trained and evaluated video or multimodal models yourself, with strong PyTorch fundamentals.
  • Fluency in the formats and tooling video work runs on: ffmpeg and decord pipelines, temporal and COCO-style annotation, WebDataset, Parquet and Arrow, and HuggingFace datasets.
  • Experience fine-tuning large video or vision-language models with the modern toolchain (HuggingFace transformers, PEFT, efficient inference), and with long-form video, streaming, temporal segmentation, or synthetic video generation.
  • A way of thinking in datasets and benchmarks: you have built evaluation sets, calibrated difficulty, and argued about what makes video data good for a given objective.
  • A track record the field recognizes: first-author publications or strong open-source contributions at venues such as CVPR, ICCV, ECCV, NeurIPS, or ICLR.
  • The ability to work directly with the research scientists at the customers and frontier labs we partner with, and to explain data and modeling decisions clearly to both expert and non-expert audiences, backed by a rigorous, reproducible approach to experiments and documentation.
  • Bonus: interest or hands-on experience in responsible-AI evaluation and red-teaming — safety and robustness testing for video and multimodal systems.

The expected salary range for this position is $160,000 - $185,000 p/year, based on experience, skills, and qualifications.

Please be aware of recruitment scams involving individuals or organizations falsely claiming to represent employers. Innodata will never ask for payment, banking details, or sensitive personal information during the application process. To learn more on how to recognize job scams, please visit the Federal Trade Commission’s guide at https://consumer.ftc.gov/articles/job-scams.

If you believe you’ve been targeted by a recruitment scam, please report it to Innodata at verifyjoboffer@innodata.com and consider reporting it to the FTC at ReportFraud.ftc.gov.

Sponsorship evidence

Why Openbound reached the conclusions above.

Visa sponsorship evidence

Current posting

Silent on sponsorship

Employer H-1B history

2
recent certified H-1B filings
2
new-hire petitions
0
filings for similar roles
2
so far in FY2026
More evidence details
  • 2 recent certified H-1B filings across the employer
  • Still filing this year — 2 filings in FY2026
  • We checked 4 filing titles for this employer and none describe work like this role
  • The posting says nothing about sponsorship either way
  • No filing activity is recorded for this job's location

Green card sponsorship evidence

Employer PERM history

No recent PERM filings on record

This employer's identity is verified, and no certified PERM case was found for it in FY2023, FY2024, FY2025.

Filing history reflects past employer behavior; it isn't a promise for this opening.

All open roles at Innodata
How Openbound evaluates sponsorship

Visa history uses official U.S. Department of Labor H-1B LCA disclosure data and USCIS H-1B petition history. Green card history uses DOL PERM disclosure data. Each is read for the employer as a whole, for roles like this one, and for this location, weighted toward the most recent fiscal years.

An employer is matched to its filing entities by verified legal name and reviewed aliases; a match is never made on a name resemblance alone. Where no verified entity can be matched, the page says so and draws no conclusion from the absence. 4 filing titles were examined for this employer.

What this posting states outranks history in both directions, and an employer's published policy outranks past filings. Filing history reflects past behavior; it is not a promise of sponsorship for this opening, and none of this is legal advice.