Skip to content
Job Details
Full-time

Research Scientist in Multimodal Interaction and World Model

ByteDance · San Jose, CA

Tailor My Resume

Start free. No credit card.

About the team

The Seed Multimodal Interaction and World Model team is dedicated to developing models that have human-level multimodal understanding and interaction capabilities. The team is working to advance the exploration and development of multimodal assistant products.

Responsibilities

  • Develop multimodal foundation models integrating vision, language, audio, and environment signals.
  • Design and optimize world models for reasoning, planning, and interaction.
  • Build training pipelines including data curation, alignment, and reinforcement learning.
  • Improve agent capabilities such as perception, memory, decision-making, and tool use.
  • Explore next-generation interaction paradigms between humans and intelligent systems.

Minimum Qualifications

  • Currently pursuing a Bachelor's or Master's degree in computer science, mathematics, engineering, or a related field, with an expected graduation date in 2027 and the ability to commit to an onboarding date by the end of 2027.
  • Excellent coding ability, data structures, and fundamental algorithm skills, proficient in C/C++ or Python, etc.
  • Demonstrated interest or project experience in relevant areas.

Preferred Qualifications

  • Experience in multimodal learning, reinforcement learning, or agent systems through internships is preferred.
  • Strong problem-solving and collaboration skills.

Posted

3 days ago

Job Type

Full-time

Location

  • San Jose, CA