Skip to content
Job Details
Full-time

Software Engineer Graduate

ByteDance · San Jose, CA

Tailor My Resume

Start free. No credit card.

About the Job

The Speech team's mission is to empower interaction and creation using speech & audio related technologies. The team focuses on cutting-edge R&D in areas like speech & audio, music processing, natural language understanding and multimodal deep learning. We are looking for top talents to work on these exciting technologies, integrate them into various products and ultimately bring joy to our global user base!

We are looking for talented individuals to join our team. As a graduate, you will get opportunities to pursue bold ideas, tackle complex challenges, and unlock limitless growth. Successful candidates must be able to commit to an onboarding date by the end of the year. Please state your availability and graduation date clearly in your resume. Candidates can apply to a maximum of two positions and will be considered for jobs in the order you apply. The application limit is applicable to our Company and its affiliates' jobs globally. Applications will be reviewed on a rolling basis - we encourage you to apply early.

We are seeking a passionate AI Model Optimization Engineer to join our team. In this role, you will design and implement cutting-edge techniques to make AI models faster, more efficient, and easier to deploy at scale. You will collaborate across research and engineering to push the limits of AI performance in production environments.

Responsibilities

  • Develop and implement algorithms for model optimization, including quantization, pruning, knowledge distillation, and efficient architectures.
  • Build and maintain performance benchmarking frameworks for large-scale training and inference.
  • Optimize training and inference pipelines on GPUs and across distributed systems.
  • Collaborate with ML researchers to transition optimized models into production.
  • Stay current with the latest research in model efficiency, compilers, and systems.

Minimum Qualifications

  • Individuals who are completing or have recently completed a Bachelor's/ Master's degree in computer engineering or a related discipline.
  • Strong coding skills in Python and C++.

Preferred Qualifications

  • Experience with deep learning frameworks and distributed training systems.
  • Solid understanding of computer architecture, parallel computing, and GPU acceleration.
  • Familiarity with GPU programming (CUDA, Triton, or similar) is a plus.
  • Familiarity with ML compilers (e.g., TVM, XLA, TensorRT) is a plus.
  • Strong analytical skills and ability to work in a fast-paced team environment.

As a condition of employment, all successful candidates must be able to establish authorization to work in the United States. For this position, the Company does not provide sponsorship or any immigration-related benefits.

Posted

3 days ago

Job Type

Full-time

Location

  • San Jose, CA