Skip to content
Job Details
Full-time

Senior AI Engineer

DV Trading · Chicago, IL +1 more

Tailor My Resume

Start free. No credit card.

About Us

Founded 20 years ago and headquartered in Chicago, the DV Group of financial services firms has grown to more than 600 people operating throughout North America, Europe and Asia. Since spinning out of a large brokerage firm in 2016, DV Trading has rapidly scaled as an independent proprietary trading firm utilizing its own capital, trading strategies, and risk management methodologies to provide liquidity to worldwide financial markets and hedging opportunities to commodity producers and users. Now, DV group affiliates include two broker dealers, a cryptocurrency market making firm, and a bourgeoning investment adviser.

Overview

DV Trading is building a centralized AI function and is now hiring for the model layer. The long-term goal is for DV to own its model capability — not to be permanently dependent on what frontier providers choose to offer, at what price, for how long. This role is how that happens: fine-tuning and distilling open-weight models for DV-specific tasks, operating the inference infrastructure to run them on-prem, and building the model gateway that routes intelligently across open and closed providers. The near-term result is lower cost and better latency. The long-term result is a firm that controls its own AI stack.

Job Responsibilities

  • Build and operate a model gateway routing inference across open and closed models with cost, latency, and quality tracking
  • Design and run distillation pipelines: use frontier model outputs to generate training data for task-specific open models
  • Fine-tune and evaluate open-weight models (Llama, Qwen, Mistral, or similar) for DV-specific tasks
  • Deploy and maintain on-prem inference infrastructure (vLLM, TGI, or equivalent) on KubernetesBuild model evaluation frameworks for quality, cost, latency, and regression
  • Define criteria and tooling for model selection: when open models are production-ready vs. when to use closed APIs
  • Partner with the agent engineering team to ensure the model layer meets agent workload

Requirements

  • 5+ years software engineering; strong Python
  • Production fine-tuning or distillation of open-weight models (not just inference API wrappers)
  • Experience serving LLMs on-prem (vLLM, TGI, Triton, or equivalent)
  • Experience managing GPU infrastructure (provisioning, scheduling, utilization monitoring) in a production environment
  • Model evaluation and regression testing in production
  • Kubernetes and GPU workload management
  • Strong grasp of the tradeoffs between open and closed models across cost, quality, latency, and data sensitivity

Preferred

  • Quantization, PEFT/LoRA, or other efficient training techniques
  • Model gateway or inference proxy design (routing, fallback, rate limiting)
  • Financial services or other regulated/sensitive-data environments
  • Familiarity with the open model ecosystem (Hugging Face, model cards, licensing

Benefits

  • Discretionary bonus eligibility
  • Medical, dental, and vision insurance
  • HSA, FSA, and Dependent Care Options
  • Employer Paid Group Term Life and AD&D insurance
  • Voluntary LTD, Life & AD&D insurance
  • Flexible Vacation policy
  • Retirement plan with employer match

Posted

4 days ago

Job Type

Full-time

Salary

$200,000 – $300,000 USD

Locations

  • Chicago, IL
  • New York, NY