Loading jobs…
Loading jobs…
DV Trading Inc. — Chicago, Illinois
About Us : Founded 20 years ago and headquartered in Chicago, the DV Group of financial services firms has grown to more than 600 people operating throughout North America, Europe and Asia. Since spinning out of a large brokerage firm in 2016, DV Trading has rapidly scaled as an independent proprietary trading firm utilizing its own capital, trading strategies, and risk management methodologies to provide liquidity to worldwide financial markets and hedging opportunities to commodity producers and users. Now, DV group affiliates include two broker dealers, a cryptocurrency market making firm, and a bourgeoning investment adviser.
Overview: DV Trading is building a centralized AI function and is now hiring for the model layer. The long-term goal is for DV to own its model capability — not to be permanently dependent on what frontier providers choose to offer, at what price, for how long. This role is how that happens: fine-tuning and distilling open-weight models for DV-specific tasks, operating the inference infrastructure to run them on-prem, and building the model gateway that routes intelligently across open and closed providers.
The near-term result is lower cost and better latency. The long-term result is a firm that controls its own AI stack.
Responsibilities
: Build and operate a model gateway routing inference across open and closed models with cost, latency, and quality tracking Design and run distillation pipelines: use frontier model outputs to generate training data for task-specific open models Fine-tune and evaluate open-weight models (Llama, Qwen, Mistral, or similar) for DV-specific tasks Deploy and maintain on-prem inference infrastructure (vLLM, TGI, or equivalent) on KubernetesBuild model evaluation frameworks for quality, cost, latency, and regression Define criteria and tooling for model selection: when open models are production-ready vs.
Requirements
: 5+ years software engineering; strong Python Production fine-tuning or distillation of open-weight models (not just inference API wrappers) Experience serving LLMs on-prem (vLLM, TGI, Triton, or equivalent) Experience managing GPU infrastructure (provisioning, scheduling, utilization monitoring) in a production environment Model evaluation and regression testing in production Kubernetes and GPU workload management Strong grasp of the tradeoffs between open and closed models across cost, quality, latency, and data sensitivity Preferred: Quantization, PEFT/LoRA, or other efficient training techniques Model gateway or inference proxy design (routing, fallback, rate limiting) Financial services or other regulated/sensitive-data environments Familiarity with the open model ecosystem (Hugging Face, model cards, licensing
Benefits
: Discretionary bonus eligibility Medical, dental, and vision insurance HSA, FSA, and Dependent Care Options Employer Paid Group Term Life and AD D insurance Voluntary LTD, Life AD D insurance Flexible Vacation policy Retirement plan with employer match DV is not accepting unsolicited resumes from search firms. Only search firms with valid, written agreements with DV should submit resumes in response to DV ’s posted positions. All resumes submitted by search firms to DV via e-mail, the Internet, personal delivery, facsimile, or any other method without a valid written agreement shall be deemed the sole property of DV , and no fee will be paid in the event the candidate is hired by DV .
DV is proud to be an equal opportunity employer and committed to creating an inclusive environment for all employees. The range below reflects the expected base salary for this position.
Compensation
determined by your experience, education, skills, and performance throughout the interview process.