Drweng
HPC Specialist
Montreal
Sponsorship not specifiedDetected 23 days ago
PythonBashNode.jsDistributed SystemsKubernetesTerraformAnsibleLinuxPrometheusGrafanaDevOpsSite Reliability EngineeringMachine LearningDeep LearningLLMsEmbedded SystemsSystems EngineeringTCP/IPFirewallResearchCommunication
About the role
- DRW is a diversified trading firm with over 3 decades of experience bringing sophisticated technology and exceptional people together to operate in markets around the world.
- Headquartered in Chicago with offices throughout the U.S., Canada, Europe, and Asia, we trade a variety of asset classes including Fixed Income, ETFs, Equities, FX, Commodities and Energy across all major global markets.
Responsibilities
- Deploy, maintain, and optimize GPU infrastructure for large-scale LLM inference workloads, including provisioning, configuration, and deployment of GPU server fleets.
- Architect and implement distributed serving solutions for multi-node, multi-GPU model deployments.
- Manage GPU-enabled Kubernetes clusters for LLM and ML workloads.
- Implement and optimize storage solutions for model weights and inference caches.
- Collaborate with ML engineers to profile model performance and implement inference acceleration techniques.
- Drive reliability improvements through monitoring, alerting, capacity planning, and incident response.
Requirements
- Bachelor's or Master's degree in Computer Science, Systems Engineering, or related field.
- 5+ years in DevOps, SRE, or infrastructure engineering roles.
- Strong experience with GPU infrastructure, model serving frameworks (vLLM, SGLang), and GPU driver management.
- Experience with infrastructure as code tools (Ansible, Terraform, or similar).
- Strong understanding of distributed systems, networking protocols (TCP/IP, HTTP/2), and load balancing.
- Proficiency in Python and Bash scripting for automation.
- Experience with monitoring and observability tools (Prometheus, Grafana, or similar).
- We have also leveraged our expertise and technology to expand into three non-traditional strategies: real estate, venture capital and cryptoassets.
Skills
- Research and evaluate emerging GPU technologies, model serving frameworks, and infrastructure optimizations.
This listing is sourced directly from Drweng's careers page and normalized into a canonical job model.