Knowtex
Applied ML Engineer
San Francisco
Sponsorship not specifiedDetected 140 days ago
PythonAWSCloud PlatformsCI/CDMachine LearningTensorFlowPyTorchNLPLLMsComplianceManual TestingHIPAAICD-10Research
About the role
- We're at an inflection point where cutting-edge AI meets real clinical impact, giving clinicians hours back each day to focus on what matters most - their patients.
- This role bridges research and engineering - transforming models into reliable, low-latency, production-grade systems deployed across enterprise healthcare environments.
Responsibilities
- Optimize inference pipelines for low latency and high throughput
- Build automated evaluation and regression testing frameworks for LLM outputs
- Implement monitoring systems for model performance and drift detection
- Collaborate with Backend teams to integrate ML services into APIs and workflows
- Support specialty-level model evaluation and performance analysis
Requirements
- Strong proficiency in Python and PyTorch (or TensorFlow)
- Experience deploying ML models in production environments
- Familiarity with transformer architectures and large language models
- Experience with model optimization techniques (quantization, distillation, pruning)
Nice to have
- Experience with speech recognition systems or NLP pipelines
- Experience with Triton Inference Server or similar deployment frameworks
- Experience working in regulated environments (HIPAA, GovCloud, etc.)
- Python, PyTorch / TensorFlow
- Transformer-based LLM architectures
- AWS (SageMaker, ECS, Lambda, S3)
- CI/CD pipelines for ML deployment
- Observability tools for performance and drift monitoring
Compensation
- Meaningful equity compensation
Benefits
- 3-7+ years of experience in machine learning engineering or applied ML roles
Apply directly at Knowtex →Create a free account for alerts like thisView Knowtex immigration profile
This listing is sourced directly from Knowtex's careers page and normalized into a canonical job model.