Altera

Altera

Senior MLOps & AI Infrastructure Engineer

San Jose, California, United States · Senior

Sponsorship not specified$149k-$216kDetected 20 days ago
PythonBashNode.jsSQLAWSGCPAzureDockerKubernetesTerraformCI/CDPrometheusGrafanaDevOpsMachine LearningDeep LearningTensorFlowPyTorchscikit-learnAirflowData EngineeringData ScienceLLMsRAG

About the role

  • We are looking for a Senior MLOps & AI Infrastructure Engineer to architect, build, and operationalize machine learning systems at scale. This role sits at the intersection of data science, software engineering, and infrastructure - combining deep ML expertise with the DevOps/MLOps discipline required to ship models reliably into production.
  • You will partner closely with software, data, and infrastructure teams to design end-to-end ML pipelines, automate model lifecycle management, and deliver AI-powered capabilities across our EDA, HPC, and cloud environments.
  • We are looking for a Senior MLOps & AI Infrastructure Engineer to architect, build, and operationalize machine learning systems at scale.

Responsibilities

  • Design, build, and maintain scalable ML pipelines for training, evaluation, and deployment across cloud and on-prem HPC environments
  • Build MLOps infrastructure including experiment tracking, model registry, feature stores, and automated retraining workflows
  • Implement CI/CD/CT (Continuous Training) pipelines for ML models using tools such as Kubeflow, MLflow, Airflow, or similar
  • Implement A/B testing frameworks and shadow deployments for safe model rollout
  • Benchmark and optimize model inference performance on GPU/TPU clusters
  • Build and maintain data pipelines for large-scale structured and unstructured datasets (terabyte-scale)
  • Collaborate with data teams to design feature engineering systems and maintain data quality for ML training
  • Implement data versioning and lineage tracking (DVC, Delta Lake, or similar)
  • Manage cloud ML infrastructure on AWS (SageMaker), Azure (AML), or GCP (Vertex AI) with cost and performance optimization
  • Support HPC schedulers (LSF, Slurm) for large-scale distributed training jobs

Requirements

  • 10+ years of experience across ML engineering, data science, and MLOps - including frameworks (PyTorch, TensorFlow, JAX, Hugging Face) and production model deployment at scale
  • 8+ years of experience experience with parallelism strategies (FSDP, DeepSpeed, data/model parallelism)
  • 10+ years of experience and proficiency in Python programming
  • 8+ years of experience in cloud ML platforms (AWS, GCP, Azure), Docker/Kubernetes, and CI/CD pipelines
  • 5+ years of hands-on experience with MLflow, W&B, or Neptune for tracking and reproducibility
  • Applicants must be eligible for any required U.S. export authorizations.
  • Bachelor's or Master's degree in Computer Science, Machine Learning, Statistics, or related field and 10+ years of industry experience

Nice to have

  • Familiarity with HPC schedulers such as LSF or Slurm and GPU cluster management for training workloads
  • Knowledge of LLM fine-tuning, Retrieval-Augmented Generation (RAG) architectures, and AI agent frameworks such as LangChain or AutoGen
  • Exposure to zero-trust security, DevSecOps, and compliance automation for ML systems
  • Experience working with large-scale simulation pipelines and synthetic data generation
  • Experience at organizations such as NVIDIA, AMD, Intel, Google DeepMind, or similar AI/HPC-focused companies
  • Published research or open-source contributions in ML, MLOps, or AI for EDA
  • Familiarity with Synopsys, Cadence, or Siemens EDA toolchains and associated data formats

Skills

  • About Altera
  • TensorFlow
  • Hugging Face
  • scikit-learn
  • XGBoost
  • Terraform
  • Docker
  • Kubernetes
  • Slurm / LSF

Compensation

  • The pay range below is for Bay Area California only.
  • Actual salary may vary based on a number of factors including job location, job-related knowledge, skills, experiences, trainings, etc.
  • We also offer incentive opportunities that reward employees based on individual and company performance.
  • $149,100 - $215,925 USD

Benefits

  • Build monitoring, alerting, and observability systems for model performance drift, data quality, and system health
  • Develop, fine-tune, and deploy large-scale models including LLMs, GNNs, and reinforcement learning agents for EDA and chip design applications
  • Apply advanced techniques: transfer learning, quantization, pruning, distillation, and RLHF for production-grade model efficiency

This listing is sourced directly from Altera's careers page and normalized into a canonical job model.