Ifm Us

Ifm Us

Inference Optimization Intern – Performance Modeling

Sunnyvale, CA · Intern · Internship

Sponsorship not specifiedDetected 27 days ago
C++SassMachine LearningDeep LearningPyTorchNLPSystems EngineeringElectrical EngineeringResearchCommunication

About the role

  • About the Institute of Foundation Models The Institute of Foundation Models is dedicated to advancing the science and engineering of large-scale AI systems.
  • This internship provides hands-on experience in low-level GPU performance analysis, kernel optimization, and hardware-aware inference acceleration.
  • This intensive internship offers a unique opportunity to contribute to the development of a simulator and profiling framework for foundation model inference on NVidia GPUs.

Responsibilities

  • Develop analytical performance models for GPU kernels and inference workloads.
  • Build and validate a simulator to estimate theoretical hardware performance limits.
  • Collaborate with researchers and engineers to optimize inference kernels for transformer-based models.
  • Design profiling methodologies for Hopper and Blackwell architectures.
  • Document findings and provide actionable recommendations for performance improvements.

Requirements

  • Currently pursuing a degree in Computer Science, Computer Engineering, Electrical Engineering, Artificial Intelligence, High-Performance Computing, or a related quantitative discipline.

Nice to have

  • Experience with CUDA programming and GPU kernel development.
  • Understanding of NVIDIA GPU architecture and memory hierarchy.
  • Familiarity with performance profiling tools such as Nsight Systems and Nsight Compute.
  • Knowledge of PTX, SASS, and low-level GPU execution.
  • Experience optimizing CUDA kernels for throughput and latency.
  • Understanding of roofline analysis, performance modeling, and hardware utilization metrics.
  • Strong programming skills in C++, CUDA, and Python.
  • Performance engineering mindset.

This listing is sourced directly from Ifm Us's careers page and normalized into a canonical job model.