Ifm Us

Ifm Us

Eval360 - Error Analysis Engineer

Sunnyvale, CA

H1B sponsorship available$150k-$450kDetected 32 days ago
JavaScriptTypeScriptPythonJavaC#ReactAngularHTMLCSSNode.jsDjangoFastAPIFlaskDistributed SystemsFull-Stack DevelopmentAlgorithmsGitPostgreSQLMySQLMongoDBCI/CDMachine LearningPandasNumPy

About the role

  • This person will focus specifically on error analysis: understanding where models fail, why they fail, how those failures should be categorized, and how evaluation systems can better detect, measure, and prevent these issues before models are released.
  • The ideal candidate is comfortable working across the stack, including front-end interfaces for reviewing errors, back-end evaluation pipelines, data analysis workflows, model evaluation infrastructure, databases, dashboards, and APIs.

Responsibilities

  • Build and improve Eval360 as an evaluation service that acts as a quality gate for model development, model comparison, and model release decisions.
  • Perform deep error analysis on model outputs, including identifying failure patterns, categorizing issues, tracing root causes, and proposing improvements to evaluation methodology.
  • Develop tools, workflows, and dashboards that make it easier for researchers and engineers to inspect model failures, compare model behavior, and understand quality regressions.
  • Design and implement client-side and server-side architecture for evaluation review systems, error analysis interfaces, reporting tools, and internal evaluation applications.
  • Develop responsive, usable interfaces that support error triage, annotation review, evaluation debugging, and model quality investigation.
  • Build and maintain back-end services, APIs, data pipelines, and integrations that support evaluation execution, results storage, analysis, and reporting.
  • Create and maintain security, access control, and data protection settings for evaluation data, model outputs, annotations, and internal tooling.
  • Contribute to the design of error taxonomies, evaluation rubrics, quality thresholds, regression detection methods, and model readiness criteria.

Requirements

  • This person should have strong software engineering skills, excellent analytical judgment, and the ability to turn ambiguous model failures into structured insights that improve evaluation quality.

Nice to have

  • Experience with large language models, foundation models, multimodal models, or model evaluation systems.
  • Experience designing or using error taxonomies, evaluation rubrics, benchmark datasets, human evaluation workflows, or automated grading systems.
  • Experience with Python-based data analysis tools such as pandas, NumPy, Jupyter, or similar.
  • Experience with visualization or dashboarding tools for model quality analysis.
  • Experience with distributed systems, job queues, workflow orchestration, or large-scale data processing.
  • Experience working in a research environment or with fast-moving AI product and model teams.
  • Visa Sponsorship This position is eligible for visa sponsorship.
  • Employee assistance program

Compensation

  • $150k-$450k

Benefits

  • Comprehensive medical, dental, and vision benefits
  • Generous paid time off, sick leave, and holidays
  • Paid parental leave
  • Life insurance and disability insurance
  • Collaborate with researchers, machine learning engineers, data scientists, product managers, and internal stakeholders to implement innovative software solutions for Eval360 and related model evaluation workflows.
  • Work with researchers, data scientists, analysts, and machine learning engineers to improve evaluation quality, model diagnostics, and failure-mode visibility.

Visa & Work Authorization

  • This position is eligible for visa sponsorship.

This listing is sourced directly from Ifm Us's careers page and normalized into a canonical job model.