Ifm Us
Eval360 - Error Analysis Engineer
Sunnyvale, CA
H1B sponsorship available$150k-$450kDetected 32 days ago
JavaScriptTypeScriptPythonJavaC#ReactAngularHTMLCSSNode.jsDjangoFastAPIFlaskDistributed SystemsFull-Stack DevelopmentAlgorithmsGitPostgreSQLMySQLMongoDBCI/CDMachine LearningPandasNumPy
About the role
- This person will focus specifically on error analysis: understanding where models fail, why they fail, how those failures should be categorized, and how evaluation systems can better detect, measure, and prevent these issues before models are released.
- The ideal candidate is comfortable working across the stack, including front-end interfaces for reviewing errors, back-end evaluation pipelines, data analysis workflows, model evaluation infrastructure, databases, dashboards, and APIs.
Responsibilities
- Build and improve Eval360 as an evaluation service that acts as a quality gate for model development, model comparison, and model release decisions.
- Perform deep error analysis on model outputs, including identifying failure patterns, categorizing issues, tracing root causes, and proposing improvements to evaluation methodology.
- Develop tools, workflows, and dashboards that make it easier for researchers and engineers to inspect model failures, compare model behavior, and understand quality regressions.
- Design and implement client-side and server-side architecture for evaluation review systems, error analysis interfaces, reporting tools, and internal evaluation applications.
- Develop responsive, usable interfaces that support error triage, annotation review, evaluation debugging, and model quality investigation.
- Build and maintain back-end services, APIs, data pipelines, and integrations that support evaluation execution, results storage, analysis, and reporting.
- Create and maintain security, access control, and data protection settings for evaluation data, model outputs, annotations, and internal tooling.
- Contribute to the design of error taxonomies, evaluation rubrics, quality thresholds, regression detection methods, and model readiness criteria.
Requirements
- This person should have strong software engineering skills, excellent analytical judgment, and the ability to turn ambiguous model failures into structured insights that improve evaluation quality.
Nice to have
- Experience with large language models, foundation models, multimodal models, or model evaluation systems.
- Experience designing or using error taxonomies, evaluation rubrics, benchmark datasets, human evaluation workflows, or automated grading systems.
- Experience with Python-based data analysis tools such as pandas, NumPy, Jupyter, or similar.
- Experience with visualization or dashboarding tools for model quality analysis.
- Experience with distributed systems, job queues, workflow orchestration, or large-scale data processing.
- Experience working in a research environment or with fast-moving AI product and model teams.
- Visa Sponsorship This position is eligible for visa sponsorship.
- Employee assistance program
Compensation
- $150k-$450k
Benefits
- Comprehensive medical, dental, and vision benefits
- Generous paid time off, sick leave, and holidays
- Paid parental leave
- Life insurance and disability insurance
- Collaborate with researchers, machine learning engineers, data scientists, product managers, and internal stakeholders to implement innovative software solutions for Eval360 and related model evaluation workflows.
- Work with researchers, data scientists, analysts, and machine learning engineers to improve evaluation quality, model diagnostics, and failure-mode visibility.
Visa & Work Authorization
- This position is eligible for visa sponsorship.
This listing is sourced directly from Ifm Us's careers page and normalized into a canonical job model.