Govworx

Govworx

AI Evaluation Engineer

United States · Full-time

Sponsorship not specifiedDetected 5 days ago
PythonSQLAWSMachine LearningData ScienceData VisualizationLLMsLLMOpsA/B TestingManual TestingCommunication

About the role

  • This role sits at the intersection of AI engineering, prompt engineering, and data science.
  • You'll work closely with data scientists, data engineers, and product managers to ensure our AI systems remain accurate, reliable, and trustworthy in real-world public safety environments.

Responsibilities

  • Design, build, and maintain automated AI evaluation pipelines for production LLM applications
  • Develop prompt engineering strategies and iterate on prompts and compare LLMs using quantitative evaluation methods
  • Build offline evaluation datasets and regression testing frameworks to measure AI performance over time
  • Design experiments, A/B tests, and benchmarking methodologies for evaluating prompt and model changes
  • Develop dashboards and reporting that communicate AI quality, reliability, and performance metrics
  • Partner with engineering and product teams to safely deploy and monitor improvements to production AI systems
  • You'll own the evaluation and continuous improvement of production AI systems, developing automated evaluation pipelines, designing prompt experiments, analyzing model performance, and building tooling that enables rapid iteration.

Requirements

  • Must have US citizenship and pass FBI fingerprint and background check in multiple states
  • Experience designing evaluation metrics and interpreting AI model performance
  • Strong Python development experience
  • Experience with prompt engineering and systematic prompt evaluation
  • 3+ years of experience in software engineering, machine learning, data science, or a related technical field
  • Understanding of statistical methods including hypothesis testing and experiment design
  • Strong SQL skills with experience analyzing large datasets
  • Experience building or supporting production LLM or Generative AI applications

Nice to have

  • Experience using AI evaluation or observability platforms such as Langfuse, LangSmith, MLflow, or Label Studio
  • Experience with AWS services such as Bedrock, Lambda, S3, Glue, or SageMaker
  • Knowledge of Responsible AI principles and evaluation methodologies
  • Work on cutting-edge LLM technologies and help shape the future of Responsible AI
  • Solve technically challenging problems with real-world impact on public safety
  • Influence AI strategy and evaluation practices across a growing technology company
  • WHY JOIN GOVWORX?

Company info

  • We're looking for an experienced AI Evaluation Engineer to help build and improve the next generation of AI systems used by public safety agencies across the country.

Visa & Work Authorization

  • Clearance: Must have US citizenship and pass FBI fingerprint and background check in multiple states

This listing is sourced directly from Govworx's careers page and normalized into a canonical job model.