Govworx
AI Evaluation Engineer
United States · Full-time
Sponsorship not specifiedDetected 5 days ago
PythonSQLAWSMachine LearningData ScienceData VisualizationLLMsLLMOpsA/B TestingManual TestingCommunication
About the role
- This role sits at the intersection of AI engineering, prompt engineering, and data science.
- You'll work closely with data scientists, data engineers, and product managers to ensure our AI systems remain accurate, reliable, and trustworthy in real-world public safety environments.
Responsibilities
- Design, build, and maintain automated AI evaluation pipelines for production LLM applications
- Develop prompt engineering strategies and iterate on prompts and compare LLMs using quantitative evaluation methods
- Build offline evaluation datasets and regression testing frameworks to measure AI performance over time
- Design experiments, A/B tests, and benchmarking methodologies for evaluating prompt and model changes
- Develop dashboards and reporting that communicate AI quality, reliability, and performance metrics
- Partner with engineering and product teams to safely deploy and monitor improvements to production AI systems
- You'll own the evaluation and continuous improvement of production AI systems, developing automated evaluation pipelines, designing prompt experiments, analyzing model performance, and building tooling that enables rapid iteration.
Requirements
- Must have US citizenship and pass FBI fingerprint and background check in multiple states
- Experience designing evaluation metrics and interpreting AI model performance
- Strong Python development experience
- Experience with prompt engineering and systematic prompt evaluation
- 3+ years of experience in software engineering, machine learning, data science, or a related technical field
- Understanding of statistical methods including hypothesis testing and experiment design
- Strong SQL skills with experience analyzing large datasets
- Experience building or supporting production LLM or Generative AI applications
Nice to have
- Experience using AI evaluation or observability platforms such as Langfuse, LangSmith, MLflow, or Label Studio
- Experience with AWS services such as Bedrock, Lambda, S3, Glue, or SageMaker
- Knowledge of Responsible AI principles and evaluation methodologies
- Work on cutting-edge LLM technologies and help shape the future of Responsible AI
- Solve technically challenging problems with real-world impact on public safety
- Influence AI strategy and evaluation practices across a growing technology company
- WHY JOIN GOVWORX?
Company info
- We're looking for an experienced AI Evaluation Engineer to help build and improve the next generation of AI systems used by public safety agencies across the country.
Visa & Work Authorization
- Clearance: Must have US citizenship and pass FBI fingerprint and background check in multiple states
Apply directly at Govworx →Create a free account for alerts like thisView Govworx immigration profile
This listing is sourced directly from Govworx's careers page and normalized into a canonical job model.