Lifted, an Upwork Company™
Senior Python Developer (AI Evaluation & Benchmarking)
Texas City, TX, United States · Senior
Sponsorship not specified$80k-$100kDetected 23 days ago
JavaScriptPythonGoC++Distributed SystemsCI/CDMachine LearningData EngineeringCybersecuritypytestMochaTest AutomationResearchProblem Solving
About the role
- The selected consultants will contribute to AI research by designing programming benchmarks, evaluating AI-generated code, and helping improve the performance, reasoning, and reliability of frontier AI models.
- This is an excellent opportunity for experienced software engineers who enjoy solving complex technical problems while contributing to the future of Generative AI.
- This opportunity is ideal for senior software engineers with strong Python expertise who enjoy writing high-quality code, reviewing technical solutions, and working on AI-related projects.
Responsibilities
- Design and develop coding benchmarks used to evaluate frontier AI models.
- Analyze AI-generated code for correctness, reliability, efficiency, and edge cases.Build and maintain scalable data pipelines that support AI evaluation workflows.
- Create structured programming scenarios to test reasoning, debugging, and code quality.
- Collaborate with teams focused on improving how AI models understand, generate, and evaluate software.
- An enterprise client is seeking experienced Senior Python Developers to help build the next generation of Artificial Intelligence systems.
Requirements
- 4+ years of professional software engineering experience (required).
- Expert-level proficiency in Python.
- Experience working at a high-growth technology company or top-tier software organization.
- Proficiency in at least one additional programming language such as JavaScript, Go, C++, or similar.
- Experience with CI/CD pipelines and automated testing frameworks such as pytest, Mocha, or JUnit.
- Strong understanding of software engineering best practices, debugging, and code quality.
- Experience with AI/ML evaluation, model benchmarking, or Generative AI.
- Experience working with large-scale distributed systems or enterprise software platforms.
Compensation
- Experience with AI/ML evaluation, model benchmarking, or Generative AI.
- Background in security engineering.
- Significant contributions to open-source software projects.
- Experience working with large-scale distributed systems or enterprise software platforms.
- Fully remote contract opportunity.
- Compensation ranges from $80-$100 USD per hour.
Company info
- This opportunity support the client who is a leading AI platform that enables organizations to build intelligent applications through high-quality human feedback, AI evaluation, and model alignment.
Apply directly at Lifted, an Upwork Company™ →Create a free account for alerts like thisView Lifted, an Upwork Company™ immigration profile
This listing is sourced directly from Lifted, an Upwork Company™'s careers page and normalized into a canonical job model.