Lifted, an Upwork Company™

Lifted, an Upwork Company™

Senior Python Developer (AI Evaluation & Benchmarking)

Texas City, TX, United States · Senior

Sponsorship not specified$80k-$100kDetected 23 days ago
JavaScriptPythonGoC++Distributed SystemsCI/CDMachine LearningData EngineeringCybersecuritypytestMochaTest AutomationResearchProblem Solving

About the role

  • The selected consultants will contribute to AI research by designing programming benchmarks, evaluating AI-generated code, and helping improve the performance, reasoning, and reliability of frontier AI models.
  • This is an excellent opportunity for experienced software engineers who enjoy solving complex technical problems while contributing to the future of Generative AI.
  • This opportunity is ideal for senior software engineers with strong Python expertise who enjoy writing high-quality code, reviewing technical solutions, and working on AI-related projects.

Responsibilities

  • Design and develop coding benchmarks used to evaluate frontier AI models.
  • Analyze AI-generated code for correctness, reliability, efficiency, and edge cases.Build and maintain scalable data pipelines that support AI evaluation workflows.
  • Create structured programming scenarios to test reasoning, debugging, and code quality.
  • Collaborate with teams focused on improving how AI models understand, generate, and evaluate software.
  • An enterprise client is seeking experienced Senior Python Developers to help build the next generation of Artificial Intelligence systems.

Requirements

  • 4+ years of professional software engineering experience (required).
  • Expert-level proficiency in Python.
  • Experience working at a high-growth technology company or top-tier software organization.
  • Proficiency in at least one additional programming language such as JavaScript, Go, C++, or similar.
  • Experience with CI/CD pipelines and automated testing frameworks such as pytest, Mocha, or JUnit.
  • Strong understanding of software engineering best practices, debugging, and code quality.
  • Experience with AI/ML evaluation, model benchmarking, or Generative AI.
  • Experience working with large-scale distributed systems or enterprise software platforms.

Compensation

  • Experience with AI/ML evaluation, model benchmarking, or Generative AI.
  • Background in security engineering.
  • Significant contributions to open-source software projects.
  • Experience working with large-scale distributed systems or enterprise software platforms.
  • Fully remote contract opportunity.
  • Compensation ranges from $80-$100 USD per hour.

Company info

  • This opportunity support the client who is a leading AI platform that enables organizations to build intelligent applications through high-quality human feedback, AI evaluation, and model alignment.

This listing is sourced directly from Lifted, an Upwork Company™'s careers page and normalized into a canonical job model.