Innopeaktech

Innopeaktech

Test Engineer-AI/LLM

Palo Alto, California, United States · Contract

Sponsorship not specified$100k-$200kDetected 388 days ago
PythoniOSAndroidGitSQLNoSQLAWSGCPAzureCloud PlatformsHelmMachine LearningData AnalysisData EngineeringData ScienceData VisualizationNLPLLMsMLOpsStatisticsSeleniumpytestPostmanAPI Testing

About the role

  • OPPO US Research Center is seeking a full-time meticulous and innovative AI/LLM Test Engineer to join our cutting-edge AI team.
  • In this critical role, you will evaluate the performance, reliability, and safety of Large Language Models (LLMs) in real-world product scenarios and test end-to-end generative AI solutions.
  • Your work will directly shape how users experience AI-powered features by ensuring robustness, accuracy, and alignment with product goals.

Responsibilities

  • You will help implement test strategies, execute evaluation workflows, and assist in model performance validation across diverse generative AI use cases.
  • Perform manual and automated test execution on APIs and LLM-integrated user interfaces.
  • Collabration & Documentation Work with QA leads, product managers, and ML engineers to understand test goals and criteria.
  • Report defects, compile evaluation summaries, and maintain testing logs.

Requirements

  • Bachelor's degree in Computer Science, Data Science, Engineering, or a related technical field, or equivalent practical experience.
  • 1+ years of experience in software testing, data science, or ML validation, with exposure to AI/ML systems.
  • Proficiency in Python and testing frameworks (e.g., PyTest, Selenium).
  • Hands-on experience evaluating LLMs in production environments (e.g., GPT, Claude, Llama, Gemini).
  • Familiarity with cloud platforms (GCP, Azure, or AWS) and MLOps tooling (e.g., MLflow, Weights & Biases).
  • Experience with version control (Git) and agile development methodologies.
  • Bachelor's degree or equivalent work experience in a technical field (e.g., Computer Science, Engineering, Data Science).
  • 6+ months experience in software QA, data labeling, LLM evaluation, or ML testing projects.
  • Basic Python proficiency, especially for data processing and automation tasks.
  • Familiarity with LLMs (e.g., GPT, Claude, Gemini) and prompt-based outputs.

Nice to have

  • Expertise in prompt engineering, LLM fine-tuning (e.g., LoRA, RLHF), or optimization techniques.
  • Experience with automated evaluation tools (e.g., LangChain, TruLens) or LLM-specific test suites.
  • Knowledge of data pipelines, SQL/NoSQL databases, and API testing (e.g., Postman).
  • Background in statistics, quantitative analysis, or data visualization for test insights.
  • Contributions to AI safety/ethics initiatives or open-source LLM evaluation projects.
  • Experience testing mobile-integrated AI solutions (Android/iOS).
  • Contractor position
  • Run scripted evaluations to assess outputs for factuality, coherence, and safety.

Skills

  • Conduct end-to-end testing of integrated generative AI solutions, including APIs, data pipelines, and user interfaces.
  • Analyze model failures, edge cases, and adversarial inputs to identify risks and improvement areas.
  • Benchmark LLM performance against industry standards and product-specific KPIs.
  • Advocate for AI ethics and safety through rigorous testing of fairness, bias mitigation, and content moderation.
  • Stay current with advancements in generative AI testing, including red-teaming techniques and evaluation frameworks (e.g., HELM, Dynabench).
  • Propose novel testing strategies for emerging challenges (e.g., hallucinations, context drift).
  • Use existing internal tools or frameworks to automate test runs and result collection.
  • Contribute to prompt generation, input templating, or result tagging processes.

Compensation

  • $100k-$200k

Benefits

  • We are committed to equal employment opportunity regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity or Veteran status.
  • The US base salary range for this full-time position is $100,000-$200,000 + bonus + long term incentives benefits.

Company info

  • We are also seeking a Contractor based LLM Evaluation & QA Engineer to support the testing and validation of large language model (LLM)-powered applications.

Equal opportunity

  • equal opportunity workplace.

Visa & Work Authorization

  • al employment opportunity regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity or Veteran status

This listing is sourced directly from Innopeaktech's careers page and normalized into a canonical job model.