Elloe Ai

Elloe Ai

LLM Red Team Intern (Evaluation Systems)

Austin, USA · Intern · Internship

Sponsorship not specified$12k-$16kDetected 376 days ago
Machine LearningNLPLLMsDetection EngineeringLogisticsResearchMentoring

About the role

  • This isn't just prompt tuning - it's forensic risk mapping.
  • You'll be helping define how AI gets deployed responsibly - with traceability, transparency, and real-time protection.

Responsibilities

  • We trace every hallucination, enforce every policy boundary, and create an audit trail for every critical LLM interaction.
  • About the Role You'll red team real-world LLM deployments, design eval harnesses, and help scale Elloe's output-level safety layer.
  • You'll work directly with product and safety leads to uncover failure patterns and codify guardrails for GenAI systems under real-world scrutiny.
  • What You'll Own 1.
  • Evaluation Design Build truthsets and scoring rubrics tied to factuality, policy, or ethical standards Benchmark Elloe's modules across model types (Claude, GPT-4, Gemini, open models) Collaborate with product to refine and expand our eval harnesses 3.
  • You'll red team real-world LLM deployments, design eval harnesses, and help scale Elloe's output-level safety layer.

Compensation

  • $12k-$16k

Benefits

  • Research stipend
  • Research stipend Location: Remote-first; flexible for global candidates To Apply: Share a jailbreak or eval idea you'd love to run against GPT-4.

This listing is sourced directly from Elloe Ai's careers page and normalized into a canonical job model.