Elloe Ai
LLM Red Team Intern (Evaluation Systems)
Austin, USA · Intern · Internship
Sponsorship not specified$12k-$16kDetected 376 days ago
Machine LearningNLPLLMsDetection EngineeringLogisticsResearchMentoring
About the role
- This isn't just prompt tuning - it's forensic risk mapping.
- You'll be helping define how AI gets deployed responsibly - with traceability, transparency, and real-time protection.
Responsibilities
- We trace every hallucination, enforce every policy boundary, and create an audit trail for every critical LLM interaction.
- About the Role You'll red team real-world LLM deployments, design eval harnesses, and help scale Elloe's output-level safety layer.
- You'll work directly with product and safety leads to uncover failure patterns and codify guardrails for GenAI systems under real-world scrutiny.
- What You'll Own 1.
- Evaluation Design Build truthsets and scoring rubrics tied to factuality, policy, or ethical standards Benchmark Elloe's modules across model types (Claude, GPT-4, Gemini, open models) Collaborate with product to refine and expand our eval harnesses 3.
- You'll red team real-world LLM deployments, design eval harnesses, and help scale Elloe's output-level safety layer.
Compensation
- $12k-$16k
Benefits
- Research stipend
- Research stipend Location: Remote-first; flexible for global candidates To Apply: Share a jailbreak or eval idea you'd love to run against GPT-4.
Apply directly at Elloe Ai →Create a free account for alerts like thisView Elloe Ai immigration profile
This listing is sourced directly from Elloe Ai's careers page and normalized into a canonical job model.