Crogl, Inc.

Crogl, Inc.

AI Engineer

United States

Sponsorship not specifiedDetected 21 days ago
GitVector DatabasesMachine LearningLLMsRAGAgentic AILangGraphA/B TestingCybersecuritySOC OperationsManual TestingResearchCommunication

About the role

  • This role is ideal for early-career and mid-level engineers who are passionate about AI, enjoy shipping products, and want to work at the forefront of LLMs and agent systems.
  • Understanding whether they are actually improving is equally important.
  • A significant portion of this role involves designing and maintaining evaluation systems for AI agents operating in real-world security workflows.

Responsibilities

  • Build LLM-powered features, workflows, and agentic systems that solve real customer problems.
  • Design and implement evaluation frameworks to measure agent quality, reliability, and business impact.
  • Create automated benchmarks, regression tests, and datasets for evaluating AI behavior.
  • Investigate agent failures and develop systematic approaches to improve performance.
  • Build infrastructure that supports rapid experimentation, evaluation, deployment, and monitoring.
  • Building agents is only half the challenge.
  • You'll build benchmarks, datasets, automated evaluations, regression testing pipelines, and observability systems that help us continuously improve agent performance.
  • Solid software engineering fundamentals, including testing, debugging, and system design.
  • Experience building applications, projects, or products using LLMs and modern AI tools.
  • Ability to design experiments, interpret results, and make data-driven decisions.

Requirements

  • Experience designing evaluations, benchmarks, or testing frameworks for AI systems.
  • Familiarity with OpenAI, Anthropic, Gemini, or open-source LLM ecosystems.
  • Experience with retrieval systems, vector databases, and RAG architectures.
  • Familiarity with LangGraph, OpenAI Agents SDK, MCP, or similar agent frameworks.
  • Experience with observability, tracing, and production monitoring for AI systems.
  • You may have experience as a software engineer, ML engineer, researcher, AI engineer, or founder.
  • What matters most is your ability to learn quickly, work independently, and ship impactful systems.

Skills

  • Is the agent producing accurate investigations?
  • Are changes to prompts, tools, or models actually improving outcomes?
  • Which failure modes matter most?
  • Built and deployed AI agents.
  • Created evaluation frameworks for LLM applications.
  • Published AI projects, demos, or open-source contributions.
  • Developed tools that other people actively use.
  • Strong opinions about what makes AI systems reliable and useful.
  • Built AI products, agents, or developer tools that people actually use.
  • Created evaluation frameworks, benchmarks, or testing systems for LLM applications.
  • Contributed to open-source AI projects.
  • Conducted independent research or published technical writing.

Company info

  • Strong programming skills, preferably in Python.
  • Curiosity, ownership, and a desire to learn quickly.
  • Work closely with customers and internal teams to understand workflows and identify opportunities for AI automation.
  • A core part of this role: evaluating AI systems
  • How do we detect regressions before customers experience them?
  • What makes you stand out from others:

This listing is sourced directly from Crogl, Inc.'s careers page and normalized into a canonical job model.