TechTorch

TechTorch

AI-Enabled Data Engineer

United States

Sponsorship not specifiedDetected 56 days ago
PythonSQLSnowflakeRedshiftDatabricksVector DatabasesAWSAzureCI/CDDevOpsKafkaMachine LearningSparkAirflowdbtData EngineeringLLMsRAGAgentic AIIncident ResponseLeadership

About the role

  • We combine the agility of a scale-up with the discipline and rigor demanded by the most sophisticated investors and operators.
  • We work across the full data stack - from ingestion and modeling to AI-ready data products - and we move fast by letting AI do the heavy lifting wherever it can.
  • This role sits at the intersection of deep data engineering craft and modern AI capability.

Responsibilities

  • Design, build, and maintain scalable data pipelines and ETL/ELT workflows across cloud and on-prem environments.
  • Model data with dbt: write modular SQL transformations, manage dependencies, enforce data contracts, and maintain documentation.
  • Build and maintain semantic layers that serve consistent, governed metrics to downstream consumers.
  • Design data warehouse schemas and data lake structures that balance performance, cost, and queryability.
  • Implement data quality frameworks - testing, validation, alerting, and lineage - as first-class citizens in every pipeline.
  • Own the reliability of data products end-to-end - monitoring, alerting, incident response, and root cause analysis.
  • Work across AWS and Azure cloud services (S3, Glue, ADLS, ADF, Synapse, Redshift) to design cost-effective, scalable architectures.
  • Build data pipelines that feed AI systems - including RAG ingestion workflows, vector store loading, document chunking, and embedding pipelines.
  • TechTorch is a high-growth enterprise technology consultancy that partners with the world's leading private equity-backed businesses.
  • We deliver AI-powered solutions, accelerators, and data-driven transformation initiatives that drive measurable value at speed and scale.

Requirements

  • Core Data Engineering: ETL/ELT Design · Data Modeling · Data Quality & Testing · Data Lineage · Batch & Incremental Loads
  • Data Platforms: Snowflake · Databricks · Apache Spark / PySpark · Delta Lake · Data Warehouses · Data Lakes
  • Transformation & Modeling: dbt Core / dbt Cloud · SQL (advanced) · Semantic Layer · Dimensional Modeling
  • Orchestration: Apache Airflow · Dagster / Prefect · Azure Data Factory · Databricks Workflows
  • AI-Enabled Engineering: RAG & Vector Store Pipelines · AI-Augmented ETL · MCP / Agent Data Tools · AI-Paired Programming · LLM Integration in Pipelines
  • Cloud & DevOps: AWS (S3, Glue, Redshift) · Azure (ADLS, ADF, Synapse) · CI/CD for Data · Infrastructure as Code · Python
  • Nice to Have
  • Building or contributing to internal data tooling, frameworks, or accelerators.

Nice to have

  • Experience with streaming architectures: Kafka, Spark Streaming, or Flink.
  • Exposure to feature stores (Feast, Tecton) or ML platform data pipelines.
  • Hands-on with vector databases: Pinecone, Weaviate, Qdrant, or pgvector.
  • Familiarity with data mesh or data product ownership models.
  • Experience with Snowpark or Databricks AI/BI tooling.

Skills

  • Orchestrate workflows across Airflow, Dagster/Prefect, Azure Data Factory, and Databricks Workflows - choosing the right tool for each job.
  • dbt Core / dbt Cloud · SQL (advanced) · Semantic Layer · Dimensional Modeling
  • RAG & Vector Store Pipelines · AI-Augmented ETL · MCP / Agent Data Tools · AI-Paired Programming · LLM Integration in Pipelines
  • AWS (S3, Glue, Redshift) · Azure (ADLS, ADF, Synapse) · CI/CD for Data · Infrastructure as Code · Python

Company info

  • Our mission is to redefine enterprise technology consulting for private equity.
  • TechTorch was founded by seasoned leaders - including former Bain consultants, CIOs, and tech executives - with deep expertise in technology, transformation, and value creation.

This listing is sourced directly from TechTorch's careers page and normalized into a canonical job model.