TechTorch
AI-Enabled Data Engineer
United States
Sponsorship not specifiedDetected 56 days ago
PythonSQLSnowflakeRedshiftDatabricksVector DatabasesAWSAzureCI/CDDevOpsKafkaMachine LearningSparkAirflowdbtData EngineeringLLMsRAGAgentic AIIncident ResponseLeadership
About the role
- We combine the agility of a scale-up with the discipline and rigor demanded by the most sophisticated investors and operators.
- We work across the full data stack - from ingestion and modeling to AI-ready data products - and we move fast by letting AI do the heavy lifting wherever it can.
- This role sits at the intersection of deep data engineering craft and modern AI capability.
Responsibilities
- Design, build, and maintain scalable data pipelines and ETL/ELT workflows across cloud and on-prem environments.
- Model data with dbt: write modular SQL transformations, manage dependencies, enforce data contracts, and maintain documentation.
- Build and maintain semantic layers that serve consistent, governed metrics to downstream consumers.
- Design data warehouse schemas and data lake structures that balance performance, cost, and queryability.
- Implement data quality frameworks - testing, validation, alerting, and lineage - as first-class citizens in every pipeline.
- Own the reliability of data products end-to-end - monitoring, alerting, incident response, and root cause analysis.
- Work across AWS and Azure cloud services (S3, Glue, ADLS, ADF, Synapse, Redshift) to design cost-effective, scalable architectures.
- Build data pipelines that feed AI systems - including RAG ingestion workflows, vector store loading, document chunking, and embedding pipelines.
- TechTorch is a high-growth enterprise technology consultancy that partners with the world's leading private equity-backed businesses.
- We deliver AI-powered solutions, accelerators, and data-driven transformation initiatives that drive measurable value at speed and scale.
Requirements
- Core Data Engineering: ETL/ELT Design · Data Modeling · Data Quality & Testing · Data Lineage · Batch & Incremental Loads
- Data Platforms: Snowflake · Databricks · Apache Spark / PySpark · Delta Lake · Data Warehouses · Data Lakes
- Transformation & Modeling: dbt Core / dbt Cloud · SQL (advanced) · Semantic Layer · Dimensional Modeling
- Orchestration: Apache Airflow · Dagster / Prefect · Azure Data Factory · Databricks Workflows
- AI-Enabled Engineering: RAG & Vector Store Pipelines · AI-Augmented ETL · MCP / Agent Data Tools · AI-Paired Programming · LLM Integration in Pipelines
- Cloud & DevOps: AWS (S3, Glue, Redshift) · Azure (ADLS, ADF, Synapse) · CI/CD for Data · Infrastructure as Code · Python
- Nice to Have
- Building or contributing to internal data tooling, frameworks, or accelerators.
Nice to have
- Experience with streaming architectures: Kafka, Spark Streaming, or Flink.
- Exposure to feature stores (Feast, Tecton) or ML platform data pipelines.
- Hands-on with vector databases: Pinecone, Weaviate, Qdrant, or pgvector.
- Familiarity with data mesh or data product ownership models.
- Experience with Snowpark or Databricks AI/BI tooling.
Skills
- Orchestrate workflows across Airflow, Dagster/Prefect, Azure Data Factory, and Databricks Workflows - choosing the right tool for each job.
- dbt Core / dbt Cloud · SQL (advanced) · Semantic Layer · Dimensional Modeling
- RAG & Vector Store Pipelines · AI-Augmented ETL · MCP / Agent Data Tools · AI-Paired Programming · LLM Integration in Pipelines
- AWS (S3, Glue, Redshift) · Azure (ADLS, ADF, Synapse) · CI/CD for Data · Infrastructure as Code · Python
Company info
- Our mission is to redefine enterprise technology consulting for private equity.
- TechTorch was founded by seasoned leaders - including former Bain consultants, CIOs, and tech executives - with deep expertise in technology, transformation, and value creation.
Apply directly at TechTorch →Create a free account for alerts like thisView TechTorch immigration profile
This listing is sourced directly from TechTorch's careers page and normalized into a canonical job model.