MeridianLink

MeridianLink

Data Scientist, AI Data Foundations

US Remote

Sponsorship not specifiedDetected 77 days ago
PythonData StructuresAlgorithmsSQLDatabricksVector DatabasesAWSAzureMachine Learningscikit-learnPandasNumPyData EngineeringData ScienceNLPLLMsRAGStatisticsUnityCommunicationPublic Speaking

About the role

  • Reporting into the Data Engineering organization, the Data Scientist is responsible for designing and building the curated data structures that AI and ML applications consume across MeridianLink.
  • You will own the vector stores behind our RAG systems, the feature store that powers model training and inference, and the graph databases that capture relationships across applicants, products, and decisions.
  • You will also lead targeted data discovery work, surfacing hidden trends in our lending and account-opening data that inform both AI use cases and the broader business.

Responsibilities

  • Build and maintain vector stores for RAG: Design embedding pipelines, chunking strategies, indexing approaches, and refresh patterns for the vector stores powering retrieval-augmented generation across MeridianLink products.
  • Own the feature store: Design, build, and operate feature store assets used for model training and online/offline inference, including feature definitions, freshness SLAs, lineage, point-in-time correctness, and reuse across teams.
  • Design graph data structures: Build graph databases that model relationships between applicants, applications, products, lenders, decisions, and outcomes - and make them queryable for both AI use cases and analytical investigations.
  • Lead data discovery: Profile our lending, deposit, and behavioral datasets to identify hidden trends, segments, anomalies, and potential model drivers
  • Engineer for AI consumption: Build the curated, AI-ready datasets that downstream model builders, application engineers, and analysts rely on - with appropriate quality, documentation, and governance baked in.
  • Partner with model builders: Work closely with ML engineers and applied scientists to make sure the data structures you build accelerate their work rather than slow it down.
  • Champion responsible data use: Partner with governance, security, and compliance to ensure that AI-facing data assets respect data classification, customer consent, and regulatory boundaries from day one.
  • Lead data discovery: Profile our lending, deposit, and behavioral datasets to identify hidden trends, segments, anomalies, and potential model drivers; turn findings into actionable hypotheses for product, risk, and growth teams.
  • Design embedding pipelines, chunking strategies, indexing approaches, and refresh patterns for the vector stores powering retrieval-augmented generation across MeridianLink products.

Requirements

  • 4-7 years of experience in a data science, ML engineering, or applied data role, with a meaningful portion of that time spent building data assets that other people's models or applications consumed.
  • Strong proficiency in Python (pandas, NumPy, scikit-learn, PySpark) and SQL
  • Practical experience with embedding models and LLM tooling (e.g., Hugging Face transformers, OpenAI / Azure OpenAI APIs, LangChain or similar) in a production or near-production context.

Nice to have

  • Familiarity with open-source vector databases such as pgvector, Pinecone, Weaviate, Chroma, or FAISS, and a clear point of view on when to use which.
  • Bachelor's or Master's degree in Computer Science, Statistics, Mathematics, Engineering, or a related quantitative field, or equivalent professional experience.

Skills

  • Strong written and verbal communication skills; able to write up findings for both technical and business audiences.
  • profiling messy real-world datasets, surfacing non-obvious patterns, validating findings statistically, and explaining them clearly.
  • Experience working in a SaaS or FinTech environment, particularly with lending, deposit, credit, fraud, or KYC/AML data.
  • Experience with Databricks-native AI/ML tooling: Databricks Vector Search, Databricks Feature Store, MLflow, and Unity Catalog.
  • Experience with Microsoft Azure data and AI services (Azure OpenAI, Azure AI Search, ADLS Gen2).
  • Experience evaluating RAG systems end-to-end (recall@k, faithfulness, answer quality, hallucination measurement).
  • Exposure to graph algorithms (community detection, link prediction, centrality) applied to real business problems.
  • Our Data & AI Stack
  • Lakehouse: Azure Databricks, Delta Lake, Unity Catalog, PySpark, SQL
  • AI Data Foundations: Databricks Vector Search, Databricks Feature Store, MLflow

This listing is sourced directly from MeridianLink's careers page and normalized into a canonical job model.