Sayari
Data Engineer (Remote, US)
Remote - US
Sponsorship not specified$90k-$120kDetected 11 days ago
PythonScalaDjangoCode ReviewGitPostgreSQLBigQueryMachine LearningSparkAirflowData EngineeringLLMsAI OrchestrationRoadmappingRecruitingCommunication
About the role
- About Sayari: Sayari is the judgment infrastructure for trustworthy AI in economic security and commercial risk.
- The Sayari Commercial World Model resolves 11.7B+ primary-source records from 250+ jurisdictions forming the ground truth of global commerce.
- Trusted by U.S. Customs and Border Protection, HM Revenue & Customs, and Fortune 500 enterprises, Sayari is used by thousands of professionals across 35+ countries to secure supply chains and dismantle illicit networks.
Responsibilities
- Design, build, and maintain scalable data pipelines using Python, Spark, and Airflow to support our core data acquisition and entity resolution engines.
- Collaborate cross-functionally with AI/ML and Product teams to implement new features and AI-native products.
- Proactively identify and resolve bottlenecks in our complex ETL processes, bringing a fresh perspective to refine and optimize our existing codebase.
- Contribute to a robust engineering culture through rigorous code reviews, unit testing, and clear communication of design decisions.
- Own the end-to-end delivery of roadmap tasks within two-week sprints, ensuring work meets high standards for quality, documentation, and performance.
- Participate in roadmap planning and story refinement, eventually taking ownership of major epics that drive our long-term product defensibility.
- A Judgment Ontology, encoding over a decade of investigative tradecraft, and Superconductor, an agentic orchestration platform, deliver AI that reasons like an expert analyst, shows its work, and traces every finding to its source.
Skills
- Professional proficiency in Python and experience contributing to shared codebases using Git (branching, PRs, code reviews).
- 3+ years of experience working in Data Engineering
- Demonstrated experience working with relational databases (PostgreSQL/BigQuery) and an interest in or familiarity with graph databases.
- Familiarity with distributed computing (Spark) or a strong desire to master it.
- Strong collaborative skills and the ability to work effectively in an Agile, sprint-based environment.
- A "self-directed" orientation: ability to move tasks from "assigned" to "complete" with high autonomy and clear communication.
- Experience with Django, Scala, or Scrapy.
- Hands-on experience with workflow orchestration tools like Airflow.
- Experience or strong interest in LLM tuning, deployment, and AI engineering best practices.
- Experience working with international or non-English datasets.
- Prior experience working with high-scale, complex data pipelines.
Compensation
- $90,000 - $120,000 USD
Benefits
- 100% fully paid medical, vision, and dental for employees and their dependents
- Generous time off
- we observe all US federal holidays, close our office for a winter break (12/24-12/31), in addition to granting 18 PTO days and 10 sick days
- A strong commitment to diversity, equity, and inclusion
- Eligibility to participate in additional benefits such as 401k match up to 5%, 100% paid life insurance (up to $100,000 coverage),, and parental leave
- Limitless growth and learning opportunities
Equal opportunity
- equal opportunity employer and strongly encourages diverse candidates to apply.
- We believe diversity and inclusion mean our team members should reflect the diversity of the United States.
This listing is sourced directly from Sayari's careers page and normalized into a canonical job model.