LawZero

LawZero

Senior ML Data Processing Developer

Montreal · Senior

Sponsorship not specifiedDetected 19 days ago
PythonDockerKubernetesMachine LearningSparkAirflowData EngineeringNLPLLMsAgentic AIAI OrchestrationResearch

About the role

  • Its scientific direction is based on new research and methods proposed by Professor Yoshua Bengio, the most cited AI researcher in the world.
  • Such AI systems could be used to accelerate scientific discovery, to provide oversight for agentic AI systems, and to advance the understanding of AI risks and how to avoid them.
  • LawZero believes that AI should be cultivated as a global public good-developed and used safely towards human flourishing.

Responsibilities

  • Partner with the Research team to define, build, automate, scale, and manage data pipelines that transform raw web-scale data into training datasets for the Scientist AI.
  • Build and maintain data processing pipelines, including deduplication, model-based quality scoring, heuristic filtering, toxicity removal, PII scrubbing, metadata extraction, and proprietary data transformations, with full dataset versioning and provenance tracking, optimizing for throughput and cost at scale.
  • Design and maintain strict leakage detection mechanisms to guard against evaluation contamination across all stages of the data processing pipeline.
  • Build internal tooling and interfaces that let researchers explore, query, and understand available datasets with minimal friction.
  • LawZero is a non-profit organization committed to advancing research and creating technical solutions that enable safe-by-design AI systems.

Requirements

  • Experience with data privacy implementation (PII scrubbing), content-safety filtering (toxicity, bias), and evaluation-contamination prevention.
  • Demonstrated ability to work across Research, Engineering, and/or Legal/Governance teams, translating varied requirements into concrete pipeline work.

Nice to have

  • Experience training, fine-tuning, or deploying ML models for data-quality tasks (classifiers, LLM-based evaluators) and familiarity with LLM inference optimization (e.g. vLLM, SGLang).
  • Familiarity with containerized deployment (Docker, Kubernetes) and infrastructure-as-code practices.
  • Familiarity with ML experiment tracking tools (e.g. Weights and Biases).
  • Experience with data licensing workflows or web-scale data acquisition.
  • Contributions to open-source data processing or NLP tooling.
  • The chance to contribute meaningfully to a globally critical initiative
  • A team of passionate world-class experts in their field
  • A collaborative and inclusive work environment in our vibrant office space in the heart of Little Italy, in the trendy Mile-Ex district, close to public transportation

Skills

  • Degree in computer science, software engineering, or a related field.
  • Hands-on experience with distributed processing frameworks (e.g., Spark, Ray, Flink), designing and optimizing high-throughput pipelines.
  • Strong Python proficiency, including experience writing production-grade data-processing code.
  • Experience with pipeline orchestration frameworks (e.g., Airflow, Prefect, Dagster).
  • Ensure all ingested data meets compliance requirements, internal Data Governance policies, and legal obligations.

Compensation

  • 20 days of vacation per year upon start

Benefits

  • Comprehensive health benefits (including mental health and wellness management account)
  • 20 days of vacation per year upon start
  • Employer contribution of 4% to your retirement savings, with no required employee match
  • Additional compensation totaling 8% of your salary to apply towards additional retirement savings or bonuses (independent of group and individual performance)

Company info

  • We welcome applications from highly qualified individuals interested in working towards our mission in a respectful, inclusive and collaborative setting.

This listing is sourced directly from LawZero's careers page and normalized into a canonical job model.