END

END

Platform Engineer

Washington, GB

Sponsorship not specifiedDetected 41 days ago
GitAzureTerraformCI/CDGitHub ActionsDatadogDevOpsSite Reliability EngineeringPlatform EngineeringGraphQLCybersecurityIncident ResponseStakeholder ManagementERPCadenceFirewallShopify

About the role

  • Service ownership & SLOs: define ownership boundaries, SLIs/SLOs (availability, latency, data freshness), alert thresholds, and regular reporting for platform and integration services.
  • Integration runtime operations: operational ownership of Shopify APIs/webhooks and Patchworks/D365 flows-rate-limit safe patterns, idempotent processing, retries/DLQs, replay tooling, reconciliation, and data correctness checks.
  • D365 environment management: oversee Dev/Sandbox, UAT and Production environments (access, refresh cadence, config hygiene, and release readiness).

Responsibilities

  • CI/CD maintenance & governance: maintain vendor-delivered pipelines for themes/apps/integration services; manage access, configurations, release controls, and quality gates; ensure pipelines remain reliable, secure, and well-documented (not building net-new from scratch).
  • own the change calendar, release readiness checks, environment parity checks, and exception processes during freeze/peak windows.
  • define and maintain dashboards/alerts across storefront experience (RUM/synthetics where available) and integration services; ensure actionable alerting, clear ownership, and on-going tuning.
  • lead triage, stakeholder comms, mitigation, RCA and preventative actions; maintain a stability backlog and drive repeat-incident themes to closure to improve MTTR and reduce recurrence.
  • implement operational controls for key flows (orders, inventory, pricing, fulfilment) including mismatch detection, audit trails, and repeatable replay/backfill procedures.
  • manage D365 version upgrades by testing in lower environments first; coordinate regression validation and go/no-go decisioning.
  • manage user/licence reviews and access governance; monitor and optimise storage/capacity; track and report cost/consumption drivers and optimisation actions.
  • run periodic access reviews across Shopify apps, Patchworks, D365, and monitoring tooling; maintain joiner/mover/leaver processes and audit evidence.
  • maintain runbooks, operational standards, and "how-to-operate" documentation; ensure onboarding for Engineering/Support is clear and current.
  • environment management, release coordination, upgrade rehearsals/testing in lower environments, and go/no-go support.

Requirements

  • Experience with SLO tooling and incident/problem management practices (post-incident reviews, stability backlogs, automation/runbook maturity).
  • Experience implementing synthetic monitoring and RUM for ecommerce journeys, including peak-readiness monitoring uplift.
  • If you're passionate, driven, and believe you can contribute to our future success, we encourage you to apply.

Nice to have

  • understand Shopify-managed protections and operate the process to engage Shopify Plus Support (e.g., bot protection schedules) when required.
  • queues/backpressure, rate-limiting, idempotency, retries/DLQs, replay/backfill, reconciliation, and data integrity monitoring.
  • access reviews (joiner/mover/leaver), least privilege, audit evidence, and secure operational practices across Shopify/Patchworks/D365 and monitoring tools.
  • Experience with D365 licensing and capacity/storage management, plus cost observability and optimisation.
  • Shopify Plus security posture (platform-managed + escalation paths): understand Shopify-managed protections and operate the process to engage Shopify Plus Support (e.g., bot protection schedules) when required.

Skills

  • permissions, secrets, environment configs, release controls/quality gates, and troubleshooting.

Company info

  • Inspired by a shared love of fashion, sneakers and the surrounding culture, END. was founded in Newcastle upon Tyne back in 2005.
  • Our first UK store brought together a curated mix of celebrated designers and emerging brands, many of which were previously hard to find outside of London.
  • From day one, we've nurtured a community of like-minded individuals united by a love for this ever-evolving culture.
  • Today, END. serves over 2 million customers worldwide through a seamless blend of online and physical retail.
  • Our industry-leading stores in Newcastle, Glasgow, Manchester, London, and Milan reflect our commitment to innovation, influence, and inspiration.
  • We offer a carefully curated selection of menswear, womenswear, sneakers, homeware, and lifestyle products to a global audience.
  • At the heart of END. is our customer and it's our people and culture that make the difference.
  • With over 600 employees across our HQ, offices, and retail locations, our customer-first mindset continues to guide everything we do.
  • The Role: Own the reliability, security, performance, and day-to-day operational excellence of our Shopify Plus platform and its critical integration ecosystem (Shopify ↔ Patchworks ↔ D365).
  • You'll lead observability, incident response, stability and problem management, and the operational governance that keeps trading safe and predictable.
  • The role also takes BAU ownership of vendor-delivered CI/CD pipelines-maintaining release controls, access and configuration, and continuous improvement-while operating effectively within Shopify's SaaS constraints (Shopify-managedhosting/CDN/WAF, with clear escalation paths where required).
  • Here's a breakdown of what you'll be doing: CI/CD maintenance & governance: maintain vendor-delivered pipelines for themes/apps/integration services; manage access, configurations, release controls, and quality gates; ensure pipelines remain reliable, secure, and well-documented (not building net-new from scratch).
  • Incident management, stability & problem management: lead triage, stakeholder comms, mitigation, RCA and preventative actions; maintain a stability backlog and drive repeat-incident themes to closure to improve MTTR and reduce recurrence.
  • Data reconciliation & control reporting: implement operational controls for key flows (orders, inventory, pricing, fulfilment) including mismatch detection, audit trails, and repeatable replay/backfill procedures.
  • D365 upgrades & release support: manage D365 version upgrades by testing in lower environments first; coordinate regression validation and go/no-go decisioning.
  • D365 licence, storage & cost observability: manage user/licence reviews and access governance; monitor and optimise storage/capacity; track and report cost/consumption drivers and optimisation actions.
  • Security & compliance controls (operational): secrets management, least privilege, auditability, dependency hygiene, and secure configuration for apps/services; coordinate security reviews and remediation plans.
  • Peak readiness & resilience: trading-event readiness plans (monitoring uplift, incident playbooks, change risk controls/freeze coordination, validation checklists) and business continuity/degraded-mode procedures for integration failures.
  • Vendor/platform escalation: manage Shopify Plus / Patchworks / D365 support escalations with evidence packs (logs, timelines, impact) and track actions to closure.
  • Access governance & audit: run periodic access reviews across Shopify apps, Patchworks, D365, and monitoring tooling; maintain joiner/mover/leaver processes and audit evidence.
  • Platform documentation & guardrails: maintain runbooks, operational standards, and "how-to-operate" documentation; ensure onboarding for Engineering/Support is clear and current.
  • Who we're looking for: Platform/SRE/operations experience supporting high-traffic ecommerce, with strong incident management and operational discipline.
  • Experience maintaining CI/CD pipelines delivered by others (e.g., GitHub Actions/Azure DevOps): permissions, secrets, environment configs, release controls/quality gates, and troubleshooting.
  • Strong observability and reliability engineering experience (New Relic/Sentry/Datadog/etc.): defining SLIs/SLOs, improving alert quality, incident analysis, RCA/problem management, and measurable MTTR/incident reduction.

This listing is sourced directly from END's careers page and normalized into a canonical job model.