Powered by pgvector · cosine kNN
iterate
We're hiring a Site Reliability Engineer with a strong observability focus to help keep our clients stores, warehouses, and digital channels running when it matters most. You'll own our Prometheus and Grafana stack, t…
Your match
See how you fit
Scored against this job in seconds
Your account
Sign in to apply
Your profile and your match for this job appear right here.
sign in above to apply · via LinkedIn
About the role
We're hiring a Site Reliability Engineer with a strong observability focus to help keep our clients stores, warehouses, and digital channels running when it matters most. You'll own our Prometheus and Grafana stack, turning telemetry from checkout, payments, stock, and store systems into the dashboards, alerts, and SLOs that keep the business trading especially during Black Friday, Christmas, and other peak moments.
What You'll Do:
Build and maintain our Prometheus/Grafana observability stack across Kubernetes, cloud, and our store and warehouse systems.Define golden signals, SLOs, and error budgets for checkout, payments, stock, and store connectivity.Build low-noise, actionable alerting and drive down MTTD/MTTR.Extend tracing with OpenTelemetry across the full customer journey, from click to payment to stock update.Lead peak trading readiness; capacity planning, load testing, and live trading dashboards for Black Friday, Christmas, and flash sales.Run blameless postmortems and turn every incident into a lasting fix.Automate remediation and routine platform maintenance wherever possible.
What You'll Bring:
Hands-on Prometheus and Grafana experience in production, ideally at scale.Kubernetes and cloud experience (Azure, AWS, or GCP).Working knowledge of OpenTelemetry, TSDBs, and open metrics.Comfort with Terraform, CI/CD, Git, and scripting (Bash, Python, or PowerShell).Understanding of DORA metrics and experience using them to drive improvement.Retail, e-commerce, or other seasonally-spiky industry experience is a plus.
The CultureThey treat incidents as data, not blame. Reliability is everyone's job, not just the SRE team's. They build tools people want to use, not tools they're forced to use. On-call is shared and sustainable, and we plan hard in the quiet months so peak trading stays calm.
Sound Like You?
We'd love to hear from you. Apply with your CV and a short note on an observability problem you're proud of solving.
sign in above to apply · via LinkedIn
CMC Markets
If you also believe that everyone should be able to achieve their financial potential, then you’ll love contributing to CMC Markets’ company vision of providing the ultimate trading experience. Seize the opportunity t…
CMC Markets ANZ
If you also believe that everyone should be able to achieve their financial potential, then you’ll love contributing to CMC Markets’ company vision of providing the ultimate trading experience. Seize the opportunity t…
Salient Group
Observability & Reliability Engineer | Melbourne 💰 $150k–$170k + Super + Bonus📍 Melbourne | Hybrid When you’re building a platform that sits behind high-stakes decisions, knowing what to work on and when is importan…
CMC Markets
If you also believe that everyone should be able to achieve their financial potential, then you’ll love contributing to CMC Markets’ company vision of providing the ultimate trading experience. Seize the opportunity t…
Cover Genius
About the Company Cover Genius is a Series E Insurtech that protects the global customers of the world’s largest digital companies including Booking Holdings, owner of Priceline, Kayak and Booking.com, Intuit, Hopper,…
Tyro Payments
Why Tyro? At Tyro, we’re into business big time. Through our integrated payments, banking and lending solutions, we’re here to ensure nothing stands in the way of Australian business success. With over 21 years' exper…
Your job hunt, handled
Ask about any role and get a straight answer on your fit. Then stop searching: new matches land in your WhatsApp the moment they’re listed.
Free for jobseekers