Team Lead, Site Reliability Engineering
Moniepoint, BRM · Lagos
Job description
About the role
We are seeking a Team Lead, SRE to guide a squad of site reliability engineers responsible for the reliability of our highly distributed financial platform. You will design high‑level reliability architecture, mentor engineers, define the technical roadmap, and drive the culture of Site Reliability Engineering. The role balances strategic leadership with deep technical work to ensure our systems and people can scale linearly with hyper‑growth.
Key responsibilities
- Set the technical direction for the SRE team, architect self‑healing systems, define reliability standards (Production Readiness Reviews), and promote observability‑as‑code and automation best practices.
- Define and enforce end‑to‑end system visibility standards, guide deep instrumentation (logging, tracing, metrics), and govern the monitoring ecosystem to keep alerts actionable while minimizing noise.
- Lead, mentor, and grow a team of Senior and Associate SREs; conduct code reviews, facilitate technical workshops, and foster a culture of engineering excellence.
- Act as the ultimate escalation point for major incidents, refine incident‑management processes, ensure efficient RCAs, and partner with Engineering Managers and Product Leads to define business‑aligned SLOs.
Required profile
- Minimum 6 years of experience in SRE or Backend Engineering, with at least 2 years in a Lead or Senior/Staff role mentoring others.
- Expert‑level proficiency in Java, Go, Rust, or Python, setting the standard for code quality.
- Mastery of distributed systems patterns, capable of designing scalable architectures and debugging complex micro‑service interactions.
- Deep expertise with Google Cloud Platform (GCP) or AWS, extensive experience running Kubernetes (GKE) at scale and troubleshooting infrastructure issues.
- Proven experience defining observability strategies for large teams, architecting the full telemetry stack from custom instrumentation to monitoring and actionable alerting.
- Strong communication skills with the ability to de‑escalate high‑pressure war rooms with calm authority.
Required skills
- Java
- Go
- Rust
- Python
- Google Cloud Platform (GCP)
- AWS
- Kubernetes (GKE)
- Distributed systems design
- Observability and telemetry
- Logging, tracing, metrics
- Monitoring and alerting
What we offer
- Culture – people‑first environment that values well‑being, inclusivity, and mutual respect.
- Learning – strong focus on development with knowledge sharing, training, and regular internal technical talks.
- Compensation – attractive salary, pension, health insurance, annual bonus, and additional benefits.
Questions fréquentes
Why are you reporting this job?
Explore further
Salaries, guides and searches in Nigeria.
Salaries by job title
Apply in 30 seconds
Enter your email to apply. An account will be created automatically.
By continuing, you accept our terms of use.
Already have an account? Login
Published 1 day ago
Expires 1 month from now
3 views · 0 interested
Boost your chances
Upload your CV — we will match you with relevant openings.
Analyzing your CV...
Moniepoint, BRM
Lagos