Senior Site Reliability Engineer
Moniepoint, BRM
Job description
About the role
We are seeking an experienced SRE to engineer the reliability of our highly distributed platform. You will combine deep knowledge of distributed systems with strong coding skills to define SLOs, lead incident response, and build automation and self‑healing mechanisms into our systems. You will balance immediate operational stability with long‑term strategic engineering to ensure our services scale linearly with our hyper‑growth.
Key responsibilities
- Participate in on‑call rotations as the primary technical lead and act as Incident Commander during major incidents, initiating war rooms, coordinating cross‑functional teams, and providing clear status updates.
- Instrument code to expose high‑cardinality metrics and distributed traces, and collaboratively define, measure, and defend Service Level Objectives (SLOs) and error budgets with product owners.
- Write high‑quality, production‑ready code (Java, Go, or Python) to build internal tooling, automation platforms, and self‑healing mechanisms that eliminate manual operator intervention.
- Partner with Product Engineering teams during design to ensure new services are built with reliability, scalability, and observability patterns such as circuit breakers, rate limiting, backpressure, and fallback strategies.
- Analyze system performance and traffic patterns to model future capacity needs, conduct load testing and chaos engineering experiments to verify system resilience under failure conditions.
Required profile
- Minimum of 5 years of experience in SRE or Backend Engineering with strong ability to write clean, performant, and tested code in Java, Go, Rust, or Python.
- Deep understanding of distributed systems architecture, microservices fundamentals, and event‑driven architectures required to build systems that scale.
- Extensive experience with Google Cloud Platform (GCP) or similar cloud providers (AWS/Azure) and proficiency in running production workloads on Kubernetes (GKE/EKS) and troubleshooting cluster/infrastructure issues.
- Experience designing observability strategies using OpenTelemetry, Prometheus, New Relic, Datadog, or SigNoz to improve system visibility.
- Familiarity with operating and tuning production data stores (PostgreSQL, MySQL) and streaming platforms (Kafka, RabbitMQ) in high‑throughput environments.
Required skills
- Programming languages: Java, Go, Python, Rust.
- Cloud & container platforms: GCP, AWS, Azure, Kubernetes (GKE, EKS).
- Observability tools: OpenTelemetry, Prometheus, New Relic, Datadog, SigNoz.
- Data stores & streaming: PostgreSQL, MySQL, Kafka, RabbitMQ.
What we offer
- Culture that puts people first, values every opinion, and fosters a supportive environment.
- Learning opportunities with knowledge sharing, training, and regular internal technical talks.
- Competitive compensation including attractive salary, pension, health insurance, annual bonus, and other benefits.
Questions fréquentes
Why are you reporting this job?
Explore further
Salaries, guides and searches in Nigeria.
Salaries by job title
Apply in 30 seconds
Enter your email to apply. An account will be created automatically.
By continuing, you accept our terms of use.
Already have an account? Login
Published 9 hours ago
Expires 1 month from now
4 views · 0 interested
Boost your chances
Upload your CV — we will match you with relevant openings.
Analyzing your CV...
Moniepoint, BRM
Related job offers
-
Chief Technology Officer (CTO)
Seismic Consulting Group Abuja -
Chief Technology Officer (CTO) – Remote, Equity‑Based
DivCex Abuja -
Technical Advisor II – Systems Integration
Catholic Relief Services -
Account Manager – IT Solutions & Junior Pre‑Sales Security Engineer
Click IT Lagos -
iOS Developer (Contract, On-site)
Conclase Ikeja