Job Description
A leading organization is seeking a Senior Site Reliability Engineer with expertise in cloud-native observability, Kubernetes operations, and modern monitoring technologies. This position focuses on improving system reliability, operational visibility, and platform performance across AWS environments through advanced observability practices, automation, and infrastructure optimization.
The ideal candidate combines strong DevOps and SRE experience with deep technical knowledge of Kubernetes, OpenTelemetry, eBPF, and the Grafana ecosystem.
Responsibilities
- Design, implement, and support enterprise observability solutions across cloud-native platforms.
- Deploy, configure, and optimize Grafana-based monitoring and telemetry environments.
- Implement observability strategies using Grafana Beyla, Grafana Alloy, and OpenTelemetry instrumentation.
- Leverage eBPF technologies to enable low-overhead monitoring, distributed tracing, performance analysis, and infrastructure visibility.
- Administer and optimize Amazon EKS environments to support scalability, availability, and operational reliability.
- Develop monitoring, alerting, visualization, and reporting solutions that provide actionable platform insights.
- Collaborate with engineering teams to improve incident response processes, platform performance, and capacity planning initiatives.
- Drive automation and continuous improvement efforts that enhance operational maturity and system resilience.
- Establish and promote site reliability and DevOps best practices across platform and engineering teams.
Required Experience and Skills
- 5–8 years of experience in Site Reliability Engineering, DevOps Engineering, or related cloud infrastructure roles.
- Strong experience supporting production cloud environments.
- Deep expertise with Kubernetes administration and Amazon EKS.
- Hands-on experience implementing and managing the Grafana observability ecosystem, including Grafana, Grafana Beyla, and Grafana Alloy.
- Strong knowledge of eBPF for application monitoring, infrastructure observability, tracing, and performance optimization.
- Experience implementing OpenTelemetry and modern telemetry frameworks.
- Proficiency with AWS services, Infrastructure as Code methodologies, CI/CD practices, and cloud-native operations.
- Strong troubleshooting and analytical skills focused on reliability, scalability, automation, and performance improvement.
- Experience designing and maintaining monitoring, alerting, and observability solutions in production environments.
- Experience working in healthcare or other regulated industries is preferred.
FAQ
1. What does a Senior Site Reliability Engineer specializing in cloud observability do?
A Senior Site Reliability Engineer (SRE) designs and operates systems that improve the reliability, performance, and availability of cloud-based applications. The role emphasizes observability, automation, incident response, system health monitoring, and reducing operational risks across distributed environments.
2. What does cloud observability involve in an SRE role?
Cloud observability involves collecting and analyzing metrics, logs, traces, and events to understand application and infrastructure behavior. SREs use this visibility to identify performance issues, investigate incidents, detect anomalies, and improve system reliability.
3. Which observability signals are commonly monitored?
The primary signals include metrics, logs, traces, and service events. Engineers correlate these signals to identify root causes of failures, monitor service health, and understand how applications and infrastructure behave under different workloads.
4. What cloud technologies are commonly used in this position?
Depending on the environment, engineers may work with AWS, Azure, or Google Cloud alongside Kubernetes, containers, infrastructure-as-code tools, and monitoring platforms. Observability stacks may include tools such as Prometheus, Grafana, OpenTelemetry, Datadog, or similar platforms.
5. How does this role contribute to system reliability?
The engineer establishes service level objectives (SLOs), monitors service level indicators (SLIs), automates operational tasks, and identifies recurring failure patterns. These practices help teams improve availability, reduce recovery times, and build more resilient cloud systems.
6. What role does incident response play in this position?
Senior SREs lead or support incident investigations, coordinate technical response efforts, identify root causes, and implement corrective actions. Post-incident reviews are used to identify systemic improvements and reduce the likelihood of repeat failures.
7. How does automation support cloud observability?
Automation reduces manual operational work by standardizing monitoring configuration, alerting, remediation, deployment, and infrastructure management. Engineers may use scripting, infrastructure-as-code, and automated response mechanisms to improve consistency and reliability.
8. What challenges are common in cloud observability?
Common challenges include alert fatigue, incomplete telemetry, high monitoring costs, noisy or poorly configured alerts, complex distributed architectures, and difficulty correlating issues across multiple services. Effective observability requires balancing visibility, performance, and operational cost.
Apply for this position
**If you have already submitted your resume for another Job Opening please do not re-apply to a different role. You can email through Contact Us about your interest in other roles.