Senior Systems Engineer - AWS in India, India is listed on Jobeax. Browse 30,000+ vacancies available.
Senior DevOps Engineer
As a Senior DevOps Engineer - Observability Platform , you will be responsible for building and maintaining scalable, reliable infrastructure and deployment pipelines with a strong emphasis on observability - metrics, logs, and traces - across systems running on Kubernetes and Azure. You will work closely with development teams to improve development velocity while ensuring system reliability, security, and performance. This role is critical in providing a standardized, observability platform that gives both internal engineering teams and external, customer-facing services deep, reliable visibility into system health, performance, and reliability
3+ years of academic or work experience with Programming Language such as C, C++, Java, Python, etc.
- Infrastructure Management: Design, implement, and maintain cloud-based infrastructure using Infrastructure as Code principles
- Automation: Develop automation scripts and tools to streamline operations and eliminate manual processes
- Manage containerization strategies and orchestration using Docker and Kubernetes
- Observability Platform: Design, build, and operate a standardized, self-service metrics, logs, and tracing platform (Grafana, Prometheus, Loki, OpenTelemetry) serving both internal teams and external, customer-facing services running on Kubernetes and Azure
- Partner with engineering teams to instrument applications and infrastructure, standardizing telemetry collection with OpenTelemetry
- Performance Optimization: Use observability data to analyze and optimize system performance, scalability, and cost-efficiency
- Create and maintain thorough documentation for infrastructure, deployment processes, and operational procedures
- Incident Response & Escalation: Provide second-tier engineering escalation during business hours and own the telemetry, SLO, and alerting tooling that powers incident detection and reduces MTTD/MTTR front-line 24/7 on-call is owned by the dedicated SRE team, not observability engineers. 5+ years of experience in DevOps, Observability, or similar roles, including hands-on production experience operating Kubernetes based stack.
- Cloud Platforms: Extensive hands-on experience with Azure, including its observability services
Infrastructure as Code: Proficiency with Terraform, AWS CloudFormation, or similar IaC tools
Advanced knowledge of Docker and Kubernetes ecosystem
Hands-on experience with Grafana, Prometheus, Loki, Tempo or Jaeger, OpenTelemetry, and Alertmanager experience scaling metrics storage with Thanos, Mimir, or Cortex
Programming/Scripting: Strong coding skills in Python, Bash, or Go