Lead Site Reliability Engineer (AWS) in Bengaluru, India - Jobeax
Vacancy description
Lead Site Reliability Engineer (AWS) in Bengaluru, India
FIS
Full-timeStandard weekly hours
India, Bengaluru
Lead Site Reliability Engineer (AWS) in Bengaluru, India is listed on Jobeax. Browse 30,000+ vacancies available.
Full time Type Of Hire : Experienced (relevant combo of work and education) Education Desired : We are hiring a Senior Lead Site Reliability Engineer to define, build, and operate always-on, low-latency , and highly secure payment platforms that power large-scale financial transactions. This is a senior technical role, not a pure operations position. You will operate at the intersection of distributed systems engineering, cloud platforms, and reliability architecture, setting technical direction and driving reliability outcomes across mission-critical, regulated systems in Payments and FinTech. You will work across multiple teams and domains, influencing architecture, engineering practices, and operational maturity while remaining hands-on with the most complex reliability challenges. Own and drive reliability outcomes at scale for real-time, distributed payment and transaction processing platforms with strict SLAs, SLOs, and regulatory requirements. Design and evolve enterprise-grade observability platforms (metrics, logs, traces, SLOs/SLIs) that provide actionable insights into system health, customer experience, and business impact. Lead and coordinate response to high-severity production incidents , acting as a technical authority during major events and driving deep root-cause analysis and long-term systemic fixes. Set strategy and drive adoption of SRE best practices including error budgets, capacity modeling, resilience testing, graceful degradation, and operational readiness. Architect automation and self-service platforms that eliminate toil, reduce operational risk, and enable safe, frequent production releases across teams. Partner with senior engineering, product, and platform leaders to influence architectural decisions, cloud migration strategy, disaster recovery posture, and long-term platform evolution. Deep software engineering expertise with a proven track record of building and operating large-scale, distributed, API-driven systems in production. Expertise in observability, alerting, and reliability engineering , using tools such as Prometheus, Grafana, Datadog, Splunk, ELK, or equivalent ecosystems. Strong command of cloud platforms and open systems ( AWS, Azure, or GCP ), including infrastructure-as-code , platform automation, and cloud-native design patterns. Significant experience running mission-critical systems in Payments, FinTech, Banking, or similarly regulated environments , where availability, correctness, and security are non-negotiable. Hands-on experience across Linux (RHEL), Windows , databases (e.g., Oracle RDBMS), and complex enterprise stacks with strong system-level troubleshooting skills. Demonstrated leadership in incident management, post-incident reviews, and continuous reliability improvement , with the ability to influence behavior and standards across teams. Strong automation and scripting skills using Python, Bash, Ansible, or similar tools. Experience building or scaling CI/CD platforms and release automation in high-risk production environments. Prior ownership of reliability strategy or platform initiatives spanning multiple teams or business units. Experience modernizing legacy financial systems into cloud-native or hybrid architectures with a focus on resilience and compliance. Join a culture that values engineering excellence, technical leadership, automation, and continuous learning. For specific information on how FIS protects personal information online, please see the Online Privacy Notice .
... requirements for new telematics API capabilities and enhancements. Develop business cases and requirements to address recurring API issues and improve overall API reliability and customer experience. Lead and coordinate Field Follow Programs for new API endpoints, features, and functionality launches, ensuring successful adoption ...
... container technologies including Docker, Kubernetes Proficiency in Python and Shell scripting to enable automation, platform operations, and continuous improvement Experience with monitoring, observability, Site Reliability Engineering (SRE), networking, governance, and DevSecOps principles within enterprise cloud environments
... neutral infrastructure that enables organizations to safely embrace this new era. This is an opportunity to do career-defining work. As a Senior Site Reliability Engineer you will champion all things pertaining to reliability at Okta for Auth0. Working closely with the Product Engineers, Quality Engineers, Platform Engineers ...
Title: Senior Site Reliability Engineer - I, Product Area Focus Noida (Hybrid) Own availability, the most important product feature, by continually striving for sustained operational excellence of Sumo's planet-scale observability and security products. Work alongside your global SRE team, executing on projects in your ...
About the Role :Site Reliability Engineers are responsible for ensuring the availability, reliability, scalability, and performance of the firms most critical, customer-facing microservices that power all eCommerce channels. This role applies Google-inspired SRE principles to balance feature velocity and system reliability ...
... lifecycle (Shift-Left Security). Research, recommend, and implement best practices for DevSecOps and Kubernetes operations. 5+ years of experience in DevOps, Site Reliability Engineering, or Platform Engineering roles. - 2+ years of hands-on Kubernetes experience, including cluster provisioning, scaling, and troubleshooting. ...
You work as a member of a high-energy, top-performing team of engineers, working alongside technology leaders to shape the vision and drive the execution of ground-breaking compute and data platforms that make a real impact. As a site reliability engineer, you will be responsible for building, maintaining and operating ...
Join us in bringing joy to customer experience. Five9 is a leading provider of cloud contact center software, bringing the power of cloud innovation to customers worldwide. We celebrate diversity and foster an inclusive environment, empowering our employees to be their authentic selves. The Devops Engineer would be an active ...
... comfortable working in distributed teams and enjoy improving systems through thoughtful https://jobeax.com/link/cUAbOSDobkfnp7aS working as a DevOps, Platform, or Site Reliability Engineer Hands‑on experience with public cloud platforms such as AWS, Azure, or GCP Experience with containerization and orchestration technologies ...
We're looking for a seasoned Senior DevOps Engineer to lead cloud-native infrastructure initiatives across AWS, Azure and GCP. You'll architect scalable CI/CD pipelines using GitLab, manage containerized workloads with Kubernetes, Docker and Helm, and drive automation, security, and governance across multi-cloud environments. ...
... in Computer Science, Engineering, or a related field. Minimum of 3 to 4 years of experience as a DevOps or Site Reliability Engineer. Minimum of 5 years of engineering experience. Strong expertise in Jenkins, Terraform, Ansible, Kubernetes, Python, PostgreSQL, MSK, Kafka, RDS, Airflow, and AWS. Comfortable with deployments ...
Site Reliability Engineer (Private Cloud / Virtualization) Cisco is transforming its platforms to run the next generation of cloud-native and multi-cloud services. This role offers a superb opportunity to transform how infrastructure platforms are developed and managed with full software automation. This team is responsible ...
... application issues, and implementing reliability improvements. Contributes to maintaining stable and resilient software services through collaboration with Software Engineering and Site Reliability Engineering teams while ensuring compliance with organizational policies, industry standards, and applicable regulatory requirements. ...
Full time Type Of Hire : Experienced (relevant combo of work and education) Site Reliability Engineer – 4 - 6 Yrs – Pune Location FIS empowers the financial world with payment processing and banking solutions, including software, services and technology outsourcing. FIS’ more than 55,000 worldwide employees are passionate ...
... and infrastructure while always thinking about reliability, scalability, resilience, security, and https://jobeax.com/link/Im0ECCXXkw7iwTBY : - Help build a Site Reliability Engineering culture by sharing best practices, approaches, documentation, and code with other engineering teams.- Apply automation and software to ...
... and infrastructure while always thinking about reliability, scalability, resilience, security, and https://jobeax.com/link/Im0ECCXXkw7iwTBY : - Help build a Site Reliability Engineering culture by sharing best practices, approaches, documentation, and code with other engineering teams.- Apply automation and software to ...
... to identify reliability risks, strengthen service-level performance, and continuously improve operational excellence. Bachelor's degree in Computer Science, Engineering, or a related technical discipline. - 6+ years of experience in Site Reliability Engineering, infrastructure engineering, DevOps, systems engineering, or ...
We are seeking a Site Reliability Engineer (SRE) to support and maintain a 24×7 Azure cloud environment, ensuring high availability, reliability, and performance of infrastructure and hosted services. This role requires the engineer to operate across L1 and L2 support responsibilities, combining proactive monitoring with ...