Site Reliability Engineer/DevOps in Thiruvananthapuram, India is listed on Jobeax. Browse 30,000+ vacancies available.
Zafin is an AI platform company helping regulated institutions modernize how critical work is designed, governed, and delivered. Our technology enables organizations to move faster while maintaining the governance, accountability, and control required in highly regulated environments.
Our portfolio includes Zafin AIOS , an agent orchestration platform for governed AI work; the Zafin Banking Platform , which helps banks modernize product, pricing, offers, billing, loyalty, and relationship management; and Zafin IO , an integration platform that connects data, systems, and workflows across complex enterprise environments.
Headquartered in Toronto, Canada, Zafin partners with leading financial institutions across North America, Europe, the Middle East, Africa, and Asia-Pacific. As AI transforms the future of financial services, we're building the platforms that help regulated organizations adopt AI responsibly and at scale.
Cloud Site Reliability Engineer I (CSRE I)
Zafin is seeking a Cloud Site Reliability Engineer I (CSRE I) to lead strategic initiatives in ensuring the reliability, scalability, and performance of our cloud infrastructure and applications. This advanced role requires mastery in cloud technologies, strategic planning, and incident management to drive innovative solutions and operational excellence.
Manage the resolution of complex technical issues involving Zafin's products and Azure cloud environment.
Conduct in-depth Root Cause Analysis (RCA) for high-severity incidents and drive initiatives to reduce error recurrence.
Represent the organization in external client escalation calls, providing expert guidance and solutions.
Optimize cloud infrastructure for high performance, scalability, and cost-effectiveness.
Oversee the implementation of advanced monitoring solutions and integrate predictive analytics for proactive issue resolution.
Create and maintain comprehensive documentation of cloud architectures, processes, and incident management strategies.
Bachelor's degree in computer science, Engineering, or a related field (Master's degree preferred).
- 8+ years of experience in cloud support, operations, or a related role.
- Advanced expertise in Microsoft Azure (preferred) or equivalent cloud platforms.
- Demonstrated experience in designing and scaling container orchestration systems like AKS or OpenShift.
- Proven leadership in managing automated deployment pipelines, including Azure DevOps.
- Mastery in enterprise monitoring platforms (e.g., Advanced scripting skills with PowerShell, Python, or similar languages.
- Extensive experience in incident management and defining SLAs for global production environments.
- In-depth knowledge of database management, particularly Postgres.
Advanced certifications in cloud platforms (e.g., Azure Solutions Architect Expert).
Experience with ITSM tools and processes (e.g., ServiceNow).
Comprehensive understanding of security and compliance in cloud environments.
Visionary approach to operational innovation and strategic planning.
Joining our team means being part of a culture that values diversity, teamwork, and high-quality work. We offer competitive salaries, annual bonus potential, generous paid time off, paid volunteering days, wellness benefits, and robust opportunities for professional growth and career advancement. Zafin welcomes and encourages applications from people with disabilities. Accommodations are available on request for candidates taking part in all aspects of the selection process.
The methods by which Zafin contains uses, stores, handles, retains, or discloses applicant information can be accessed by reviewing Zafin's privacy policy at By submitting a job application, you confirm that you agree to the processing of your personal data by Zafin described in the candidate privacy notice.