Vacancy description
prophecy technologies
Bombay, भारत
Key Responsibilities :1. Cloud Operations & Incident Management :- Provide hands-on support, maintenance, and troubleshooting for high-availability production environments across AWS and GCP.- Lead incident management workflows, performing root-cause analysis (RCA), resolving critical production outages, and driving post-mortem reviews.- Manage production uptime, capacity planning, performance tuning, and service availability SLAs.2. Configuration Management & Infrastructure as Code (IaC) :- Utilize SaltStack (Mandatory) to enforce configuration standards, automate state management, and orchestrate large-scale node environments.- Provision, update, and manage cloud infrastructure using Terraform following IaC best practices.- Maintain consistency and security postures across multi-cloud infrastructure environments.3. Observability, Automation & Scripting :- Build, maintain, and enhance enterprise monitoring, logging, and observability platforms (e.g., Datadog, Prometheus, Grafana, CloudWatch).- Automate operational tasks and routine maintenance using scripting languages (Python, Bash, or Shell).- Implement automated proactive alerting mechanisms to reduce MTTR (Mean Time to Resolution).4. Technical Governance & Collaboration :- Author comprehensive system documentation, standard operating procedures (SOPs), and architectural runbooks.- Collaborate closely with DevOps and Engineering teams to transition new applications into production https://jobeax.com/link/oQag9rPtOErxPUeK Qualifications :Required Experience & Skills : Experience : 6+ years in Cloud Operations, Cloud Infrastructure, or Site Reliability Engineering (SRE).Multi-Cloud Support : Proven hands-on production support across both AWS and https://jobeax.com/link/OsmuR2Rg2guUdAW7 Management : Expert-level, mandatory proficiency in SaltStack for automated server configuration and https://jobeax.com/link/Ec3uSvyA2weok83b as Code : Hands-on proficiency implementing IaC using https://jobeax.com/link/AN8035IpuHs8VAmc Readiness : Extensive experience in production support, high-severity incident management, and complex system https://jobeax.com/link/fy5oM7jzYH8tDIPp : Practical experience deploying and managing modern monitoring, logging, and metrics https://jobeax.com/link/LiXqOnai3LPCPpvm : Strong automation mindset with proven scripting skills (Python, Bash, etc.).Soft Skills : Excellent written/verbal communication and documentation https://jobeax.com/link/x19nFe4r7kZDCXV3 / Nice-to-Have Skills : - Infrastructure provisioning with AWS CloudFormation.- Hands-on Kubernetes Administration and container orchestration.- Experience with DevOps & CI/CD pipeline development (GitLab CI, GitHub Actions, Jenkins).- Strong Linux Systems Administration foundation (RHEL/CentOS/Ubuntu).- Database Lifecycle Management : Experience handling database upgrades, maintenance, and failovers (Relational & NoSQL).- Proven involvement in Cloud Migration initiatives. (ref:hirist.tech)