NeedsO'Need Recruitment Consultancy
Platform Engineer - Monitoring/Incident Management
Pune, भारत
Vacancy description
prophecy technologies
Bombay, भारत
Key Responsibilities :1. Cloud Operations & Incident Management :- Provide hands-on support, maintenance, and troubleshooting for high-availability production environments across AWS and GCP.- Lead incident management workflows, performing root-cause analysis (RCA), resolving critical production outages, and driving post-mortem reviews.- Manage production uptime, capacity planning, performance tuning, and service availability SLAs.2. Configuration Management & Infrastructure as Code (IaC) :- Utilize SaltStack (Mandatory) to enforce configuration standards, automate state management, and orchestrate large-scale node environments.- Provision, update, and manage cloud infrastructure using Terraform following IaC best practices.- Maintain consistency and security postures across multi-cloud infrastructure environments.3. Observability, Automation & Scripting :- Build, maintain, and enhance enterprise monitoring, logging, and observability platforms (e.g., Datadog, Prometheus, Grafana, CloudWatch).- Automate operational tasks and routine maintenance using scripting languages (Python, Bash, or Shell).- Implement automated proactive alerting mechanisms to reduce MTTR (Mean Time to Resolution).4. Technical Governance & Collaboration :- Author comprehensive system documentation, standard operating procedures (SOPs), and architectural runbooks.- Collaborate closely with DevOps and Engineering teams to transition new applications into production https://jobeax.com/link/oQag9rPtOErxPUeK Qualifications :Required Experience & Skills : Experience : 6+ years in Cloud Operations, Cloud Infrastructure, or Site Reliability Engineering (SRE).Multi-Cloud Support : Proven hands-on production support across both AWS and https://jobeax.com/link/OsmuR2Rg2guUdAW7 Management : Expert-level, mandatory proficiency in SaltStack for automated server configuration and https://jobeax.com/link/Ec3uSvyA2weok83b as Code : Hands-on proficiency implementing IaC using https://jobeax.com/link/AN8035IpuHs8VAmc Readiness : Extensive experience in production support, high-severity incident management, and complex system https://jobeax.com/link/fy5oM7jzYH8tDIPp : Practical experience deploying and managing modern monitoring, logging, and metrics https://jobeax.com/link/LiXqOnai3LPCPpvm : Strong automation mindset with proven scripting skills (Python, Bash, etc.).Soft Skills : Excellent written/verbal communication and documentation https://jobeax.com/link/x19nFe4r7kZDCXV3 / Nice-to-Have Skills : - Infrastructure provisioning with AWS CloudFormation.- Hands-on Kubernetes Administration and container orchestration.- Experience with DevOps & CI/CD pipeline development (GitLab CI, GitHub Actions, Jenkins).- Strong Linux Systems Administration foundation (RHEL/CentOS/Ubuntu).- Database Lifecycle Management : Experience handling database upgrades, maintenance, and failovers (Relational & NoSQL).- Proven involvement in Cloud Migration initiatives. (ref:hirist.tech)
Similar vacancies
NeedsO'Need Recruitment Consultancy
Pune, भारत
Ahmedabad, भारत
The Boston Consulting Group
Delhi, भारत
Overture Rede
भारत
Chennai, भारत
Faridabad, भारत
Pune, भारत
Gurgaon, भारत
metro global solution center in
भारत
Amtech Software
$100 USD
Bangalore, भारत