Vacancy description
India, Chennai
Manager - Infrastructure & Site Reliability Engineering in Chennai, India is listed on Jobeax. Browse 30,000+ vacancies available.
Role : Manager Infrastructure & Site Reliability EngineeringExperience : 10 12 yrsLocation : Chennai / BangaloreRole/Responsibilities : - Lead the SRE function enforce SLOs/SLAs, error budget accountability, incident management, and post-mortem culture, with focus on availability 9s and metrics like MTTR, MTTD, and change failure rate.- Own observability, telemetry, tracking, and reporting including instrumentation, alerting logic, and custom dashboards across infrastructure and services.- Drive engineering-led reliability practices : build and maintain self-healing systems, automated runbooks, capacity models, and performance profiling frameworks to reduce toil and improve system resilience.- Serve as a hands-on technical contributor IaC, CI/CD pipelines, platform tooling, and active participation in reviews and critical incident response.- Manage the full infrastructure team scope : SRE, patching, hardware lifecycle, and facility infrastructure.- Handle compliance across the board audit readiness, access controls, and vulnerability management.- Hire, develop, and manage the team performance management, career growth, workload planning, and shift management.- Communicate infrastructure health, risk, and investment needs to stakeholders; apply AI tooling selectively to improve operational https://jobeax.com/link/GK6JLJWOsFhSBApv Skills and Experience : - Up to 10 years of experience in infrastructure and/or SRE roles, with 3+ years in a team lead or management capacity.- Hands-on cloud platform experience Azure and AWS including networking, IAM, compute, and storage.- Infrastructure as Code proficiency : Terraform, Pulumi, or CloudFormation with version-controlled, testable infra pipelines.- SRE fundamentals : SLO/SLA design, error budgets, availability 9s, and key reliability metrics (MTTR, MTTD, change failure rate); blameless post-mortem process.- Observability stack experience Datadog, Prometheus, Grafana, or equivalent; familiarity with instrumentation standards like OpenTelemetry.- Hands-on experience building self-healing systems, automated runbooks, and capacity modeling and performance profiling frameworks.- CI/CD and GitOps pipeline experience.- Experience managing shift-based operations teams.- Strong stakeholder communication translating infrastructure risk and investment needs for non-technical audiences.- Experience with security and compliance requirements in enterprise environments (SOC 2, PCI-DSS, or equivalent).Nice to Have Qualities & Skills : - Experience with FinOps cloud cost visibility, rightsizing, and chargeback models.- Background in platform engineering or developer experience (internal developer portals, self-service infra).- Prior experience in a regulated industry financial services, real estate tech, or similar.- Exposure to multi-cloud or hybrid cloud environments. (ref:hirist.tech)