Principal Site Reliability Engineer - DevOps in United… - Jobeax
Vacancy description
Principal Site Reliability Engineer - DevOps in United States of America, India
Mig Staffing
India
Principal Site Reliability Engineer - DevOps in United States of America, India is listed on Jobeax. Browse 30,000+ vacancies available.
Role Overview :We're looking for a Principal Site Reliability Engineer to join our Site Reliability & Infrastructure Engineering team. We run entirely on the cloud, and this team owns production intelligence, reliability, and operational automation - moving from reactive incident response toward systems that observe, correlate, reason, and act. We're building the next generation of SRE: not dashboards and runbooks, but agents and automation that make our platform self-aware and self-healing at https://jobeax.com/link/ThFBJVlsRbz9m53j Principal level, you won't just operate the platform - you'll define how we think about reliability across the entire engineering organization. You'll set the technical direction, establish the standards, and be the person other engineers turn to when the hardest infrastructure problems need to be https://jobeax.com/link/7rSp0MQPRkGNfSEk You'll Do :- Define the technical vision and long-term architecture for ServiceTitan's reliability and infrastructure platform - setting the standard for how we operate at scale.- Lead the design and delivery of major platform initiatives: Kubernetes infrastructure, observability frameworks, SLO programs, and incident management systems.- Partner with Engineering Managers, Staff Engineers, and product engineering leadership to review architecture and infrastructure decisions before they ship - and hold the bar on non-functional requirements across the org.- Identify systemic reliability risks before they become incidents, and drive remediation across teams with org-wide impact.- Define and operationalize SLIs, SLOs, and error budgets across ServiceTitan's platform - not just within the SRE team, but as a standard every engineering team adopts.- Drive adoption of reliability and observability best practices across engineering - through documentation, design reviews, and direct partnership with product teams.- Design and build AI-assisted operational systems - agents that correlate production telemetry, diagnose failure patterns, recommend or execute remediation, and participate safely in deployment decisions.- Partner with Infrastructure Engineering on deployment safety - progressive delivery, canary analysis, and AI-assisted promotion decisions.- Mentor Staff and Senior SREs, raising the technical ceiling of the team through code reviews, architecture feedback, and direct pairing on complex problems.- Contribute to the technical hiring bar - participate in system design interviews and help define what Principal-level looks like at https://jobeax.com/link/HRYdLLuEdNgvHb5s You'll Bring :- Kubernetes: deep, expert-level understanding of Kubernetes as a system - internals, failure modes, capacity planning, and large-scale cluster management.- SRE principles: proven experience defining and operationalizing SLIs, SLOs, and error budgets across an engineering organization.- Observability at scale: deep expertise with a modern observability stack (OpenTelemetry, Prometheus, Grafana, Datadog, or Elasticsearch).- Cloud engineering & networking: expert grounding in AWS, GCP, or Azure - including networking fundamentals, security, and cost optimization at scale.- Distributed systems depth: able to reason through complex failure modes and architect systems that degrade gracefully under pressure.- Incident management: track record of leading major incident response, driving blameless postmortems, and implementing systemic fixes.- CI/CD and delivery safety: deep experience with CI/CD systems, progressive delivery, and the ability to improve pipeline reliability and safety at org scale.- AI in operations: experience building or deploying AI-assisted operational capabilities in production.- 12+ years of relevant hands-on experience, with a clear track record of technical leadership on complex, cross-team infrastructure initiatives at product companies operating at https://jobeax.com/link/e1ygwzk0lw9Z6sRs You :- You've operated at Staff or Principal level before and know what it means to own a technical domain.- You care about reliability the way a product engineer cares about user experience.- You're drawn to the question of what operations looks like when AI agents handle the repetitive analysis and humans own the judgment https://jobeax.com/link/aHViHIguu2zEpa0v Human With Us :- Being human isn't about checking every box on a list. It's about the experiences we have, people we meet, and the perspectives we share. So, if you have the skills but are hesitant to apply because of your background, apply anyway. We need amazing people like you to help us challenge the conventional and think differently about the problems that we're solving. (ref:hirist.tech)
Full time Type Of Hire : Experienced (relevant combo of work and education) Site Reliability Engineer – 4 - 6 Yrs – Pune Location FIS empowers the financial world with payment processing and banking solutions, including software, services and technology outsourcing. FIS’ more than 55,000 worldwide employees are passionate ...
Join us as a DevOps Engineer This is an opportunity for a driven individual to take on an exciting new career challenge You'll focus on enhancing the integration of development and platform activities to improve deployment and system reliability It's a chance to have a tangible effect on our function, put your existing ...
... development lifecycle (Shift-Left Security). Research, recommend, and implement best practices for DevSecOps and Kubernetes operations. 5+ years of experience in DevOps, Site Reliability Engineering, or Platform Engineering roles. - 2+ years of hands-on Kubernetes experience, including cluster provisioning, scaling, and troubleshooting. ...
... enjoy solving complex infrastructure challenges, we'd love to hear from you. Required Skills & Experience SRE & Production Operations (mandatory) - 5+ years in a Site Reliability Engineer or production operations engineering role, operating a SaaS or always-on service at scale. - Proven hands-on experience defining and operating ...
... cloud innovation to customers worldwide. We celebrate diversity and foster an inclusive environment, empowering our employees to be their authentic selves. The Devops Engineer would be an active member within the Voice Application Services team, responsible for providing automation and test support for the SW releases of Five9. ...
... mindset. You are comfortable working in distributed teams and enjoy improving systems through thoughtful https://jobeax.com/link/cUAbOSDobkfnp7aS working as a DevOps, Platform, or Site Reliability Engineer Hands‑on experience with public cloud platforms such as AWS, Azure, or GCP Experience with containerization and orchestration ...
We're looking for a seasoned Senior DevOps Engineer to lead cloud-native infrastructure initiatives across AWS, Azure and GCP. You'll architect scalable CI/CD pipelines using GitLab, manage containerized workloads with Kubernetes, Docker and Helm, and drive automation, security, and governance across multi-cloud environments. ...
... Title: Senior Site Reliability Engineer Job Summary We are seeking a highly motivated and experienced Site Reliability Engineer to join our team. As a Site Reliability Engineer, you will be responsible for ensuring the reliability, scalability, and availability of our systems by leveraging your expertise in DevOps practices ...
... and infrastructure while always thinking about reliability, scalability, resilience, security, and https://jobeax.com/link/Im0ECCXXkw7iwTBY : - Help build a Site Reliability Engineering culture by sharing best practices, approaches, documentation, and code with other engineering teams.- Apply automation and software to ...
... and infrastructure while always thinking about reliability, scalability, resilience, security, and https://jobeax.com/link/Im0ECCXXkw7iwTBY : - Help build a Site Reliability Engineering culture by sharing best practices, approaches, documentation, and code with other engineering teams.- Apply automation and software to ...
We are seeking a Site Reliability Engineer (SRE) to support and maintain a 24×7 Azure cloud environment, ensuring high availability, reliability, and performance of infrastructure and hosted services. This role requires the engineer to operate across L1 and L2 support responsibilities, combining proactive monitoring with ...
... and infrastructure while always thinking about reliability, scalability, resilience, security, and https://jobeax.com/link/Im0ECCXXkw7iwTBY : - Help build a Site Reliability Engineering culture by sharing best practices, approaches, documentation, and code with other engineering teams.- Apply automation and software to ...
... requirements for new telematics API capabilities and enhancements. Develop business cases and requirements to address recurring API issues and improve overall API reliability and customer experience. Lead and coordinate Field Follow Programs for new API endpoints, features, and functionality launches, ensuring successful adoption ...
Role Overview We are hiring backend engineers who specialize in reliability . SRE at our company is a software engineering role — focused on designing, building, and improving highly available distributed systems through code. This is not a DevOps / CI-CD / Terraform-heavy role. What You'll Do Design and build reliable, ...
Full time Type Of Hire : Experienced (relevant combo of work and education) Education Desired : We are hiring a Senior Lead Site Reliability Engineer to define, build, and operate always-on, low-latency , and highly secure payment platforms that power large-scale financial transactions. This is a senior technical role, ...
The Principal Full Stack Engineer is a technology leader responsible for end-to-end design, backend/frontend integration, and system scalability . This role requires deep expertise in C#/.NET, Angular, SQL, and AWS platform . Full Stack Architecture & Technical Leadership Own and drive the full-stack design for scalable ...
... Markets - Wealth Management - Investment BankingLocation : Gurugram (Hybrid - 3 Days Office)Experience : 8+ YearsAbout the Role :We are seeking a highly skilled Site Reliability Engineer (SRE) to join a leading Asset Management and Investment Operations team. The ideal candidate will be responsible for ensuring application ...
... resolve and/or escalate to service teams Implement changes to enable or improve infrastructure resilience, monitoring, and alerting Experience - 3+ years as a Site Reliability Engineer or in a Cloud Operations/DevOps role - 2+ years using golang, shell scripting and terraform - 2+ years as software developer in a SaaS environment ...