Platform Engineer - Cloud Infrastructure in Pune, India - Jobeax
Vacancy description
Platform Engineer - Cloud Infrastructure in Pune, India
Consulting Pandits
HybridMix of office and remote
India, Pune
Platform Engineer - Cloud Infrastructure in Pune, India is listed on Jobeax. Browse 30,000+ vacancies available.
Key Responsibilities :- Build and operate Kubernetes clusters, with cloud-hosted control planes and AI accelerator nodes joined as workers over site-to-site connectivity.- Register, label and taint accelerator worker nodes so that inference workloads schedule onto the correct hardware class and manage device scheduling and topology constraints.- Plan and execute cluster and operating system upgrades: RKE2 version upgrades, RHEL patching and major-version migration, etcd backup and restore, and control-plane node replacement.- Own cluster networking and storage end to end: CNI, ingress, DNS, load balancing, CSI drivers, persistent volume lifecycle, backup and tested disaster recovery.- Deploy, configure and upgrade the vendor AI platform stack, which is delivered as Helm charts from an OCI registry and must be installed in a defined dependency order.- Manage platform configuration as code: Helm values files, chart versions, namespace layout, registry pull secrets, artifact credentials and service-account key rotation.- Manage TLS certificates and DNS for the inference API and console endpoints, including CA-issued and wildcard certificates and automated renewal.- Operate the supporting data services the stack depends on, including operator-managed PostgreSQL, Redis queues and the bundled identity provider.- Design and operate cloud network infrastructure: virtual networks, subnets, routing, security groups, NAT and controlled egress, with ongoing cost analysis and right-sizing.- Own our side of IPSec connectivity into the accelerator racks, including tunnel endpoints, client-side routing and failover, and keep hybrid path latency inside inference latency budgets.- Build and maintain Terraform modules and Ansible automation, and reconcile cluster and platform state from version control through a GitOps workflow.- Implement cloud IAM, Kubernetes RBAC, namespace isolation, pod security standards, secrets rotation and hardening baselines, and produce evidence for security reviews.- Deploy and operate the monitoring and logging stack, define service-level objectives and alerts tied to inference availability and latency, and track cluster and accelerator capacity.- Support model bundle and deployment configuration changes through the platform's Kubernetes custom resources, in coordination with ML systems engineers.- Lead incident response for cluster and platform faults, write root-cause analyses that result in a tracked change, and maintain runbooks as a deliverable of each https://jobeax.com/link/8s1JU2HRuPVIpaSw Requirements :- Strong Linux administration on enterprise distributions, at the level of diagnosing service, storage, network, and kernel problems without escalation.- Production Kubernetes lifecycle experience: building clusters, upgrading them and recovering them when they break. RKE2, K3s or another CNCF-certified distribution is preferred over managed-only experience.- Helm proficiency beyond installing public charts: values management, chart versioning, multi-chart upgrade and rollback, and debugging failed releases.- Deep hands-on experience with at least one major public cloud and working knowledge of a second, covering networking, identity and cost management.- Terraform and Ansible at production scale, as reusable and reviewed code rather than one-off scripts.- Networking fundamentals: routing, NAT, firewalling, DNS and TLS termination, plus the ability to debug a hybrid connectivity problem end to end.- Working knowledge of OIDC authentication and how identity providers integrate with Kubernetes and platform applications.- Practical experience running a Prometheus and Grafana monitoring stack and a centralised log pipeline.- Scripting in Python and Bash, and comfort with YAML-heavy configuration.- Strong ownership and automation instinct, clear written communication for runbooks and incident reports, and availability for a shared on-call https://jobeax.com/link/kkVMsEJ7kyGJ0poP Requirements :- Experience operating AI or HPC clusters, including accelerator-aware scheduling and node health management.- Exposure to non-GPU AI accelerators and their distinct driver, runtime and scheduling models.- Experience deploying a vendor-supplied platform product into a customer or partner environment, including handover and upgrade cycles.- Policy-as-code tooling such as OPA, Kyverno or Sentinel, and experience with air-gapped or restricted-egress deployments. (ref:hirist.tech)
... Deployment frequency & stability Alert quality and monitoring maturity Experience Required 8+ years in DevOps / Infrastructure Engineering / Site Reliability Engineering / Cloud Infrastructure roles. Experience managing production SaaS environments at scale. Strong working knowledge of at least one major cloud platform (AWS), ...
... Technology Solution Centers, we are a team of passionate engineers solving complex business challenges across the P&C insurance value chain. Our expertise spans Cloud Engineering, Application Engineering, Data Engineering, Core Engineering, Quality Engineering, and Domain capabilities. Through our Infinity Program, we invest ...
... Experience with large-scale data processing and analytics projects. Ability to work in Agile teams and collaborate effectively with various stakeholders. Preferred Qualifications Advanced AWS certifications. Experience with additional cloud platforms (e.g., Azure). Familiarity with other data engineering tools and frameworks.
... clients. Assist in building, managing, and optimizing CI/CD pipelines with tools like Jenkins, GitHub Actions, or GitLab CI. Provision, configure, and manage cloud infrastructure (AWS, Azure, or GCP) under the guidance of senior engineers. Work on containerization using Docker and support Kubernetes orchestration in development ...
... related field; or equivalent practical experience in DevOps / Cloud Engineering. Any relevant Cloud certifications (e.g., Azure Administrator, Azure DevOps Engineer Expert, Azure Solutions Architect, or Azure Data/AI certificationsto design, implement, and maintain CI/CD pipelines and cloud infrastructure. You will enable ...
... the backend infrastructure. This role requires actively engaging in the entire development lifecycle, ensuring seamless integration and functionality. The engineer collaborates with various teams to deliver innovative solutions that enhance client services, while embracing a cloud-first approach and agile methodologies ...
... performance and service delivery. Lead and mentor infrastructure, network, cloud, and support teams. Windows Server, Linux, VMware, Hyper-V Data Center Operations Cloud Platforms Microsoft Azure AWS Google Cloud Platform Networking ServiceNow / ITSM Tools Monitoring & Observability Platforms ITIL Framework Bachelor's Degree ...
... performance and service delivery. Lead and mentor infrastructure, network, cloud, and support teams. Windows Server, Linux, VMware, Hyper-V Data Center Operations Cloud Platforms Microsoft Azure AWS Google Cloud Platform Networking ServiceNow / ITSM Tools Monitoring & Observability Platforms ITIL Framework Bachelor's Degree ...
... off-the-shelf systems alone. - Partner with ML and platform teams to keep large runs alive and serving latency predictable. What We're Looking For - 5+ years in infrastructure or site reliability engineering, including 2+ years operating GPU clusters at scale.* - Demonstrated on-call ownership of infrastructure that mattered, with ...
DevOps Engineer Full-time 4-8 years of experience Qualification: BCA, MCA Online PSB Loans (OPL Innovate) is a digital credit infrastructure and fintech company focused on transforming India's lending ecosystem through technology, automation, and innovation. The company provides an end-to-end digital platform that enables ...
... of microservices and architectural components from inception and design, through deployment, operation, and refinement. Work closely with the developer infrastructure teams to expedite development infrastructure adoption of tools to advance your reliability roadmap by identifying needs for your supported engineering teams, ...
... neutral infrastructure that enables organizations to safely embrace this new era. This is an opportunity to do career-defining work. As a Senior Site Reliability Engineer you will champion all things pertaining to reliability at Okta for Auth0. Working closely with the Product Engineers, Quality Engineers, Platform Engineers ...
... access is restricted Experience building or operating break-glass access systems, session recording, or privileged access management tooling Prior work in a platform or internal developer experience team serving engineering orgs of 30+ engineers Flat, rotation-friendly engineering org (40 engineers) with real ownership and ...
... the Role :We are seeking a highly skilled Senior DevOps Engineer to join our product team within the DaVinci Data Engine domain. This role is ideal for an engineer with deep experience in distributed systems operations, Kubernetes cluster management, networking, high-availability architectures, and cloud platforms such ...
... or a related field (or equivalent experience). - 6+ years of experience in site reliability engineering, DevOps, or a related role. - Proficiency in cloud platforms (AWS, Azure, GCP) and cloud-native services. - Strong scripting and programming skills (Python, Bash, Go, or similar). - Experience with Infrastructure as ...
Site Reliability Engineer (Private Cloud / Virtualization) Cisco is transforming its platforms to run the next generation of cloud-native and multi-cloud services. This role offers a superb opportunity to transform how infrastructure platforms are developed and managed with full software automation. This team is responsible ...
Azure Migration Architect Job Description :A Senior Azure Migration Engineer/Architect is a highly skilled and experienced individual contributor at the forefront of cloud transformation https://jobeax.com/link/D4GLjfywL6FEvFfU role is pivotal for organizations looking to modernize their IT infrastructure, streamline operations, ...
Role Overview : The Senior GCP Cloud Network Engineer (C2) is a specialist Individual Contributor responsible for designing, implementing, operating, and optimizing GCP network infrastructure at https://jobeax.com/link/qAYqbkLnnVqh7KFx role focuses on deep technical execution, operational excellence, and collaboration, ...
... infrastructure on Azure (primary) and AWS Implement basic DevSecOps practices , including secrets management and policy controls Apply sound understanding of cloud security, networking, and enterprise environments Collaboration & Mentorship Collaborate effectively with engineers, testers, and platform teams Support and ...
... days a week working from the Hyderabad office during core working hours and 2 days working from home. Principal Search Engineer – Distributed Search & SaaS Platforms Nasuni is building the future of AI-powered enterprise file data infrastructure, and search is a foundational capability of that platform. This role is designed ...