Platform Engineer - Cloud Infrastructure in Pune, India - Jobeax
Vacancy description
Platform Engineer - Cloud Infrastructure in Pune, India
Consulting Pandits
HybridMix of office and remote
India, Pune
Platform Engineer - Cloud Infrastructure in Pune, India is listed on Jobeax. Browse 30,000+ vacancies available.
Key Responsibilities :- Build and operate Kubernetes clusters, with cloud-hosted control planes and AI accelerator nodes joined as workers over site-to-site connectivity.- Register, label and taint accelerator worker nodes so that inference workloads schedule onto the correct hardware class and manage device scheduling and topology constraints.- Plan and execute cluster and operating system upgrades: RKE2 version upgrades, RHEL patching and major-version migration, etcd backup and restore, and control-plane node replacement.- Own cluster networking and storage end to end: CNI, ingress, DNS, load balancing, CSI drivers, persistent volume lifecycle, backup and tested disaster recovery.- Deploy, configure and upgrade the vendor AI platform stack, which is delivered as Helm charts from an OCI registry and must be installed in a defined dependency order.- Manage platform configuration as code: Helm values files, chart versions, namespace layout, registry pull secrets, artifact credentials and service-account key rotation.- Manage TLS certificates and DNS for the inference API and console endpoints, including CA-issued and wildcard certificates and automated renewal.- Operate the supporting data services the stack depends on, including operator-managed PostgreSQL, Redis queues and the bundled identity provider.- Design and operate cloud network infrastructure: virtual networks, subnets, routing, security groups, NAT and controlled egress, with ongoing cost analysis and right-sizing.- Own our side of IPSec connectivity into the accelerator racks, including tunnel endpoints, client-side routing and failover, and keep hybrid path latency inside inference latency budgets.- Build and maintain Terraform modules and Ansible automation, and reconcile cluster and platform state from version control through a GitOps workflow.- Implement cloud IAM, Kubernetes RBAC, namespace isolation, pod security standards, secrets rotation and hardening baselines, and produce evidence for security reviews.- Deploy and operate the monitoring and logging stack, define service-level objectives and alerts tied to inference availability and latency, and track cluster and accelerator capacity.- Support model bundle and deployment configuration changes through the platform's Kubernetes custom resources, in coordination with ML systems engineers.- Lead incident response for cluster and platform faults, write root-cause analyses that result in a tracked change, and maintain runbooks as a deliverable of each https://jobeax.com/link/8s1JU2HRuPVIpaSw Requirements :- Strong Linux administration on enterprise distributions, at the level of diagnosing service, storage, network, and kernel problems without escalation.- Production Kubernetes lifecycle experience: building clusters, upgrading them and recovering them when they break. RKE2, K3s or another CNCF-certified distribution is preferred over managed-only experience.- Helm proficiency beyond installing public charts: values management, chart versioning, multi-chart upgrade and rollback, and debugging failed releases.- Deep hands-on experience with at least one major public cloud and working knowledge of a second, covering networking, identity and cost management.- Terraform and Ansible at production scale, as reusable and reviewed code rather than one-off scripts.- Networking fundamentals: routing, NAT, firewalling, DNS and TLS termination, plus the ability to debug a hybrid connectivity problem end to end.- Working knowledge of OIDC authentication and how identity providers integrate with Kubernetes and platform applications.- Practical experience running a Prometheus and Grafana monitoring stack and a centralised log pipeline.- Scripting in Python and Bash, and comfort with YAML-heavy configuration.- Strong ownership and automation instinct, clear written communication for runbooks and incident reports, and availability for a shared on-call https://jobeax.com/link/kkVMsEJ7kyGJ0poP Requirements :- Experience operating AI or HPC clusters, including accelerator-aware scheduling and node health management.- Exposure to non-GPU AI accelerators and their distinct driver, runtime and scheduling models.- Experience deploying a vendor-supplied platform product into a customer or partner environment, including handover and upgrade cycles.- Policy-as-code tooling such as OPA, Kyverno or Sentinel, and experience with air-gapped or restricted-egress deployments. (ref:hirist.tech)
... tooling, agentic workflows, or intelligent automation platforms. Prior experience as a Staff, Principal, or equivalent senior technical leader in Site Reliability Engineering, Platform Engineering, Infrastructure Engineering, or Cloud Operations. #The Okta Experience Developing Talent and Fostering Connection + Community Our global ...
... SRE/DevOps engineers; Contribute to CSRE capability building, tooling standards, and internal knowledge assets. 6 – 8 years of total experience in SRE, DevOps, or cloud infrastructure roles. - Strong hands-on expertise in Google Cloud Platform (GCP) as primary cloud — GKE, VPC, IAM, Cloud Monitoring, Pub/Sub, BigQuery, Cloud ...
... we're not done. Not even close. Role: DevOps Engineer Location: Remote Role Overview: We are hiring a DevOps Engineer to build, automate, and maintain the infrastructure that keeps Joveo's platform running reliably at scale. You will own CI/CD pipelines, cloud infrastructure, and operational tooling — enabling engineering ...
... Senior Virtualization Engineer is responsible for the design, implementation, administration, optimization, and support of enterprise virtualization and hybrid cloud infrastructure environments. This role provides technical leadership for VMware and Nutanix platforms, leads infrastructure upgrade and migration initiatives, ...
... capabilities with Azure-hosted applications, workloads, AI platforms and infrastructure services.- Translate Azure secure networking requirements into Zscaler engineering patterns, connectivity models and deployment standards.- Partner closely with EY Azure architecture and cloud engineering teams to drive secure Zscaler product ...
... capabilities with Azure-hosted applications, workloads, AI platforms and infrastructure services.- Translate Azure secure networking requirements into Zscaler engineering patterns, connectivity models and deployment standards.- Partner closely with EY Azure architecture and cloud engineering teams to drive secure Zscaler product ...
... collaborate with start-ups to develop solutions for productivity optimization, workforce management, loyalty management, payments systems, and more. Senior Platform Engineer The Senior Platform Engineer position will require heavy hands-on work on one or more public cloud platforms leveraging several PaaS services and marketplace ...
... networking and containerization e Docker Kubernetes is also crucial Relevant certifications Industry standard credentials like the Microsoft Certified DevOps Engineer Expert AZ 400 are highly preferred Cloud Platform- Azure Devops- Azure Cloud Shell,Technology- Microsoft Technologies- Windows PowerShell,Technology- Python ...
... project requirements and performance. Why Join Apply your DevOps expertise to next-generation AI systems . Work on realistic infrastructure, automation, and cloud engineering scenarios. Help improve how AI systems reason about DevOps and production infrastructure . Contribute your existing engineering expertise without ...
... project requirements and performance. Why Join - Apply your DevOps expertise to next-generation AI systems . - Work on realistic infrastructure, automation, and cloud engineering scenarios. - Help improve how AI systems reason about DevOps and production infrastructure . - Contribute your existing engineering expertise without ...
... performance and service delivery. Lead and mentor infrastructure, network, cloud, and support teams. Windows Server, Linux, VMware, Hyper-V Data Center Operations Cloud Platforms Microsoft Azure AWS Google Cloud Platform Networking ServiceNow / ITSM Tools Monitoring & Observability Platforms ITIL Framework Bachelor's Degree ...
... performance and service delivery. Lead and mentor infrastructure, network, cloud, and support teams. Windows Server, Linux, VMware, Hyper-V Data Center Operations Cloud Platforms Microsoft Azure AWS Google Cloud Platform Networking ServiceNow / ITSM Tools Monitoring & Observability Platforms ITIL Framework Bachelor's Degree ...
... off-the-shelf systems alone. - Partner with ML and platform teams to keep large runs alive and serving latency predictable. What We're Looking For - 5+ years in infrastructure or site reliability engineering, including 2+ years operating GPU clusters at scale.* - Demonstrated on-call ownership of infrastructure that mattered, with ...
... troubleshoot Azure resources, pipelines, and deployments for reliability, scalability, and cost efficiency. Stay current with Azure DevOps, IaC, and cloud engineering best practices to drive continuous improvement. What You'll Bring: Strong 5+ years of experience as a DevOps Engineer or Cloud Infrastructure Engineer in ...
... Experience: 7-12 years The Data engineer is responsible for managing and operating upon Databricks, Dbt, SSRS, SSIS, AWS DWS, AWS APP Flow, PowerBI/Tableau. The engineer will work closely with the customer and team to manage and operate cloud data platform. Resolving pipeline issues / Proactive monitoring for sensitive batches ...
... portfolio management theories, models, and practices, within an Agile or DevOps environment Strong experience with AWS, including cloud-native architecture and platform design Expertise in Terraform and Infrastructure as Code (IaC) for provisioning, managing, and automating cloud infrastructure Hands-on experience with GitLab ...
... this role, you will act as a core infrastructure anchor, managing day-to-day Application Basis administration and collaborating directly with the SAP Enterprise Cloud Services (ECS) team to govern high-performance, multi-tenant environments hosted on SAP RISE and Google Cloud Platform (GCP).The ideal candidate is a seasoned ...
... Actions, Jenkins, or Google Cloud Build) for automated testing, evaluation, and deployment. Working knowledge of Terraform for provisioning GCP-based AI infrastructure. Nice-to-Have: ● Experience building AI platform capabilities in a multi-cloud environment (GCP and Microsoft Azure), ideally supporting a 'build once, leverage ...
... Experience: Minimum of 12 years of hands-on experience in a technical support or engineering role related to Databricks Data Intelligence platform, cloud data platforms, or big data technologies. - Technical Skills: A deep understanding of Databricks architecture and Apache Spark™, along with experience in cloud platforms ...