Platform Engineer - Cloud Infrastructure in Pune, India - Jobeax
Vacancy description
Platform Engineer - Cloud Infrastructure in Pune, India
Consulting Pandits
HybridMix of office and remote
India, Pune
Platform Engineer - Cloud Infrastructure in Pune, India is listed on Jobeax. Browse 30,000+ vacancies available.
Key Responsibilities :- Build and operate Kubernetes clusters, with cloud-hosted control planes and AI accelerator nodes joined as workers over site-to-site connectivity.- Register, label and taint accelerator worker nodes so that inference workloads schedule onto the correct hardware class and manage device scheduling and topology constraints.- Plan and execute cluster and operating system upgrades: RKE2 version upgrades, RHEL patching and major-version migration, etcd backup and restore, and control-plane node replacement.- Own cluster networking and storage end to end: CNI, ingress, DNS, load balancing, CSI drivers, persistent volume lifecycle, backup and tested disaster recovery.- Deploy, configure and upgrade the vendor AI platform stack, which is delivered as Helm charts from an OCI registry and must be installed in a defined dependency order.- Manage platform configuration as code: Helm values files, chart versions, namespace layout, registry pull secrets, artifact credentials and service-account key rotation.- Manage TLS certificates and DNS for the inference API and console endpoints, including CA-issued and wildcard certificates and automated renewal.- Operate the supporting data services the stack depends on, including operator-managed PostgreSQL, Redis queues and the bundled identity provider.- Design and operate cloud network infrastructure: virtual networks, subnets, routing, security groups, NAT and controlled egress, with ongoing cost analysis and right-sizing.- Own our side of IPSec connectivity into the accelerator racks, including tunnel endpoints, client-side routing and failover, and keep hybrid path latency inside inference latency budgets.- Build and maintain Terraform modules and Ansible automation, and reconcile cluster and platform state from version control through a GitOps workflow.- Implement cloud IAM, Kubernetes RBAC, namespace isolation, pod security standards, secrets rotation and hardening baselines, and produce evidence for security reviews.- Deploy and operate the monitoring and logging stack, define service-level objectives and alerts tied to inference availability and latency, and track cluster and accelerator capacity.- Support model bundle and deployment configuration changes through the platform's Kubernetes custom resources, in coordination with ML systems engineers.- Lead incident response for cluster and platform faults, write root-cause analyses that result in a tracked change, and maintain runbooks as a deliverable of each https://jobeax.com/link/8s1JU2HRuPVIpaSw Requirements :- Strong Linux administration on enterprise distributions, at the level of diagnosing service, storage, network, and kernel problems without escalation.- Production Kubernetes lifecycle experience: building clusters, upgrading them and recovering them when they break. RKE2, K3s or another CNCF-certified distribution is preferred over managed-only experience.- Helm proficiency beyond installing public charts: values management, chart versioning, multi-chart upgrade and rollback, and debugging failed releases.- Deep hands-on experience with at least one major public cloud and working knowledge of a second, covering networking, identity and cost management.- Terraform and Ansible at production scale, as reusable and reviewed code rather than one-off scripts.- Networking fundamentals: routing, NAT, firewalling, DNS and TLS termination, plus the ability to debug a hybrid connectivity problem end to end.- Working knowledge of OIDC authentication and how identity providers integrate with Kubernetes and platform applications.- Practical experience running a Prometheus and Grafana monitoring stack and a centralised log pipeline.- Scripting in Python and Bash, and comfort with YAML-heavy configuration.- Strong ownership and automation instinct, clear written communication for runbooks and incident reports, and availability for a shared on-call https://jobeax.com/link/kkVMsEJ7kyGJ0poP Requirements :- Experience operating AI or HPC clusters, including accelerator-aware scheduling and node health management.- Exposure to non-GPU AI accelerators and their distinct driver, runtime and scheduling models.- Experience deploying a vendor-supplied platform product into a customer or partner environment, including handover and upgrade cycles.- Policy-as-code tooling such as OPA, Kyverno or Sentinel, and experience with air-gapped or restricted-egress deployments. (ref:hirist.tech)
Role Overview MontyCloud is seeking a highly skilled Senior Engineer - Platform Engineering with a robust background in Python programming and extensive experience with AWS services. With at least 6years of relevant experience, the ideal candidate will be an expert in serverless development and event-driven architecture ...
... applications. Oracle Cloud Infrastructure (OCI) is building next-generation cloud security and compliance solutions, and we're looking for skilled Software Engineers to join our Cloud Guard / Assurance Service team. Be part of our mission to create the most secure cloud environment. In this pivotal, hands-on engineering ...
... skilled DevOps Engineer with 2- 4 years of experience to join our IT team. The ideal candidate will be responsible for automating, deploying, and maintaining cloud-based infrastructure and CI/CD pipelines while ensuring system reliability and https://jobeax.com/link/xrDEIpgaaOEL82CD Certified DevOps Engineer certification ...
... skilled DevOps Engineer with 2- 4 years of experience to join our IT team. The ideal candidate will be responsible for automating, deploying, and maintaining cloud-based infrastructure and CI/CD pipelines while ensuring system reliability and https://jobeax.com/link/xrDEIpgaaOEL82CD Certified DevOps Engineer certification ...
... and AI services.- Proficiency in agile methodologies.- Fluent in English, spoken and https://jobeax.com/link/JJwJ988bLYVlQifO :- Build and maintain scalable cloud infrastructure for AI platforms using IaC.- Create and maintain CI/CD pipelines for infrastructure platforms.- Automate operational efforts and establish self-service ...
Company : Very big MNCRole : GCP Data EngineerExperience : 8 - 15 yrsNotice Period : 30 daysLocation : PAN INDIATech Stack :- GCP data engineer- Pyspark- Dataflow- Dataproc- BigQuery- AirflowKey Responsibilities :- Design, develop, and optimize endtoend data pipelines using PySpark, Dataflow, and Dataproc- Implement data ...
... and engineering teams to improve deployment speed and reliability. Follow cloud security, access control, and infrastructure best practices. Contribute to platform standards, documentation, and automation initiatives. Must-Have Skills - 7+ years of experience in DevOps, Platform Engineering, Cloud Engineering, or a similar ...
... Google Cloud Platform (GCP).- Strong hands-on experience with Google Kubernetes Engine (GKE) and Kubernetes.- Relevant Google Cloud Certification (Professional Cloud DevOps Engineer, Professional Cloud Architect, or Professional Cloud Engineer) is mandatory.- Strong experience with Terraform for Infrastructure as Code (IaC).- ...
... excellence, and ensuring the reliability of production systems. This is a hands-on role suited for someone who thrives at the intersection of customer support, cloud operations, and platform engineering. Provide expert-level AWS support to our customers, including troubleshooting, incident response, and architectural guidance. ...
... https://jobeax.com/link/SxaWoIptPtnP4xy1 Skills & Experience :- 6 - 10 years of experience in cloud engineering, infrastructure engineering, system administration, or cloud migration, with significant hands-on experience in Microsoft Azure.- Strong hands-on experience with Azure infrastructure and cloud migration projects.- Experience ...
... Compensation commensurate with candidate experience As a Platform Engineer, you'll assume a pivotal role in designing, implementing, and managing scalable and reliable cloud infrastructure solutions. Collaborating closely with cross-functional teams, you'll leverage modern DevOps practices and tools to automate infrastructure provisioning, ...
... technical accuracy, and quality across training data. - Apply your production engineering experience to help improve AI systems' understanding of real-world cloud and DevOps environments. What You Bring - 4+ years of professional experience in cloud infrastructure, DevOps, site reliability engineering, platform engineering, ...
... enterprise-scale Azure cloud solutions aligned with business and technical requirements.- Design and implement Azure infrastructure using Terraform and Infrastructure as Code (IaC).- Develop reusable Terraform modules and establish best practices for infrastructure provisioning.- Automate cloud infrastructure deployment, ...
... focus on virtualization, clustering, storage, security, resiliency and automation. Your key responsibilities The I&O Platform Infrastructure Hyperconverged Infrastructure Specialist provides technical support and engineering of complex nature for the firm's Azure Stack HCI platforms and associated infrastructure services. ...
... environments - Preferred Experience AI/ML infrastructure or large‑scale data platforms Familiarity with model training, inference, evaluation, or data pipeline infrastructure Experience with regulated cloud environments (FedRAMP, GovCloud, CMMC Level 2) Contributions to platform engineering, open‑source infrastructure, or developer ...
... https://jobeax.com/link/S0veizAgyg6cBkoP Overview :We are seeking an experienced Azure Terraform Engineer to design, automate, and manage enterprise-scale Azure cloud environments. The ideal candidate will possess strong expertise in Infrastructure as Code (Terraform), Azure cloud services, DevOps, CI/CD automation, cloud ...
... scalable, secure, highly available, and cost-effective cloud solutions. The ideal candidate will have strong expertise in AWS architecture, cloud migration, infrastructure modernization, security, and enterprise-scale cloud https://jobeax.com/link/m7FLZdCihDPQzn7w role will involve working closely with engineering, security, ...
... Practical coding fluency in both React and Angular UI frameworks (Advanced practical mastery in 1 language, or proficient across 2 coding language stacks).- Cloud & DevOps : Strong hands - on experience navigating Google Cloud Platform (GCP) infrastructures alongside automated CI/CD pipelines.- Tenure : 6+ years in IT, ...
... Analytics.- Drive account growth through cross-selling and innovation-led https://jobeax.com/link/OZWsErPURNpr0waD Excellence & Mentorship : - Mentor architects and engineers across projects.- Build strong internal communities around Snowflake, dbt/modern data stack, and cloud-native data engineering.- Promote best practices in ...