Platform Engineer - Cloud Infrastructure in Pune, India - Jobeax
Vacancy description
Platform Engineer - Cloud Infrastructure in Pune, India
Consulting Pandits
HybridMix of office and remote
India, Pune
Platform Engineer - Cloud Infrastructure in Pune, India is listed on Jobeax. Browse 30,000+ vacancies available.
Key Responsibilities :- Build and operate Kubernetes clusters, with cloud-hosted control planes and AI accelerator nodes joined as workers over site-to-site connectivity.- Register, label and taint accelerator worker nodes so that inference workloads schedule onto the correct hardware class and manage device scheduling and topology constraints.- Plan and execute cluster and operating system upgrades: RKE2 version upgrades, RHEL patching and major-version migration, etcd backup and restore, and control-plane node replacement.- Own cluster networking and storage end to end: CNI, ingress, DNS, load balancing, CSI drivers, persistent volume lifecycle, backup and tested disaster recovery.- Deploy, configure and upgrade the vendor AI platform stack, which is delivered as Helm charts from an OCI registry and must be installed in a defined dependency order.- Manage platform configuration as code: Helm values files, chart versions, namespace layout, registry pull secrets, artifact credentials and service-account key rotation.- Manage TLS certificates and DNS for the inference API and console endpoints, including CA-issued and wildcard certificates and automated renewal.- Operate the supporting data services the stack depends on, including operator-managed PostgreSQL, Redis queues and the bundled identity provider.- Design and operate cloud network infrastructure: virtual networks, subnets, routing, security groups, NAT and controlled egress, with ongoing cost analysis and right-sizing.- Own our side of IPSec connectivity into the accelerator racks, including tunnel endpoints, client-side routing and failover, and keep hybrid path latency inside inference latency budgets.- Build and maintain Terraform modules and Ansible automation, and reconcile cluster and platform state from version control through a GitOps workflow.- Implement cloud IAM, Kubernetes RBAC, namespace isolation, pod security standards, secrets rotation and hardening baselines, and produce evidence for security reviews.- Deploy and operate the monitoring and logging stack, define service-level objectives and alerts tied to inference availability and latency, and track cluster and accelerator capacity.- Support model bundle and deployment configuration changes through the platform's Kubernetes custom resources, in coordination with ML systems engineers.- Lead incident response for cluster and platform faults, write root-cause analyses that result in a tracked change, and maintain runbooks as a deliverable of each https://jobeax.com/link/8s1JU2HRuPVIpaSw Requirements :- Strong Linux administration on enterprise distributions, at the level of diagnosing service, storage, network, and kernel problems without escalation.- Production Kubernetes lifecycle experience: building clusters, upgrading them and recovering them when they break. RKE2, K3s or another CNCF-certified distribution is preferred over managed-only experience.- Helm proficiency beyond installing public charts: values management, chart versioning, multi-chart upgrade and rollback, and debugging failed releases.- Deep hands-on experience with at least one major public cloud and working knowledge of a second, covering networking, identity and cost management.- Terraform and Ansible at production scale, as reusable and reviewed code rather than one-off scripts.- Networking fundamentals: routing, NAT, firewalling, DNS and TLS termination, plus the ability to debug a hybrid connectivity problem end to end.- Working knowledge of OIDC authentication and how identity providers integrate with Kubernetes and platform applications.- Practical experience running a Prometheus and Grafana monitoring stack and a centralised log pipeline.- Scripting in Python and Bash, and comfort with YAML-heavy configuration.- Strong ownership and automation instinct, clear written communication for runbooks and incident reports, and availability for a shared on-call https://jobeax.com/link/kkVMsEJ7kyGJ0poP Requirements :- Experience operating AI or HPC clusters, including accelerator-aware scheduling and node health management.- Exposure to non-GPU AI accelerators and their distinct driver, runtime and scheduling models.- Experience deploying a vendor-supplied platform product into a customer or partner environment, including handover and upgrade cycles.- Policy-as-code tooling such as OPA, Kyverno or Sentinel, and experience with air-gapped or restricted-egress deployments. (ref:hirist.tech)
... Solutions Architect, Staff Engineer, Lead Integration Engineer, or Founding Solutions Engineer.- Experience integrating enterprise platforms such as Salesforce, SAP, Oracle, Workday, NetSuite, or ServiceNow.- Knowledge of cloud platforms (AWS, GCP, Azure), Docker, Kubernetes, and modern data infrastructure. (ref:hirist.tech)
... Experience 2 year(s) 2 year(s) Apply By Not Provided Posted 2 days ago Job Be an early applicant About the job Yotta Data Services is Indias leading sovereign AI infrastructure, cloud platform and data center services company, enabling enterprises, governments, startups, and digital platforms to build, deploy, and scale next-Generation ...
... NexaStack AI – Inference AI Infrastructure for Agentic Systems Our mission is to help enterprises transform into future-ready, data-driven organizations through Cloud Native Platforms, Decision-Driven Analytics, and AI-powered solutions . We are looking for a hands-on Agentic AI Engineer to develop and integrate AI agents ...
... experience with Linux Administration Experience working with AWS Cloud Services Good understanding of DevOps principles and practices Additional Responsibilities: ~ Location of Posting Hyderabad Preferred Skills: Technology- Cloud Platform- Amazon Webservices DevOps,Technology- Infrastructure-Server Administration- Linux Admin
... industry - Experience in Linux, Bash, Git, and programming and scripting (Python, Java, Ruby, Go, or a similar language) - Hands on in deploying and managing infrastructure within a major cloud provider (AWS, GCP, AZURE, or a similar platform) - Required experience and expertise in Chef/Ansible/Puppet, Hadoop, Terraform, Centos, ...
... GitHub for version control, CI/CD workflows, and automated testing of data models and pipelines. Cloud Platforms: Familiarity with AWS services such as S3 and cloud-based data infrastructure. Arrive, including brands like EasyPark, Flowbird, RingGo, ParkMobile and Parkopedia, is a leading global mobility platform. Present ...
... Science, Engineering, or a related field.- Minimum of 5+ years of proven experience as a DevOps Engineer or in a similar capacity.- Strong understanding of Infrastructure as Code tools using AWS CloudFormation, Terraform, Ansible, Puppet, Chef, or an equivalent.- Extensive experience with cloud computing concepts and utilizing ...
... partnering with product and engineering teams to prioritize, track, and reduce security risk in line with established remediation timelines.- Drive remediation of cloud and infrastructure security findings, partnering with Cloud Security and Engineering teams to track, prioritize, and address AWS security vulnerabilities, misconfigurations, ...
... system design and architecture. Deep experience in database architecture, query optimisation, and data pipeline design (PostgreSQL preferred). Background in cloud infrastructure, DevOps, or platform/reliability engineering. Prior experience in operations technology, logistics, supply chain, or device lifecycle management ...
... of the network infrastructure. Implement and maintain cloud network connectivity (VPCs, ExpressRoute, Direct Connect, VPNs, Interconnects). Collaborate with Cloud, DevOps, Security, and IT teams to support new deployments and infrastructure initiatives. Perform network assessments, audits, and optimisations for cost, performance, ...
... the future of secure, enterprise-grade cloud automation and system connectivity. We are looking for a Senior Software Engineer to join our global Automation Engineering team to design, scale, and govern our enterprise integration substrate using AWS cloud services and modern iPaaS platforms. This is a senior individual contributor ...
... infrastructure we own, such as Snowflake, dbt Cloud, Airflow, and Kafka, with deep ownership of governance, reliability, and cost. Your customers are other engineers, analytics engineers, data scientists, and operations teams. The platform you build is what makes their work possible. If you enjoy building developer-facing ...
... for test generation, debugging, infrastructure provisioning, and productivity improvements. - Contribute to continuous improvement initiatives across quality engineering processes, automation, and release readiness. Requirements - 4+ years of experience in networking QA, cloud platform testing, network engineering, or related ...
... experience in one or more relevant programming or command languages Has hands-on experience with configuration management tools- Has strong knowledge on AWS Cloud and its native services Has strong experience on building and deploying infrastructure using Terraform/Cloud formation- Has deep knowledge of networking concepts ...
... experiences and perspectives to build a company culture that fuels growth through innovation. Platform Science is seeking a highly skilled Senior Software Engineer to join our Orion (Fleet Management System) department in Chennai. In this role, you will be a key contributor to our cloud-native full-stack development efforts, ...
... Network, Windows, Cloud, Monitoring and ITSM platform and any AI skills will be an added advantage Program Delivery & Financial Governance Lead end-to-end infrastructure and network operations projects (Network, Windows, Monitoring platform and Cloud) Manage project scope, timelines, budgets, risks, and deliverables Drive ...
... your team so you’re continuously innovating – doing more with less while remaining secure. And that’s just the beginning. Expert Systems Engineer – Windows Infrastructure Job Description: Responsibilities: - Serve as the primary Windows Infrastructure SME supporting Windows Server, Active Directory, Cloud infra and VMware ...
... scaling infrastructure services using Amazon Web Services or Microsoft Azure Skilled with infrastructure tools like Ansible, Puppet, Chef, or Terraform for infrastructure as code, monitoring tools (e.g., Skilled in the understanding of using core cloud application infrastructure services including identity platforms, networking, ...
SAP ABAP Cloud Location: Bengaluru Experience: 8–12 Years CTC: Up to ₹23 LPA Role: Team Lead / Consultant Education: 15 Years Full-Time Education Role Overview: We are looking for an experienced SAP ABAP Cloud professional to design, develop, enhance and maintain scalable SAP applications. The role involves working with ...