Platform Engineer - Cloud Infrastructure in Pune, India - Jobeax
Vacancy description
Platform Engineer - Cloud Infrastructure in Pune, India
Consulting Pandits
HybridMix of office and remote
India, Pune
Platform Engineer - Cloud Infrastructure in Pune, India is listed on Jobeax. Browse 30,000+ vacancies available.
Key Responsibilities :- Build and operate Kubernetes clusters, with cloud-hosted control planes and AI accelerator nodes joined as workers over site-to-site connectivity.- Register, label and taint accelerator worker nodes so that inference workloads schedule onto the correct hardware class and manage device scheduling and topology constraints.- Plan and execute cluster and operating system upgrades: RKE2 version upgrades, RHEL patching and major-version migration, etcd backup and restore, and control-plane node replacement.- Own cluster networking and storage end to end: CNI, ingress, DNS, load balancing, CSI drivers, persistent volume lifecycle, backup and tested disaster recovery.- Deploy, configure and upgrade the vendor AI platform stack, which is delivered as Helm charts from an OCI registry and must be installed in a defined dependency order.- Manage platform configuration as code: Helm values files, chart versions, namespace layout, registry pull secrets, artifact credentials and service-account key rotation.- Manage TLS certificates and DNS for the inference API and console endpoints, including CA-issued and wildcard certificates and automated renewal.- Operate the supporting data services the stack depends on, including operator-managed PostgreSQL, Redis queues and the bundled identity provider.- Design and operate cloud network infrastructure: virtual networks, subnets, routing, security groups, NAT and controlled egress, with ongoing cost analysis and right-sizing.- Own our side of IPSec connectivity into the accelerator racks, including tunnel endpoints, client-side routing and failover, and keep hybrid path latency inside inference latency budgets.- Build and maintain Terraform modules and Ansible automation, and reconcile cluster and platform state from version control through a GitOps workflow.- Implement cloud IAM, Kubernetes RBAC, namespace isolation, pod security standards, secrets rotation and hardening baselines, and produce evidence for security reviews.- Deploy and operate the monitoring and logging stack, define service-level objectives and alerts tied to inference availability and latency, and track cluster and accelerator capacity.- Support model bundle and deployment configuration changes through the platform's Kubernetes custom resources, in coordination with ML systems engineers.- Lead incident response for cluster and platform faults, write root-cause analyses that result in a tracked change, and maintain runbooks as a deliverable of each https://jobeax.com/link/8s1JU2HRuPVIpaSw Requirements :- Strong Linux administration on enterprise distributions, at the level of diagnosing service, storage, network, and kernel problems without escalation.- Production Kubernetes lifecycle experience: building clusters, upgrading them and recovering them when they break. RKE2, K3s or another CNCF-certified distribution is preferred over managed-only experience.- Helm proficiency beyond installing public charts: values management, chart versioning, multi-chart upgrade and rollback, and debugging failed releases.- Deep hands-on experience with at least one major public cloud and working knowledge of a second, covering networking, identity and cost management.- Terraform and Ansible at production scale, as reusable and reviewed code rather than one-off scripts.- Networking fundamentals: routing, NAT, firewalling, DNS and TLS termination, plus the ability to debug a hybrid connectivity problem end to end.- Working knowledge of OIDC authentication and how identity providers integrate with Kubernetes and platform applications.- Practical experience running a Prometheus and Grafana monitoring stack and a centralised log pipeline.- Scripting in Python and Bash, and comfort with YAML-heavy configuration.- Strong ownership and automation instinct, clear written communication for runbooks and incident reports, and availability for a shared on-call https://jobeax.com/link/kkVMsEJ7kyGJ0poP Requirements :- Experience operating AI or HPC clusters, including accelerator-aware scheduling and node health management.- Exposure to non-GPU AI accelerators and their distinct driver, runtime and scheduling models.- Experience deploying a vendor-supplied platform product into a customer or partner environment, including handover and upgrade cycles.- Policy-as-code tooling such as OPA, Kyverno or Sentinel, and experience with air-gapped or restricted-egress deployments. (ref:hirist.tech)
System Engineer L3 Contract duration: Location: remote Shift hours: We are seeking a highly skilled and experienced Systems Engineer L3 to join our Platform Engineering Observability team. In this pivotal role, you will be at the forefront of designing, implementing, and maintaining our AWS cloud infrastructure, as well ...
As a Senior DevOps Engineer , you will play a key role in designing, building, and operating cloud infrastructure and CI/CD platforms that support HERE's products and services. You will contribute to improving system reliability, automation, observability, and security while collaborating with cross‑functional engineering ...
DevOps Engineer Job Type: Contractor Location: Remote Openings: 30 Compensation: $50–$100/hr ($104,000–$208,000/year annualized) About the Role As a DevOps Engineer, you’ll apply your expertise in cloud infrastructure, automation, deployment, and platform engineering to contribute to projects focused on training and improving ...
... such as messaging, scheduling, data The app incorporates a shared data/network layer via Kotlin Multiplatform. This is a mid-level role. About The Project The engineer will join an established mobile team working across iOS and Android platforms, reporting to an Engineering Manager. ~ concurrency, and Core Data. Develop intuitive ...
... maintenance activities while minimizing disruption to operations. 1–6 years of relevant experience in platform engineering, DevOps, site reliability, backend engineering, cloud operations, or a related technical field. - Strong expertise in Linux, Docker, Kubernetes, container infrastructure, and production system troubleshooting. ...
Job Description We are looking for an experienced NCC Charging & Platform Engineer who will be responsible for deployment, configuration, and integration of Nokia Converged Charging (NCC) solutions along with end-to-end charging business configurations. Role: Tariff/Deployment Engineer â€' NCC / Online Charging Experience: ...
Our partner is looking for a Platform Engineer based in India. You will work across Kubernetes, containerization, Linux, DevOps, CI/CD, Python, and cloud technologies to support system stability. The position combines proactive monitoring, incident response, troubleshooting, automation, and continuous improvement. You will ...
... and perspective to help EY become even better, too. Join us and build an exceptional experience for yourself, and a better working world for all. SIEM SOAR/Platform Engineer The ideal candidate will have extensive experience with Palo Alto Cortex XSOAR (formerly Demisto) and a strong background in security automation and ...
... Computer Engineering or equivalent Desired Skills And Experience Solid experience with AWS technologies such as EC2, S3, SQS, SNS, RDS, IAM, Lambda, DynamoDB and Cloud Formation Managing cloud based infrastructure such as AWS, Azure or GCP Running Docker containers on Kubernetes (EKS) and Istio Strong programming skills with ...
About the Role:We are looking for an experienced AWS Data Engineer with 6 - 8 years of hands-on experience in designing, developing, and maintaining scalable data engineering solutions on AWS. The ideal candidate should have strong expertise in Python, PySpark, SQL, and AWS data services, with experience building robust ...
... supporting RAG, Agentic Workflows, MCP Integrations, Multi-model Orchestration, Model Evaluation Frameworks, Prompt Management, and AI Observability. Design AI infrastructure leveraging cloud-native technologies. Engineering Excellence Define coding standards and AI engineering best practices. Establish AI SDLC processes. Lead ...
... or platform engineering role - Hands-on experience with CI/CD tooling, source control, and automated build and release pipelines - Experience automating infrastructure provisioning, configuration, and system administration tasks - Working knowledge of cloud or hybrid infrastructure environments and common system architectures ...
... self-service platforms for developers. Manage source control (Bitbucket/GitHub), code quality gates (SonarQube), and artifact repositories (Sonatype Nexus). Infrastructure as Code (IaC) & Cloud Engineering: Provision, configure, and manage cloud infrastructure on AWS/Azure using Infrastructure as Code (IaC) tools such as Terraform, ...
... methodologies, and tools, such as CI/CD pipelines, configuration management, and containerization. Experience with cloud platforms (e.g., AWS, Azure, GCP) and infrastructure as code (e.g., AWS Experience along with AI/ML Basics is a plus Enthusiastically follow technology trends, software engineering best practice and technologies ...
... observability engineers. 5+ years of experience in DevOps, Observability, or similar roles, including hands-on production experience operating Kubernetes based stack. - Cloud Platforms: Extensive hands-on experience with Azure, including its observability services Infrastructure as Code: Proficiency with Terraform, AWS CloudFormation, ...
We are seeking an experienced Senior Systems Engineer to join our Global Infrastructure team. In this role, you will be responsible for designing, implementing, migrating, optimizing, and supporting enterprise infrastructure across on-premises datacenters and cloud platforms. You will play a key role in cloud modernization, ...
... Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent experience. - 7+ years of experience in cloud engineering, site reliability engineering, or related roles. - Strong experience with cloud platforms (AWS, GCP, Azure) and cloud-native services. - Proficiency in infrastructure-as-code tools (Terraform), ...
... latency and correctness. Strong data-modelling capability across one or more of relational, document, key-value, search or graph data stores. Experience with cloud-native engineering on AWS, Azure or comparable cloud platforms . Strong engineering practices across automated testing, CI/CD, infrastructure automation, observability, ...
... underserved hearing care market, we want to accelerate our business transformation in order to reach more people, more effectively. Join our team as a Azure DevOps Engineer , where you will design, deploy, and manage secure and scalable Microsoft Azure cloud infrastructure. You will work across Azure networking, PaaS services, ...