Platform Engineer - Cloud Infrastructure in Pune, India - Jobeax
Vacancy description
Platform Engineer - Cloud Infrastructure in Pune, India
Consulting Pandits
HybridMix of office and remote
India, Pune
Platform Engineer - Cloud Infrastructure in Pune, India is listed on Jobeax. Browse 30,000+ vacancies available.
Key Responsibilities :- Build and operate Kubernetes clusters, with cloud-hosted control planes and AI accelerator nodes joined as workers over site-to-site connectivity.- Register, label and taint accelerator worker nodes so that inference workloads schedule onto the correct hardware class and manage device scheduling and topology constraints.- Plan and execute cluster and operating system upgrades: RKE2 version upgrades, RHEL patching and major-version migration, etcd backup and restore, and control-plane node replacement.- Own cluster networking and storage end to end: CNI, ingress, DNS, load balancing, CSI drivers, persistent volume lifecycle, backup and tested disaster recovery.- Deploy, configure and upgrade the vendor AI platform stack, which is delivered as Helm charts from an OCI registry and must be installed in a defined dependency order.- Manage platform configuration as code: Helm values files, chart versions, namespace layout, registry pull secrets, artifact credentials and service-account key rotation.- Manage TLS certificates and DNS for the inference API and console endpoints, including CA-issued and wildcard certificates and automated renewal.- Operate the supporting data services the stack depends on, including operator-managed PostgreSQL, Redis queues and the bundled identity provider.- Design and operate cloud network infrastructure: virtual networks, subnets, routing, security groups, NAT and controlled egress, with ongoing cost analysis and right-sizing.- Own our side of IPSec connectivity into the accelerator racks, including tunnel endpoints, client-side routing and failover, and keep hybrid path latency inside inference latency budgets.- Build and maintain Terraform modules and Ansible automation, and reconcile cluster and platform state from version control through a GitOps workflow.- Implement cloud IAM, Kubernetes RBAC, namespace isolation, pod security standards, secrets rotation and hardening baselines, and produce evidence for security reviews.- Deploy and operate the monitoring and logging stack, define service-level objectives and alerts tied to inference availability and latency, and track cluster and accelerator capacity.- Support model bundle and deployment configuration changes through the platform's Kubernetes custom resources, in coordination with ML systems engineers.- Lead incident response for cluster and platform faults, write root-cause analyses that result in a tracked change, and maintain runbooks as a deliverable of each https://jobeax.com/link/8s1JU2HRuPVIpaSw Requirements :- Strong Linux administration on enterprise distributions, at the level of diagnosing service, storage, network, and kernel problems without escalation.- Production Kubernetes lifecycle experience: building clusters, upgrading them and recovering them when they break. RKE2, K3s or another CNCF-certified distribution is preferred over managed-only experience.- Helm proficiency beyond installing public charts: values management, chart versioning, multi-chart upgrade and rollback, and debugging failed releases.- Deep hands-on experience with at least one major public cloud and working knowledge of a second, covering networking, identity and cost management.- Terraform and Ansible at production scale, as reusable and reviewed code rather than one-off scripts.- Networking fundamentals: routing, NAT, firewalling, DNS and TLS termination, plus the ability to debug a hybrid connectivity problem end to end.- Working knowledge of OIDC authentication and how identity providers integrate with Kubernetes and platform applications.- Practical experience running a Prometheus and Grafana monitoring stack and a centralised log pipeline.- Scripting in Python and Bash, and comfort with YAML-heavy configuration.- Strong ownership and automation instinct, clear written communication for runbooks and incident reports, and availability for a shared on-call https://jobeax.com/link/kkVMsEJ7kyGJ0poP Requirements :- Experience operating AI or HPC clusters, including accelerator-aware scheduling and node health management.- Exposure to non-GPU AI accelerators and their distinct driver, runtime and scheduling models.- Experience deploying a vendor-supplied platform product into a customer or partner environment, including handover and upgrade cycles.- Policy-as-code tooling such as OPA, Kyverno or Sentinel, and experience with air-gapped or restricted-egress deployments. (ref:hirist.tech)
... Practical coding fluency in both React and Angular UI frameworks (Advanced practical mastery in 1 language, or proficient across 2 coding language stacks).- Cloud & DevOps : Strong hands - on experience navigating Google Cloud Platform (GCP) infrastructures alongside automated CI/CD pipelines.- Tenure : 6+ years in IT, ...
... Analytics.- Drive account growth through cross-selling and innovation-led https://jobeax.com/link/OZWsErPURNpr0waD Excellence & Mentorship : - Mentor architects and engineers across projects.- Build strong internal communities around Snowflake, dbt/modern data stack, and cloud-native data engineering.- Promote best practices in ...
... Analytics.- Drive account growth through cross-selling and innovation-led https://jobeax.com/link/OZWsErPURNpr0waD Excellence & Mentorship : - Mentor architects and engineers across projects.- Build strong internal communities around Snowflake, dbt/modern data stack, and cloud-native data engineering.- Promote best practices in ...
... Analytics.- Drive account growth through cross-selling and innovation-led https://jobeax.com/link/OZWsErPURNpr0waD Excellence & Mentorship : - Mentor architects and engineers across projects.- Build strong internal communities around Snowflake, dbt/modern data stack, and cloud-native data engineering.- Promote best practices in ...
... systems/infrastructure engineering with strong Linux fundamentals and a learning-oriented mindset. Preferred Qualifications NVIDIA GPU ecosystem exposure. AI/ML infrastructure experience. Monitoring tools such as Prometheus, Grafana, Dynatrace, Datadog, or Zabbix. Hybrid cloud and datacenter infrastructure experience. NCCL and ...
... successful as a Software Engineer – Infrastructure, you should have experience with: Infrastructure design and implementation. Cloud automation and platform engineering. Cybersecurity best practices and secure infrastructure delivery. Linux, Bash, and PowerShell. Kubernetes and Docker. Basic understanding of cloud platforms. ...
About Sarvam Sarvam is building the bedrock of Sovereign AI for India. The company is developing India’s full-stack sovereign AI platform, building across research, models, infrastructure and applications with a singular focus on making AI genuinely work for India. Sarvam works with leading enterprises and public institutions ...
... predictable performance, strong isolation, and high availability at cloud scale. As a Software Engineer on the VCN Dataplane team, you will build and evolve the infrastructure, deployment automation, and operational systems that support our networking services. You will work across infrastructure-as-code, CI/CD, distributed systems, ...
... with Copado, Gearset, or Azure DevOps. Experience working in Agile teams. Nice to Have Experience with Cursor AI, Claude, or GitHub Copilot. Experience with Platform Events. Basic understanding of enterprise integrations. Skills: salesforce developer,service cloud,configuration,sales cloud,rest api,apex,lwc,integration
... Lightning Web Components (LWC), Visualforce, and Salesforce APIs. Develop and maintain integrations between Salesforce and external systems using REST, SOAP, Platform Events, and middleware platforms. Support and enhance Salesforce Sales Cloud Implement best practices for development, security, code quality, and release management. ...
... client solutions — wherever the Lead AI Platform Engineer and Field Product Managers identify agent opportunities. You'll work day-to-day alongside Suraj (AI Engineer) and under the direction of the Lead/Senior AI Platform Engineer, picking up implementation work as it's scoped and prioritized across programs. This role does ...
Role: Senior AI Platform Engineer Function: Platform Engineering / AI-ML / MLOps Type: Full-time Artificial Intelligence, Critical Infrastructure A research-first AI company incubated at the Indian Institute of Science (IISc). The company is building advanced AI for the planning and operations of critical networks. Its ...
... build pipelines). Practical exposure or functional training on Salesforce Agentforce architectures. Certifications (Required/Preferred): - Salesforce Certified Platform Developer I (PD1) is required; Platform Developer II (PD2) is a strong advantage. - Salesforce Certified Service Cloud Consultant or Experience Cloud Consultant. ...
Strong expertise in Apex, Lightning Web Components (LWC), Visualforce, and SOQL. • Experience with Salesforce integrations (REST/SOAP APIs, middleware tools). ~• Familiarity with Salesforce Sales Cloud, Service Cloud, and Marketing Cloud (must) Must Have Salesforce Platform Developer I certification
... develop Salesforce solutions using Apex, Lightning Web Components (LWC), SOQL, and declarative tools. Build and maintain features within Sales Cloud and Experience Cloud, including custom objects and Flows. Develop and support Salesforce integrations leveraging REST/SOAP APIs and MuleSoft Anypoint Platform APIs. Implement automation ...
... troubleshoot, and enhance production Salesforce solutions with a strong emphasis on Salesforce CPQ (Steelbrick), quote-to-order workflows, pricing, renewals, and platform automation. This is a senior hands-on engineering role expected to contribute to solution design, code quality, complex defect resolution, and reliable delivery. ...
... Environmental Conditions Office Job Description Job Summary We are seeking an experienced Senior Systems Infrastructure Engineer to manage and optimize enterprise infrastructure across on-premises and cloud environments. The role focuses on operational stability, VMware and Windows infrastructure, capacity management, automation, ...
... Title: Cloud Platform / Systems EngineerRole SummaryWe are seeking a highly skilled Cloud Platform / Systems Engineer with strong experience in Azure cloud infrastructure, Linux/Windows operating systems, automation, and platform engineering. The ideal candidate should possess hands-on production support and engineering experience ...