Platform Engineer - Cloud Infrastructure in Pune, India - Jobeax
Vacancy description
Platform Engineer - Cloud Infrastructure in Pune, India
Consulting Pandits
HybridMix of office and remote
India, Pune
Platform Engineer - Cloud Infrastructure in Pune, India is listed on Jobeax. Browse 30,000+ vacancies available.
Key Responsibilities :- Build and operate Kubernetes clusters, with cloud-hosted control planes and AI accelerator nodes joined as workers over site-to-site connectivity.- Register, label and taint accelerator worker nodes so that inference workloads schedule onto the correct hardware class and manage device scheduling and topology constraints.- Plan and execute cluster and operating system upgrades: RKE2 version upgrades, RHEL patching and major-version migration, etcd backup and restore, and control-plane node replacement.- Own cluster networking and storage end to end: CNI, ingress, DNS, load balancing, CSI drivers, persistent volume lifecycle, backup and tested disaster recovery.- Deploy, configure and upgrade the vendor AI platform stack, which is delivered as Helm charts from an OCI registry and must be installed in a defined dependency order.- Manage platform configuration as code: Helm values files, chart versions, namespace layout, registry pull secrets, artifact credentials and service-account key rotation.- Manage TLS certificates and DNS for the inference API and console endpoints, including CA-issued and wildcard certificates and automated renewal.- Operate the supporting data services the stack depends on, including operator-managed PostgreSQL, Redis queues and the bundled identity provider.- Design and operate cloud network infrastructure: virtual networks, subnets, routing, security groups, NAT and controlled egress, with ongoing cost analysis and right-sizing.- Own our side of IPSec connectivity into the accelerator racks, including tunnel endpoints, client-side routing and failover, and keep hybrid path latency inside inference latency budgets.- Build and maintain Terraform modules and Ansible automation, and reconcile cluster and platform state from version control through a GitOps workflow.- Implement cloud IAM, Kubernetes RBAC, namespace isolation, pod security standards, secrets rotation and hardening baselines, and produce evidence for security reviews.- Deploy and operate the monitoring and logging stack, define service-level objectives and alerts tied to inference availability and latency, and track cluster and accelerator capacity.- Support model bundle and deployment configuration changes through the platform's Kubernetes custom resources, in coordination with ML systems engineers.- Lead incident response for cluster and platform faults, write root-cause analyses that result in a tracked change, and maintain runbooks as a deliverable of each https://jobeax.com/link/8s1JU2HRuPVIpaSw Requirements :- Strong Linux administration on enterprise distributions, at the level of diagnosing service, storage, network, and kernel problems without escalation.- Production Kubernetes lifecycle experience: building clusters, upgrading them and recovering them when they break. RKE2, K3s or another CNCF-certified distribution is preferred over managed-only experience.- Helm proficiency beyond installing public charts: values management, chart versioning, multi-chart upgrade and rollback, and debugging failed releases.- Deep hands-on experience with at least one major public cloud and working knowledge of a second, covering networking, identity and cost management.- Terraform and Ansible at production scale, as reusable and reviewed code rather than one-off scripts.- Networking fundamentals: routing, NAT, firewalling, DNS and TLS termination, plus the ability to debug a hybrid connectivity problem end to end.- Working knowledge of OIDC authentication and how identity providers integrate with Kubernetes and platform applications.- Practical experience running a Prometheus and Grafana monitoring stack and a centralised log pipeline.- Scripting in Python and Bash, and comfort with YAML-heavy configuration.- Strong ownership and automation instinct, clear written communication for runbooks and incident reports, and availability for a shared on-call https://jobeax.com/link/kkVMsEJ7kyGJ0poP Requirements :- Experience operating AI or HPC clusters, including accelerator-aware scheduling and node health management.- Exposure to non-GPU AI accelerators and their distinct driver, runtime and scheduling models.- Experience deploying a vendor-supplied platform product into a customer or partner environment, including handover and upgrade cycles.- Policy-as-code tooling such as OPA, Kyverno or Sentinel, and experience with air-gapped or restricted-egress deployments. (ref:hirist.tech)
... reports from various platforms - Process Return Authorizations and Credit Memos in NetSuite - Collaborate with various teams including Support, Legal, OrderOps, Engineering etc. on advanced billing & contract issues to resolve customer disputes - Work with the Cloud Billing (Eng) Team to review/verify accounts subject to suspension ...
... Comprehensive knowledge of Source-to-Pay (S2P) and Order-to-Cash (O2C) https://jobeax.com/link/yptFbvezapmzswse Profile :- Experience : 10+ years in Oracle SCM (EBS/Cloud), with at least 3 full-lifecycle Oracle Cloud SCM implementations.- Technical-Functional Hybrid : Ability to navigate the functional setup manager and configure ...
... provisioning for lab build projects. - Audio Visual system and Network Infrastructure from design to delivery for new programs - Audio Visual system and Network Infrastructure from design to delivery for Day 2 programs - Audio Visual system and Network Infrastructure from design to delivery for replacement programs - Audio Visual ...
Role :Driven by the passion to improve quality of people's lives, Sonata Software continues to grow as market leader in the industry. We want to accelerate our business transformation in order to reach more people, more effectively. Join our IT organization supporting Sonata Software operations across India.
... solutions.- Strong command over Kubernetes and container orchestration, with the ability to manage the lifecycle of AI-heavy workloads across hybrid or multi-cloud infrastructures.- Proven ability to communicate complex technical strategies to non-technical stakeholders, ensuring alignment between engineering output and ...
About the role: We are looking for a DevOps Engineer with 4–5 years of experience to manage and enhance our cloud infrastructure, CI/CD processes, and deployment automation. The role involves working with AWS, containerized environments, and modern DevOps practices to ensure scalable, secure, and highly available systems. ...
... plan, build and execute software infrastructure for complete CI/CD for the build, deploy and testing of large-scale software products.- Manage and monitor the engineering labs, including provisioning new hardware and virtual machines, debugging problems impacting engineering and provisioning access and accounts.- Develop custom ...
As a Senior DevOps Engineer (Public Cloud – Azure & GCP), you will lead the design, automation, and operations of secure and reliable cloud platforms that enable teams to deliver software at scale. You will set engineering standards for CI/CD and Infrastructure as Code, strengthen observability and incident response, and ...
... Science, Engineering, or a related field.- Minimum of 5+ years of proven experience as a DevOps Engineer or in a similar capacity.- Strong understanding of Infrastructure as Code tools using AWS CloudFormation, Terraform, Ansible, Puppet, Chef, or an equivalent.- Extensive experience with cloud computing concepts and utilizing ...
... expertise with cloud infrastructure (AWS, GCP or Azure), containers and orchestration, CI/CD, monitoring stacks, automation - Strong experience in Site Reliability Engineering, DevOps, or Production Operations—preferably supporting large-scale systems - Solid understanding of incident management, reliability engineering, and microservice ...
The DevOps Engineer will lead the design, implementation, and maintenance of robust CI/CD pipelines, cloud infrastructure (AWS/Azure/GCP), containerized applications, and infrastructure automation solutions. The role requires advanced technical expertise in DevOps tools and scripting, infrastructure as code (Terraform), ...
... https://jobeax.com/link/ZitjVqvb84V4AOyQ to Have : - Experience working on large-scale, high-availability platforms.- Cloud platform experience.- Containerization and CI/CD experience.- Experience leading backend engineering teams or technical initiatives.- Strong understanding of observability and reliability engineering. (ref:hirist.tech)
... misconfigurations, insecure patterns, excessive permissions, control gaps, and opportunities to reduce risk. Support secure design and architecture reviews for cloud infrastructure, SaaS platforms, internal applications, product features, APIs, and integrations. Partner with engineering, product, IT, and security stakeholders ...
... data platform experience — you have built and run distributed processing and storage on infrastructure you or your organisation managed, not only on managed cloud services. - Data modelling depth — you can design a canonical model across messy sources and defend the decisions behind it. - Production pipeline engineering ...
... focus on reproducing realistic enterprise workloads and failure conditions to identify bottlenecks and reliability issues before they reach customers. The engineer will work closely with development, QA, platform, and DevSecOps teams to define measurable performance expectations and integrate non-functional testing into ...
... experience with object-oriented programming and software engineering practices, preferably using Python and C#/.NET. - Strong experience building and operating cloud infrastructure and services, preferably AWS. - Experience with modern data platforms and orchestration technologies such as Snowflake and Airflow. - Experience ...
... containment, tuning. The other half is engineering and program work: building detection content, automating response, and running security capabilities across cloud, identity, endpoint, and AI governance. You'll work across Engineering, SRE, IT, Cloud, Legal, and Compliance, and you'll own outcomes rather than tickets. Key ...
Role: As a Senior Databricks Engineer you will be responsible for automating platform operations, managing CI/CD pipelines, and supporting deployments across multiple environments. The role ensures platform stability, security, and scalability while enabling faster and more reliable software delivery. Key Responsibilities: ...