Platform Engineer - Cloud Infrastructure in Pune, India - Jobeax
Vacancy description
Platform Engineer - Cloud Infrastructure in Pune, India
Consulting Pandits
HybridMix of office and remote
India, Pune
Platform Engineer - Cloud Infrastructure in Pune, India is listed on Jobeax. Browse 30,000+ vacancies available.
Key Responsibilities :- Build and operate Kubernetes clusters, with cloud-hosted control planes and AI accelerator nodes joined as workers over site-to-site connectivity.- Register, label and taint accelerator worker nodes so that inference workloads schedule onto the correct hardware class and manage device scheduling and topology constraints.- Plan and execute cluster and operating system upgrades: RKE2 version upgrades, RHEL patching and major-version migration, etcd backup and restore, and control-plane node replacement.- Own cluster networking and storage end to end: CNI, ingress, DNS, load balancing, CSI drivers, persistent volume lifecycle, backup and tested disaster recovery.- Deploy, configure and upgrade the vendor AI platform stack, which is delivered as Helm charts from an OCI registry and must be installed in a defined dependency order.- Manage platform configuration as code: Helm values files, chart versions, namespace layout, registry pull secrets, artifact credentials and service-account key rotation.- Manage TLS certificates and DNS for the inference API and console endpoints, including CA-issued and wildcard certificates and automated renewal.- Operate the supporting data services the stack depends on, including operator-managed PostgreSQL, Redis queues and the bundled identity provider.- Design and operate cloud network infrastructure: virtual networks, subnets, routing, security groups, NAT and controlled egress, with ongoing cost analysis and right-sizing.- Own our side of IPSec connectivity into the accelerator racks, including tunnel endpoints, client-side routing and failover, and keep hybrid path latency inside inference latency budgets.- Build and maintain Terraform modules and Ansible automation, and reconcile cluster and platform state from version control through a GitOps workflow.- Implement cloud IAM, Kubernetes RBAC, namespace isolation, pod security standards, secrets rotation and hardening baselines, and produce evidence for security reviews.- Deploy and operate the monitoring and logging stack, define service-level objectives and alerts tied to inference availability and latency, and track cluster and accelerator capacity.- Support model bundle and deployment configuration changes through the platform's Kubernetes custom resources, in coordination with ML systems engineers.- Lead incident response for cluster and platform faults, write root-cause analyses that result in a tracked change, and maintain runbooks as a deliverable of each https://jobeax.com/link/8s1JU2HRuPVIpaSw Requirements :- Strong Linux administration on enterprise distributions, at the level of diagnosing service, storage, network, and kernel problems without escalation.- Production Kubernetes lifecycle experience: building clusters, upgrading them and recovering them when they break. RKE2, K3s or another CNCF-certified distribution is preferred over managed-only experience.- Helm proficiency beyond installing public charts: values management, chart versioning, multi-chart upgrade and rollback, and debugging failed releases.- Deep hands-on experience with at least one major public cloud and working knowledge of a second, covering networking, identity and cost management.- Terraform and Ansible at production scale, as reusable and reviewed code rather than one-off scripts.- Networking fundamentals: routing, NAT, firewalling, DNS and TLS termination, plus the ability to debug a hybrid connectivity problem end to end.- Working knowledge of OIDC authentication and how identity providers integrate with Kubernetes and platform applications.- Practical experience running a Prometheus and Grafana monitoring stack and a centralised log pipeline.- Scripting in Python and Bash, and comfort with YAML-heavy configuration.- Strong ownership and automation instinct, clear written communication for runbooks and incident reports, and availability for a shared on-call https://jobeax.com/link/kkVMsEJ7kyGJ0poP Requirements :- Experience operating AI or HPC clusters, including accelerator-aware scheduling and node health management.- Exposure to non-GPU AI accelerators and their distinct driver, runtime and scheduling models.- Experience deploying a vendor-supplied platform product into a customer or partner environment, including handover and upgrade cycles.- Policy-as-code tooling such as OPA, Kyverno or Sentinel, and experience with air-gapped or restricted-egress deployments. (ref:hirist.tech)
... Associate OR Databricks Certified Data Engineer Professional Additional Certifications (Preferred) - Databricks Certified Associate Developer for Apache Spark - Cloud platform certifications (Azure Data Engineer Associate, AWS Certified Data Analytics, or Google Cloud Professional Data Engineer) - Relevant data engineering ...
... (e.g., Experience with scripting languages (e.g., Python, Bash, or Go) to automate routine security tasks and threat hunting. Strong expertise in securing public cloud environments (e.g., AWS, Microsoft Azure, or Google Cloud Platform). Ability to work full-time on-site at our designated office location. Certifications (Preferred ...
... to power dashboards and insights via tools like Athena, QuickSight, or Redash. Contribute to infrastructure-as-code and CI/CD practices for deployment across cloud environments (preferably AWS). Document architecture, data flow, and support runbooks; continuously improve platform performance and resilience. Integrate with ...
... making a profound difference for the dreamers and builders in the world. We are seeking a Staff Software Engineer (IC5) to join our App Platform and Functions engineering organization. App Platform is DigitalOcean’s managed PaaS: customers ship apps from source to production on Kubernetes-backed infrastructure without managing ...
... Familiarity with CI/CD, infrastructure automation, monitoring, logging, and distributed tracing. - Experience leading large-scale technical initiatives or platform migrations. - Contributions to engineering standards, internal platforms, or open-source projects. Education Bachelor’s or Master’s degree in Computer Science, ...
About the Role : We are looking for a Big Data Engineer to design, build, and operate large-scale data pipelines and analytical infrastructure that transform high-volume raw data into reliable, query-ready datasets for analytics, reporting, and data-driven https://jobeax.com/link/HouIpt6AtbK1KWOU data platform ingests and ...
... and continuous evolution of our core platform infrastructure. In this role, you will operate at the intersection of high-throughput transactional systems and cloud-native architecture. You will be responsible for driving high-level and low-level designs (HLD/LLD), establishing engineering guardrails, optimizing complex ...
... rely on it daily. It handles protected health information at scale, making security a core product requirement — not a compliance checkbox. As a Cybersecurity Engineer, you will own the security posture of the company's AWS cloud infrastructure and its Python/Django, FastAPI, https://jobeax.com/link/JSd7UYCSg572jJDY, and React ...
Navan is looking for a Staff Android Engineer to serve as a technical architect for our mobile ecosystem. You will ensure our mobile platform can scale to support engineers and complex global travel requirements without friction. Drive the Vision: Partner with engineering and product leadership to identify, plan, and execute ...
... with modern DevOps, FinOps, security, and developer-productivity tooling to continuously improve our processes. 5–10 years of overall systems/infrastructure engineering experience, including hands-on work building and operating infrastructure platform-as-a-service capabilities. - Strong hands-on knowledge of cloud services ...
... understanding of technology and technical skills, with a focus on software. engineering. Roles will include Server Engineer, Infrastructure / SRE Engineer, Support Engineer, etc. - 10+ years of equivalent experience in a high-performing IT company - Proven track record in full-cycle recruitment processes and platform management ...
... etc.)Strong debugging using logs, APM tools, and tracing frameworksExperience with CI/CD pipelines (Jenkins, GitHub Actions, etc.)Knowledge of Infrastructure-as-Code (Terraform, CloudFormation, CDK)Preferred SkillsAWS CertificationsExperience with Kubernetes / EKSExposure to Kafka, SQS, Redshift or data streaming platforms
Data Science Data Engineer Work Type: Full Time At VIDA we're building the future of digital identity. As a Data Engineer you'll work across the stack—from platform and infrastructure to data pipelines and end-user tooling. You'll help us modernize our batch ETLs and lead our push into streaming and real-time fraud detection. ...
... role, not a research or analytics https://jobeax.com/link/kwiyzHaIipQTXHy4 will work from first principles to engineer robust, scalable, and observable GenAI platforms, owning critical components across the lifecycle - from document ingestion and retrieval to LLM orchestration, API serving, and cloud https://jobeax.com/link/OwtFYbiKAwffKwEs ...
... NexaStack AI – Inference AI Infrastructure for Agentic Systems Our mission is to help enterprises transform into future-ready, data-driven organizations through Cloud Native Platforms, Decision-Driven Analytics, and AI-powered solutions . We are looking for a hands-on Agentic AI Engineer to develop and integrate AI agents ...
... environments and software development best practices. Preferred Skills Working knowledge of Rust and ability to develop or maintain Rust-based services. Experience with cloud platforms, preferably Google Cloud Platform (GCP). Experience with Kubernetes, Docker, CI/CD pipelines, and Infrastructure as Code tools. Knowledge of blockchain ...
... Bethesda, Maryland, and Austin, Texas, and international offices in Amsterdam, London, Bangalore, and Tokyo, we are leveraging our open APIs and cyber-secure cloud infrastructure to reshape the future of physical security. Join us as we build the world's most robust platform for video intelligence and smart space automation. ...
DevOps EngineerJob Overview:We are seeking a skilled and motivated DevOps Engineer with over 6 years of experience to join our dynamic team. The ideal candidate will have a strong background in software engineering, IT operations, and cloud infrastructure management. As a DevOps Engineer, you will be responsible for automating, ...
... observability, alerting, and reliability engineering , using tools such as Prometheus, Grafana, Datadog, Splunk, ELK, or equivalent ecosystems. Strong command of cloud platforms and open systems ( AWS, Azure, or GCP ), including infrastructure-as-code , platform automation, and cloud-native design patterns. Significant experience ...