... new technologies and approaches to innovate with increasingly large data sets. Drive automation and efficiency in Data ingestion, data movement and data access workflows by innovation and collaboration. Understand, implement and enforce Software development standards and engineering principles in the Big Data space. Work ...
... Engineer will own the end-to-end operationalisation of machine learning, large language model (LLM), and agentic AI workloads on the Bajaj Finance Enterprise Data Platform — a 5PB+ medallion lakehouse built on Azure Databricks and Unity Catalog. This role sits at the intersection of data engineering, model lifecycle management, ...
... pipelines in cloud or modern data platforms. - 4+ years of hands-on experience with Databricks or similar cloud-based lakehouse platforms supporting enterprise data lakes, data warehouses, and business intelligence - 2+ years of experience using dbt (or similar transformation frameworks) including model structuring, tests, ...
... enhancing data storage and access, and ensuring seamless data consumption through APIs. The ideal candidate will work with Azure Cloud technologies to build robust data pipelines, data lakes, and marts to support business analysts and data scientists. Key Responsibilities Modern Data Platform Development: Build data lake components ...
... deploying complex data architectures within Microsoft Fabric and the broader Azure ecosystem.- Advanced proficiency in Python and SQL for building sophisticated data processing pipelines and complex analytical models.- Strong background in data engineering principles, including ETL/ELT design, data modeling, and performance ...
... have legal authorization to work in the country where this role is based on the first day of employment. Address client inquiries related to Addepar's portfolio data feeds and on general data product functionalities within established SLAs. Manage and complete requests from internal Data teams that require client outreach ...
... data pipelines and ETL workflows on AWS.- Develop efficient data processing solutions using python and pySpark.- Write complex and optimized SQL queries for data extraction, transformation, and analysis.- Build and manage ETL pipelines using AWS Glue.- Work with AWS Glue Data Catalog for metadata management and data discovery.- ...
... Databricks platform. The ideal candidate will have a robust background in backend development, Data science model basics, and cloud-based solutions to drive data-driven decision-making and innovation within our portfolio. Key skills - Python, FastAPI Framework, REST API design, Pydantic models, Async programming, Modular ...
... Own end-to-end delivery of data products—from raw source data ingestion through transformation to governed, consumption-ready datasets. Collaborate with the Data Platform team on pipeline integration, CI/CD workflows, and adherence to shared coding and deployment standards. - Data Quality & Analysis: Implement data quality ...
... with stakeholders. Key responsibilities Design and develop scalable data pipelines using AWS services. Implement ETL processes with AWS Glue, Lambda, and DBT. Work with PySpark and SQL to transform and cleanse data. Configure and manage Airflow workflows and Amazon EMR clusters. Integrate data services with API Gateway and ...
... data engineering, and DevOps automation. The ideal candidate will manage and optimize Cloudera environments, build CI/CD pipelines, and support enterprise-scale data processing workloads. Key Responsibilities Administer and support Cloudera CDP/CDH platforms, including HDFS, Hive, Spark, YARN, Hue, and CDE. Develop, deploy, ...
Roles and Responsibilities Architect and maintain enterprise-grade ELT and ETL data pipelines using Python, PySpark, Kafka, and Databricks to manage large-scale risk data. Build and deploy GenAI agents utilizing Google ADK, Google Flash 2.5+ LLMs, and Model Context Protocol (MCP) integrated with Human-in-the-Loop workflows. ...
... orchestration and automation of data workflows. Ensure the reliability, scalability, and efficiency of data pipelines for ingestion, transformation, and storage. Work with cross-functional teams to understand data needs and deliver high-quality solutions. Troubleshoot and resolve data pipeline issues in production environments. ...
... configurations, and cost optimization strategies.- Proven ability to write complex, highly optimized SQL queries and stored procedures to handle large-scale data transformations.- Strong experience in building and maintaining robust ETL/ELT workflows using modern data engineering tools and frameworks.- Exceptional communication ...
... design, API development, and data infrastructure for analytics and intelligence products. Key Responsibilities Data Pipeline Development Implement robust, scalable data pipelines using Microsoft Azure and Databricks stack Build reusable data pipeline components and frameworks Design and optimize data workflows for performance ...
... for critical database incidents. Establish governance frameworks, operational procedures, standards, and best practices for database management. Mentor junior database administrators and provide technical leadership across database initiatives. Expertise You'll Bring: - 10-15 years of experience as a SQL Server Database ...
We are seeking an experienced Lead Data Analyst to work on large-scale Energy & Utilities datasets. The ideal candidate will be responsible for analyzing tariff and pricing data, building data models, automating workflows, and delivering actionable insights to support business decisions.Role: Lead Data AnalystExperience: ...
... and generate location attributes from heterogeneous raw sources, and to reason about a feature store - versioning, reuse, freshness - as the backbone of the work.- Data fusion, hygiene, and geocoding: Integrating messy, heterogeneous datasets; imputation, anomaly/bias detection, deduplication and entity resolution; robust ...
... Databricks Platform: Practical experience with Databricks workspace, cluster management, notebooks, and job orchestration - Workspace AI Agent: Knowledge of Databricks Workspace AI Agent capabilities and integration - Data Modelling: Experience implementing data models including dimensional modeling, data vault, or lakehouse ...
... manage databases and data warehouses and optimise database performance and storage across multipul platforms, e.g. Snowflake, PostgreSQL RDS & Aurora. Ensure data quality, integrity, and reliability, whilst adhering to data security best practices. Work closely with data analysts, and other stakeholders to understand data ...