... Own end-to-end delivery of data products—from raw source data ingestion through transformation to governed, consumption-ready datasets. Collaborate with the Data Platform team on pipeline integration, CI/CD workflows, and adherence to shared coding and deployment standards. - Data Quality & Analysis: Implement data quality ...
... identify opportunities for improvement. Strong analytic skills related to working with unstructured data sets. Build processes supporting data transformation, data structures, metadata, dependency and workload management. Working knowledge of message queuing, stream processing, and highly scalable 'big data' data stores. ...
... will lead a team of data scientists, guide technical decisions, and help grow ProcDNA's data science practice through client relationships and new capabilities. Work with clients to translate business questions into well-scoped data science problems, and define the approach, data requirements and success measures. Design and ...
... and acceptance criteria. - Hands-on experience with data validation, reconciliation, source-to-target mapping, UAT, and defect management. - Understanding of data flows, cloud platforms, data repositories, and reporting/visualization layers; experience working with Engineering, Data, Analytics, and Technology teams in Agile ...
... help implement improvements to data workflows, templates, and automation to make delivery faster, more consistent, and more scalable. Hands-on experience in a data-focused role (for example, data specialist, data analyst, data migration, or similar) in a professional environment. Strong skills in Microsoft Excel, including ...
... with stakeholders. Key responsibilities Design and develop scalable data pipelines using AWS services. Implement ETL processes with AWS Glue, Lambda, and DBT. Work with PySpark and SQL to transform and cleanse data. Configure and manage Airflow workflows and Amazon EMR clusters. Integrate data services with API Gateway and ...
... Posted 2 weeks ago Job Be an early applicant About the job Job Purpose: Design and implement scalable data engineering and data warehouse solutions while managing data schemas, SQL query tuning, and code reviews. Who You Are: - 5+ years of experience in Data Engineering, with strong knowledge of Data Platforms and Data Warehousing ...
... data engineering, and DevOps automation. The ideal candidate will manage and optimize Cloudera environments, build CI/CD pipelines, and support enterprise-scale data processing workloads. Key Responsibilities Administer and support Cloudera CDP/CDH platforms, including HDFS, Hive, Spark, YARN, Hue, and CDE. Develop, deploy, ...
... years of experience evaluating and implementing data-engineering and software technologies - In addition, you have experience in programming languages and frameworks: SQL, Python, Spark, Databricks (Delta Lake) - Experience in data storages - SQL and NoSQL databases, Azure Data Lake Storage- and in developing data solutions, ...
... managing technical aspects of API framework design and development - Understanding of overall production Incidents / RCA for any feature that is delivered. \ - Working knowledge on --Java , Spring boot , microservices , AWS design , KAFKA ,Kubernetes ECS3 , Workflows such as Camunda BPM Preferred Skills: Experience with microservices ...
... including model governance and explainability- Establish data governance framework (catalog, lineage, ownership)- Ensure data quality, consistency, and compliance- Work closely with business leaders to translate requirements into data solutions- Lead cross-functional teams (data engineers, analysts, data scientists)- Manage vendors, ...
... activities and site communications. Collaborate with the project team to address issues, resolve discrepancies, and ensure trial milestones are met. Support data management activities by ensuring timely and accurate data collection and entry. Participate in audits and inspections as required, providing necessary documentation ...
... integrations, or service-oriented applications. - Experience working with cloud platforms, preferably AWS, and familiarity with deployment or CI/CD practices. - Working knowledge of software design principles, testing, debugging, performance optimization, and secure development. - Experience with SQL databases and Git-based ...
... datasets, building data pipelines, and writing efficient, scalable Python code. Key Responsibilities Develop, test, and maintain scalable Python applications. Work with large datasets to extract, transform, and analyze data. Build and optimize data pipelines and workflows. Perform data cleaning, validation, and preprocessing. ...
... efficient code using Scala and Spark - Work with Azure or on-premises Hadoop ecosystem for data processing - Solve real-world engineering problems related to data infrastructure - Collaborate with cross-functional teams to deliver data-driven solutions - Ensure data quality, reliability, and performance across data systems ...
... Skills Deep knowledge of statistics and mathematics, with experience designing rigorous analyses, testing assumptions and interpreting complex results. Strong data-wrangling and cleaning skills, including working with large or imperfect datasets, resolving quality issues and developing reproducible data pipelines. Strong ...
... Development - Experience with Azure Cloud Services (ADF, ADLS, Azure Blob Storage, Key Vault, etc.) - Strong knowledge of Data Ingestion, Connectors & Integration Frameworks - Advanced SQL and Performance Tuning skills - Experience building and optimizing scalable data pipelines - Understanding of Data Architecture & Data Warehouse ...
... role involves building deployment pipelines using CI/CD tools and supporting ML productionization while mentoring fellow team members. 5+ years of experience in data engineering with significant hands-on work in Databricks - Strong proficiency in PySpark, Spark SQL, Delta Lake, and Delta Live Tables - Advanced skills in Python ...
... happens to in-flight work when a worker or a database node dies - Hands-on with a distributed processing engine ( Spark/PySpark or equivalent) on non-trivial data volumes - Experience with an orchestrator ( Dagster , Airflow, Prefect, or equivalent) and a cloud platform (AWS/Azure) - Data-quality mindset : you build validation ...
Roles and Responsibilities Architect and maintain enterprise-grade ELT and ETL data pipelines using Python, PySpark, Kafka, and Databricks to manage large-scale risk data. Build and deploy GenAI agents utilizing Google ADK, Google Flash 2.5+ LLMs, and Model Context Protocol (MCP) integrated with Human-in-the-Loop workflows. ...