... management plans and documentation, ensuring compliance with industry standards and regulatory requirements. - Participate in the identification and resolution of data discrepancies and issues, working to streamline data flow and enhance data quality throughout the study lifecycle. - Collaborate closely with cross-functional ...
... design, API development, and data infrastructure for analytics and intelligence products. Key Responsibilities Data Pipeline Development Implement robust, scalable data pipelines using Microsoft Azure and Databricks stack Build reusable data pipeline components and frameworks Design and optimize data workflows for performance ...
... high-frequency market data, translating market behavior hypotheses into quantitative signals. Implement signals within our internal simulation/backtesting framework, iterating between exploratory data analysis and framework-based implementation. Validate features through backtesting across historical data, checking behavior ...
... enhancing data storage and access, and ensuring seamless data consumption through APIs. The ideal candidate will work with Azure Cloud technologies to build robust data pipelines, data lakes, and marts to support business analysts and data scientists. Key Responsibilities Modern Data Platform Development: Build data lake components ...
... configurations, and cost optimization strategies.- Proven ability to write complex, highly optimized SQL queries and stored procedures to handle large-scale data transformations.- Strong experience in building and maintaining robust ETL/ELT workflows using modern data engineering tools and frameworks.- Exceptional communication ...
... Hands-on experience with Azure Databricks (Delta Live Tables, Unity Catalog preferred) Strong Apache Spark skills (PySpark / Spark SQL) Experience migrating workloads from legacy data warehouse or Synapse environments to a Databricks Lakehouse Ability to re-implement governance and control frameworks natively in Databricks ...
Data Science Data Engineer Work Type: Full Time At VIDA we're building the future of digital identity. As a Data Engineer you'll work across the stack—from platform and infrastructure to data pipelines and end-user tooling. You'll help us modernize our batch ETLs and lead our push into streaming and real-time fraud detection. ...
... data pipelines and ETL workflows on AWS.- Develop efficient data processing solutions using python and pySpark.- Write complex and optimized SQL queries for data extraction, transformation, and analysis.- Build and manage ETL pipelines using AWS Glue.- Work with AWS Glue Data Catalog for metadata management and data discovery.- ...
... techniques. Hands-on experience integrating and optimizing Starburst/Trino, including connecting Starburst from Lambda and Glue ETL jobs. Experience with NoSQL databases such as DynamoDB, MongoDB. Experience working with data formats including Avro, Parquet, JSON, XML, and CSV. Comfortable challenging your peers and leadership ...
... details related to data engineering or a relevant https://jobeax.com/link/cwwOL3BLJkve6nkU Skills :- Azure Databricks, Python, pyspark, sql, ADF- Proficiency in data pipeline design and implementation- Strong understanding of data factory and orchestration toolsPreferred Skills:- Familiarity with advanced data processing techniques- ...
... S3, AWS Glue, Athena, Redshift Experience designing cloudnative data lakes and data warehouse architectures on AWS Deep understanding of batch and streaming data pipelines Experience building scalable, faulttolerant data ingestion and transformation workflows SQL & Python (Mandatory) Strong SQL expertise Writing complex ...
... using - Azure Data Factory, Azure Pipelines, Azure DevOps, and Git for version control. - Proven experience working with Azure Databricks and PySpark/Spark for data engineering tasks and Power BI as data analytics work - In-depth knowledge of performance optimization techniques for large-scale data processing, including code ...
... role involves building deployment pipelines using CI/CD tools and supporting ML productionization while mentoring fellow team members. 5+ years of experience in data engineering with significant hands-on work in Databricks - Strong proficiency in PySpark, Spark SQL, Delta Lake, and Delta Live Tables - Advanced skills in Python ...
... develop, and implement predictive models and algorithms using advanced statistical techniques and machine learning frameworks. Data Analysis: Conduct thorough data exploration, cleaning, and transformation to prepare datasets for modeling efforts, ensuring high data quality and accuracy. Collaboration: Work with cross-functional ...
... Feature Store for Data Science implementations Overall experience of 0-2 yrs. in Advanced Analytics/ Business Intelligence/Data Warehousing Strong Background in Data Transformation Skills – SQL (Mandatory) & Python . Able to understand and align to long term vision of team A phenomenal work environment with massive ownership ...
... — comfortable using LLMs/AI copilots as part of daily workflow (analysis, documentation, communication, process design), not just as a novelty. - Experience working with AI/ML teams and understanding of how annotated data feeds into model training. - Strong understanding of quality frameworks (QA/QC) for spatial data and ...
... identify opportunities for improvement. Strong analytic skills related to working with unstructured data sets. Build processes supporting data transformation, data structures, metadata, dependency and workload management. Working knowledge of message queuing, stream processing, and highly scalable 'big data' data stores. ...
... Certification – Data Migration Track . Minimum 8+ years of total experience, with at least 5+ years in Guidewire Data Migration. Proven expertise with Guidewire Data Migration Framework, Mammoth ETL, and related Guidewire tools. Strong knowledge of Guidewire PolicyCenter data model and its integration with external legacy ...
GCP Data Engineer Mandatory Skills Relevant work experience between 3 to 7 years. Proficient Experience on designing, building and operationalizing large-scale enterprise data solutions using at least four GCP services among Data Flow, Data Proc, Pub Sub, BigQuery, Cloud Functions, Composer, GCS Proficient hands-on programming ...