... System design involving large Petabytes of data with Databricks Lakehouse - Experience in modern AI/Data infrastructure patterns, Semantics layer Organizing data for AI agents (metadata, context) - AWS, GCP, Snowflake and Data pipeline frameworks like Airbyte/Airflow - Programming languages like Python, SQL, and potentially ...
... LLMs, AutoGen, OCR, Databricks, PySpark, Python, Azure DevOps, and enterprise AI platforms. The role requires strong hands-on experience across SQL Server, Azure Data Factory, Azure Data Lake, Microsoft Fabric, Logic Apps,Power BI, GitHub, and Data Modeling. Required Skills: - 10+ years of Software Engineering / Data Engineering ...
... modeling) and modern AI systems (LLMs, Generative AI and agentic frameworks). The role requires strong applied modeling skills across structured and unstructured data, and experience deploying models into production environments. This position involves end-to-end ownership of data science initiatives, from problem framing and ...
... interface performance during hypercare. Data From Process to Data Lake Data Validation & Process Improvement Support automation of CSV-Data upload by validation and data quality checks Perform data validation and cross-checks to ensure data quality and consistency Validate raw data and transform structured data to Data Model as ...
... ~ Manage budgets related to data sourcing, annotation operations, and vendor contracts. 8+ years of overall experience in GIS, remote sensing, or geospatial data operations; Strong hands-on knowledge of satellite imagery sourcing (optical, SAR, multispectral) and remote sensing data pipelines. - Proven experience managing ...
... processes. - Experience translating business requirements into features, user stories, data requirements, and acceptance criteria. - Hands-on experience with data validation, reconciliation, source-to-target mapping, UAT, and defect management. - Understanding of data flows, cloud platforms, data repositories, and reporting/visualization ...
... evaluate model performance and behavior. Investigate data distributions, model outputs, failure modes, and edge cases relevant to benchmark tasks. Write and run Python and SQL code to analyze data, create reports, and support evaluation workflows. Validate data quality, consistency, and correctness across datasets and experiments. ...
... inquiries related to Addepar's portfolio data feeds and on general data product functionalities within established SLAs. Manage and complete requests from internal Data teams that require client outreach and/or action to resolve data verification issues. Investigate client reported bugs and data processing issues, and triage ...
... drives / memory / transceivers and other hardware related tasks, as required by our server and network teams. Assisting with planning equipment moves or new data center build-outs. Coordinating the packing, shipping, and logistics of equipment to and from remote colocation sites. Maintaining data center documentation. ...
... Data Engineer will design, build, and maintain scalable data pipelines and solutions, ensuring high performance and reliability across COSMOS DB and related data platforms. Daily responsibilities include modeling and structuring data, implementing ETL processes, optimizing data storage and retrieval, and collaborating ...
... patterns. - Develop robust ingestion, ETL/ELT and integration pipelines using SQL, Python or PySpark as appropriate. - Design and maintain dimensional and analytical data models, including star schemas, facts, dimensions and Power BI semantic models. - Implement automated data-quality, validation, reconciliation and schema-change ...
... Point - Ability to prioritize tasks, manage multiple responsibilities and ensure deadlines are met without compromising on quality - Basic data handling and Data interpretation skills - Should be comfortable with 24x7 rotational shifts - Ability to pull data from numerous databases (using Excel and other data management ...
... factual accuracy, or comparing responses — when projects are available. While each project involves unique tasks, contributors may: - Carefully review provided data (text, images, or videos); - Label or classify content based on project guidelines; - Identify and flag factually incorrect, sensitive, inappropriate, or unclear ...
... factual accuracy, or comparing responses — when projects are available. While each project involves unique tasks, contributors may: - Carefully review provided data (text, images, or videos); - Label or classify content based on project guidelines; - Identify and flag factually incorrect, sensitive, inappropriate, or unclear ...
... factual accuracy, or comparing responses — when projects are available. While each project involves unique tasks, contributors may: - Carefully review provided data (text, images, or videos); - Label or classify content based on project guidelines; - Identify and flag factually incorrect, sensitive, inappropriate, or unclear ...
... Experience with MDM / Data Governance platforms such as Syniti, SAP MDG, Informatica MDM, or any home grown Strong understanding of data governance operating models, data ownership, data stewardship, workflow approvals, data quality rules, metadata, and data lifecycle management. Experience in data profiling, cleansing, harmonization, ...
... transformations, and performance tuning. - Good experience with Python for data processing, automation, and Lambda development. - Strong understanding of data warehousing, data modeling, ETL/ELT, and data lake concepts. - Experience building and supporting production data pipelines. - Experience troubleshooting data and pipeline performance ...
... factual accuracy, or comparing responses — when projects are available. While each project involves unique tasks, contributors may: - Carefully review provided data (text, images, or videos); - Label or classify content based on project guidelines; - Identify and flag factually incorrect, sensitive, inappropriate, or unclear ...
... knowledge of PostgreSQL. Experience developing REST APIs using FastAPI. Hands-on experience with Splink for entity resolution and record linkage. Strong SQL and Python programming skills. Experience with ETL/ELT pipeline development. Familiarity with cloud platforms like AWS, Azure, or GCP. Knowledge of data engineering best ...
... dashboards that visualize key performance indicators, enabling management to monitor business health in real-time.- Partner with internal business units to translate ambiguous requirements into structured data models and analytical frameworks that solve critical operational challenges.- Conduct deep-dive exploratory data analysis ...