... understanding of relational and NoSQL databases, including optimization and management techniques.- Experience with distributed data processing frameworks (e.g., Hadoop, Spark, Flink).- Strong foundation in data modeling, ETL processes, and data warehousing principles.- Excellent written and verbal communication skills, with the ability ...
... Continuously learn mindset and apply new Hadoop ecosystem tools and data technologies. Required Skills and Experience - Proficiency in Hadoop ecosystems such as Spark, HDFS, Hive, Iceberg, Spark SQL. - Extensive experience with Apache Kafka, Apache Flink, and other relevant streaming technologies. - Proven ability to design ...
... experience) Shift : (GMT+05:30) Asia/Kolkata (IST) Opportunity Type : Remote Placement Type : Full Time Indefinite Contract(40 hrs a week/160 hrs a month) (*Spark, Generative AI models, LLM, rag, AWS, Docker, GCP, Kafka, Kubernetes, Machine Learning, Python, SQL We are looking for a Senior Machine Learning Engineer who ...
... Excellent problem-solving and debugging skills. Strong communication and collaboration skills. Desired Skills: Experience with Scala frameworks like Play, Akka, or Spark. Experience with cloud platforms (AWS, Azure, GCP). Experience with containerization technologies (Docker, Kubernetes). Experience with Agile development methodologies ...
... services or data intensive environments. - Strong experience building and orchestrating data pipelines with Apache Airflow. • Solid working knowledge of Apache Spark for large-scale data processing. - Strong Python and SQL skills, with a focus on clean, maintainable, production-quality code. Experience designing data models ...
... high-throughput ETL/ELT pipelines that ingest data from various sources into our Lakehouse. Code Quality & Tooling: Drive the adoption of dbt for transformation and PySpark for heavy lifting. You will be responsible for writing modular, reusable code that follows strict CI/CD practices. Implement 'Data SLAs.' You will build the ...
... development. Familiarity with cloud platforms like AWS, Azure, or GCP. Knowledge of data engineering best practices and scalable architectures. Experience with Apache Spark or PySpark. Knowledge of Docker, Kubernetes, and CI/CD pipelines. Experience with data governance and metadata management. Exposure to machine learning data ...
... platform engineering, data engineering, and analytics. Technical Skills: Proficiency in Python, Scala, or Java; strong experience with Databricks APIs, Apache Spark, and Delta Lake. Data Engineering Background: Solid experience in data warehousing, ETL processes, and data governance frameworks. AI/ML Exposure: Practical ...
... relational and NoSQL databases, including database design, optimization, and administration. - Experience with distributed data processing frameworks such as Hadoop, Spark, or Flink. - Solid understanding of data modeling, ETL processes, and data warehousing principles. - Strong troubleshooting and analytical skills for diagnosing ...
... relational and NoSQL databases, including database design, optimization, and administration. - Experience with distributed data processing frameworks such as Hadoop, Spark, or Flink. - Solid understanding of data modeling, ETL processes, and data warehousing principles. - Strong troubleshooting and analytical skills for diagnosing ...
... & Storage Technologies: Familiarity with AWS or Azure environments Python & scripting for automation and data processing Data streaming technologies (Kafka, Spark, Flink) Big Data platforms (Snowflake, Databricks, Hadoop, Teradata) Exposure to AI/ML and advanced analytics Experience with NoSQL databases Ability to design ...
... geometry simplification - to control runtime and cost.- Big-data and pipeline fluency: Advanced SQL plus distributed processing for large spatial workloads (Spark or Dask), and building reliable, repeatable data pipelines.- Productionizing models: Experience turning models into deployable, real-time APIs in collaboration ...
... common data structures and algorithms. Excellent in one of the following languages: Go/Python Mysql/Redis/Message Queue/Nosql. Familiarity with ElasticSearch OR Spark (for Data Feeds Team) Experience in architecture and developing large-scale distributed systems. (Excellent logic analysis capabilities, able to abstract and ...
... quantitative analysis methods or approaches in relation to credit models - Strong programing skills with 2+ years' hands-on and proven experience utilizing Python, Spark, SAS, SQL, AWS, Data Lake to perform statistical analysis and manage complex or large amounts of data and 4+ years of relevant experience Knowledge of Credit ...
... deduplication logic, and multi-source joins to produce clean, growth-analytics-ready tables consumed by Segment, and downstream marketing systems. Implement Databricks PySpark jobs for capturing growth data signals — including Spark streaming with Change Data Capture patterns for behavioral events with checkpoint-based fault tolerance ...
... data platforms. Advanced proficiency in SQL for data extraction, transformation, reconciliation, and performance optimization. Hands-on experience in Python and Spark for data engineering and automation. Strong experience with Azure Data Factory, Azure Databricks, Azure Storage, and related Azure data services. Experience ...
... and optimize end-to-end data pipelines on Databricks, following the Medallion Architecture principles. Build robust and scalable ETL/ELT pipelines using Apache Spark and Delta Lake to transform raw (bronze) data into trusted curated (silver) and analytics-ready (gold) data layers. Apply schema evolution and data versioning ...
... CI/CD via Azure DevOps. - Translate ambiguous business asks into reliable, documented, maintainable data products. 3+ years in data engineering, with significant Spark / PySpark at production scale. - Strong SQL and dimensional data modeling (Kimball-style star schemas, SCD patterns, surrogate keys). - Solid grasp of lakehouse ...
... services. Experience with event-driven and synchronous data pipelines and at least 1 year designing complex distributed systems. Experience with Golang, Python, AWS, Spark, and storage or data pipeline systems. Demonstrated expertise with Kafka or a comparable high-throughput messaging and streaming platform. Hands-on experience ...
... microservices in distributed systems - Proven experience designing and operating large-scale, data-intensive systems involving batch and stream processing (e.g., Spark, Kafka, Airflow or similar) - Deep understanding of system design fundamentals including scalability, fault tolerance, data consistency, and performance optimization ...