... a related field. - Strong hands-on programming experience with Python for data processing, automation, and pipeline development. - Strong expertise in Apache Spark , particularly PySpark and/or Spark SQL. - Deep working knowledge of the AWS data ecosystem , including S3, Glue, Redshift, Athena, EMR, Kinesis, Lambda, and ...
... data engineering team to build, maintain, and optimize scalable data pipelines for large-scale data processing. Develop and implement ETL/ELT processes using PySpark, Spark, and other relevant tools to move and transform data from various sources. Assist in designing and deploying solutions in major cloud platforms such as ...
... experience) Shift : (GMT+05:30) Asia/Kolkata (IST) Opportunity Type : Remote Placement Type : Full Time Indefinite Contract(40 hrs a week/160 hrs a month) (*Spark, Generative AI models, LLM, rag, AWS, Docker, GCP, Kafka, Kubernetes, Machine Learning, Python, SQL We are looking for a Senior Machine Learning Engineer who ...
... tooling. You'll help us modernize our batch ETLs and lead our push into streaming and real-time fraud detection. We run on AWS and make extensive use of Scala Spark, dbt, and Terraform. Build and expose the data lake/lakehouse so teams can understand business and product performance and make better decisions. Design and ...
... drive both member satisfaction and business growth. Key Responsibilities: - Oversee daily operations, ensuring a seamless experience for all members - Maintain a sparkling clean, secure, and efficient workspace - Become the friendly face of The https://jobeax.com/link/J4ymvmVu0lkw6kdt, fostering a sense of community, and building ...
... design for analytics and reporting Performance optimization and scalability Preferred Experience Databricks Delta Lake experience Azure Synapse Analytics Python for data engineering Spark SQL optimization Real-time data streaming Data governance and metadata management Agile/Scrum development model CI/CD pipeline experience
... analytics. - Strong SQL and database fundamentals, with experience across relational and NoSQL technologies. - Experience with technologies such as Kafka, Flink, Spark, Trino, ClickHouse, OpenSearch/Elasticsearch , or comparable platforms is advantageous. - Experience building commercial cybersecurity products or platforms ...
... deduplication logic, and multi-source joins to produce clean, growth-analytics-ready tables consumed by Segment, and downstream marketing systems. Implement Databricks PySpark jobs for capturing growth data signals — including Spark streaming with Change Data Capture patterns for behavioral events with checkpoint-based fault tolerance ...
... on-prem systems to AWS, leveraging native AWS transformation technologies Key Responsibilities Implement end-to-end ETL pipelines using AWS native services like Spark, Step Function, EventBridge, Glue (PySpark, SQL), and Lambda for data extraction, transformation, and loading. Use pre-created utility & for seamless migration, ...
... Develop and optimize data warehouses/lakes using Redshift, BigQuery, Snowflake, or Delta Lake. Big Data & Streaming Work with distributed systems like Apache Spark, Kafka, or Flink for real-time and large-scale data processing. Manage feature stores for machine learning pipelines. Work closely with Data Scientists and ML ...
... CI/CD via Azure DevOps. - Translate ambiguous business asks into reliable, documented, maintainable data products. 3+ years in data engineering, with significant Spark / PySpark at production scale. - Strong SQL and dimensional data modeling (Kimball-style star schemas, SCD patterns, surrogate keys). - Solid grasp of lakehouse ...
... data engineering team to build, maintain, and optimize scalable data pipelines for large-scale data processing. Develop and implement ETL/ELT processes using PySpark, Spark, and other relevant tools to move and transform data from various sources. Assist in designing and deploying solutions in major cloud platforms such as ...
... business objectives. Experience with distributed computing frameworks and big data processing. Strong programming skills in languages commonly used with Apache Spark such as Scala, Java, or Python. Knowledge of data pipeline architecture and optimization techniques. Familiarity with cloud platforms and deployment of scalable ...
... similar certifications. Machine Learning: Knowledge of machine learning concepts and experience with popular ML libraries. Knowledge of big data processing (e.g., Spark, Hadoop, Hive, Kafka) Data Orchestration: Apache Airflow. Knowledge of CI/CD pipelines and DevOps practices in a cloud environment. Experience with ETL tools ...
... understanding: ETL/ELT pipelines, data warehousing, data lakes/lakehouses Advanced SQL (query optimization, not just SELECT statements) Hands-on with ONE of: Hive/Spark/Hadoop/HDFS, Snowflake/Databricks/BigQuery, or Airflow/DBT/Kafka/Trino Understands how data systems work end-to-end (not just BI layer) Problem-First Thinking ...
... Mastercard Pune 3-5 years Today $25.3K–38.6K/yr Full-time Onsite Skills Required LLM Gen AI Agentic AI RAG LangChain Vector Database Prompt Engineering Python Apache Spark SQL Azure Databricks AWS Description Mastercard powers economies and empowers people worldwide by providing secure, simple, and accessible digital payment solutions. ...
Are you a passionate, creative thinker who's ready to take charge of digital brand stories Join our team and here your ideas are legit gonna spark real impact, and your leadership shapes the creative team. We believe in collaboration, flexibility, and professional growth. Every voice matters and every campaign is a chance ...
... have skills : NA Minimum 3 Year(s) Of Experience Is Required Educational Qualification : 15 years full time education Key Skills: Databricks + SQL + Python/PySpark + Data Engineering Role Overview We are seeking an 6 Y of experienced Databricks Data Engineer with strong expertise in building scalable data pipelines, managing ...
... Scientist with e-commerce expertise to join our dynamic team. Must-Have Skills: Machine Learning Algorithms, Python or R Programming, Big Data Technologies (Hadoop, Spark, or equivalent), Data Visualization Tools (Tableau, Power BI, or similar), SQL and Data Modeling Expertise. Data Analysis & Insights Generation Machine Learning ...
... end-users. Qualifications: Core Skills: Python + SQL: Strong Python and SQL proficiency, including advanced queries and query optimization. Big Data: Apache Spark / PySpark for large-scale distributed data processing. Data Pipelines: Building ETL/ELT pipelines using Azure Data Factory, Fabric Data Flows, and Fabric Notebooks. ...