Big Data Engineer - Hadoop/Spark/Scala in Bengaluru, India
HybridMix of office and remote
India, Bengaluru
Big Data Engineer - Hadoop/Spark/Scala in Bengaluru, India is listed on Jobeax. Browse 30,000+ vacancies available.
Role Overview : We are seeking a seasoned Big Data Engineer to join our high-performing data platform team. In this role, you will be responsible for architecting, developing, and maintaining robust data pipelines that process massive datasets to fuel our analytical engines. You will work closely with cross-functional teams, including data scientists, product managers, and infrastructure engineers, to translate complex business requirements into scalable technical solutions. By optimizing our Hadoop and Spark ecosystems, you will directly influence the speed and reliability of our data-driven decision-making processes, ensuring that our stakeholders have access to high-quality, actionable insights that drive business https://jobeax.com/link/KgVLdM5cThorPSD1 Responsibilities : - Design and implement scalable data pipelines using Scala and Spark to ingest and transform large-scale datasets, ensuring high availability and performance for downstream analytics.- Optimize complex SparkSQL queries and Hadoop jobs to reduce processing latency, directly improving the efficiency of our data infrastructure.- Collaborate with engineering teams to integrate real-time data streams using Kafka, enabling low-latency data availability for critical business applications.- Maintain rigorous data quality standards by implementing automated testing and monitoring frameworks, ensuring reliability for all internal and external data consumers.- Partner with stakeholders to identify data bottlenecks and implement innovative engineering solutions that support the long-term scalability of our Big Data https://jobeax.com/link/1dkmzToQOwmnlB8O Skillset : - Demonstrated expertise in building distributed data systems using Hadoop, Spark, and Scala, with a deep understanding of performance tuning and memory management.- Proven ability to write complex SQL queries and develop efficient data models that support diverse analytical use cases.- Strong proficiency in Python for scripting and automation, complemented by hands-on experience in managing data flows through Kafka.- Exceptional communication skills with the ability to articulate technical concepts to non-technical stakeholders and work effectively within a collaborative, fast-paced environment.- A Bachelors or Masters degree in Computer Science, Engineering, or a related quantitative field, reflecting a strong foundation in data structures and algorithms.- Ability to thrive in a hybrid work model across our Chennai, Bangalore, Pune, or Mumbai offices, demonstrating high levels of self-motivation and adaptability to evolving project requirements.- 5 - 8 years of professional experience in Big Data Engineering, with a track record of delivering high-impact data solutions in enterprise environments. (ref:hirist.tech)
Senior Data Engineer Starts Immediately Competitive salary Competitive salary Experience We are looking for a Senior Data Engineer to join our team. In this role, you will design, build, and optimize scalable data pipelines on the Databricks Lakehouse Platform. You will partner with data science, analytics, and business ...
... platform ecosystem including Workbench, Connect, and integration patterns with R and Python environments Proficient in distributed computing technologies and big data analytics, including hands-on expertise with Python, Spark, SQL, and data transformation pipelines Strong understanding of MLOps/ModelOps principles, practices, ...
Job Title: Data Engineer - ML Training Data Pipeline Notice period: 0-30 Days Experience : 5+ Years Location: Hyderabad OR Pune We are looking for Data Engineer - ML Training Data Pipeline who can Build and maintain the data pipeline that transforms raw production traces into high-quality training datasets for LLM fine-tuning-ingestion, ...
... development. Familiarity with cloud platforms like AWS, Azure, or GCP. Knowledge of data engineering best practices and scalable architectures. Experience with Apache Spark or PySpark. Knowledge of Docker, Kubernetes, and CI/CD pipelines. Experience with data governance and metadata management. Exposure to machine learning data ...
... SQL, followed by expertise in data storage, modeling, cloud, data warehousing, and data lakes, ETL tools. The Data Engineer Lead will work closely with customer Data Architects, Data Scientists and BI Engineers to design and maintain scalable data models and pipelines. Data engineer lead will transfer requirements to the offshore ...
... you do in this role - Collaborates on designing, building, and maintaining scalable data pipeline architectures to ingest, process, and deliver high-quality data to business stakeholders. - Develops and optimizes batch-processing pipelines, ensuring collected data is structured and formatted for immediate, analytical consumption. ...
Project Role : Data Engineer Project Role Description : Design, develop and maintain data solutions for data generation, collection, and processing. Create data pipelines, ensure data quality, and implement ETL (extract, transform and load) processes to migrate and deploy data across systems. AWS AI Services Minimum 3 Year(s) ...
... Experience with system integration, implementation, upgrades, migrations, or technology modernization. - Good understanding of application, infrastructure, cloud, database, networking, or platform technologies relevant to the role. - Familiarity with automation, monitoring, DevOps, CI/CD, cloud platforms, or modern engineering ...
... members of IT, including business analysts, database administrators, and managers, to design and develop Cognos reporting capability- Exhibit a commitment to data quality by validating results against sources specified in the requirements- Develop and maintain stored procedures- Write and support queries against databases ...
... translating new tooling and best practices into scalable improvements across the enterprise. Requirements - Bachelor’s degree in Information Security, Computer Engineering, or a related field (or equivalent experience). - 5+ years of experience in data security, privacy engineering, or security engineering roles. - Hands-on ...
... documentation carrying the quality bar between reviews. Stack: Python, PostgreSQL, Django, server-rendered frontend (htmx). What You Will Own Core Responsibilities / Engineering Focus Deep third-party data ingestion Build reliable ingestion pipelines for external data sources Account for silently dropped webhooks through reconciliation ...
Project Role : Data Engineer Project Role Description : Design, develop and maintain data solutions for data generation, collection, and processing. Create data pipelines, ensure data quality, and implement ETL (extract, transform and load) processes to migrate and deploy data across systems. https://jobeax.com/link/6WcBICn0LM558SP0 ...
Company : Very big MNCRole : GCP Data EngineerExperience : 8 - 15 yrsNotice Period : 30 daysLocation : PAN INDIATech Stack :- GCP data engineer- Pyspark- Dataflow- Dataproc- BigQuery- AirflowKey Responsibilities :- Design, develop, and optimize endtoend data pipelines using PySpark, Dataflow, and Dataproc- Implement data ingestion, ...
... Engineering, or a related field. - 10+ years of experience in data engineering or a similar role. - Enterprise SaaS software solutions with high availability and scalability - Solution handling large scale structured and unstructured data from varied data sources - Experience in building and maintaining data platform systems ...
... Experience with cloud platforms (AWS / Azure / GCP).- Familiarity with MLOps tooling and practices (e.g., model versioning, CI/CD for ML, monitoring).- Experience with big data technologies (Spark, Databricks, etc.).- Bachelor's / Master's / PhD in Computer Science, Statistics, Mathematics, or a related quantitative https://jobeax.com/link/sEOksQVqU173JbV8 ...
... increase feature velocity of our developers Work closely with multiple teams to optimize the operations of their microservices - and improve the lives of the engineers within your product area engineering teams. Support the engineering teams within your product area by maintaining and executing a reliability roadmap of opportunities ...
... validate, and optimize AI models to improve accuracy, robustness, scalability, and business value. Work with structured and unstructured datasets to prepare data for AI model training, validation, and inference. Design and optimize data pipelines supporting AI and machine learning workflows. Perform feature engineering, ...
CDP ETL & Database Engineer Primary Job Responsibilities: The CDP ETL & Database Engineer will specialize in architecting, designing, and implementing solutions that are sustainable and scalable. The ideal candidate will understand CRM methodologies, with an analytical mindset, and a background in relational modeling in ...
... Airflow workflows and production data science infrastructureIDEAL PROFILE:Looking for candidates with strong hands-on experience in:- Data Operations- Production Data Engineering- Data Platform Support- Data Pipeline Monitoring & Troubleshooting- Not looking for purely development-focused Data Engineers or DevOps/SRE https://jobeax.com/link/V6ZwWogXLIs5s3d8 ...